A method for coordinated scheduling of power distribution network sources and loads considering multiple types of flexible resources

By establishing a multi-type flexible resource collaborative scheduling model and the PDQN algorithm, the scheduling of fixed energy storage, mobile energy storage, and demand response resources is optimized, solving the problem of low grid operation efficiency and realizing the efficient consumption of new energy and the improvement of grid economy.

CN121507853BActive Publication Date: 2026-03-31SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing scheduling methods fail to fully integrate the spatiotemporal differences of demand response mechanisms and flexible resources such as mobile energy storage devices, resulting in low grid operating efficiency, inability to effectively regulate the intermittency and load fluctuations of photovoltaic power generation, and inability to optimize resource allocation.

Method used

A collaborative scheduling model for multiple types of flexible resources is established. By combining fixed energy storage, mobile energy storage devices and demand response resources, the scheduling strategy is optimized using Markov decision process and parameterized deep Q network algorithm (PDQN) to achieve simultaneous optimization of discrete and continuous actions.

Benefits of technology

It has improved the system's ability to regulate under load fluctuations, promoted the consumption of new energy sources, reduced grid losses, improved the economic efficiency of grid operation, and achieved broader and more flexible grid support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121507853B_ABST
    Figure CN121507853B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of new energy optimization scheduling, and discloses a power distribution network source-load collaborative scheduling method considering multiple types of flexible resources. The method considers the diversity and space-time difference of flexible resources, introduces multiple types of flexible resources, enriches the resource types of source-load bilateral collaboration, and expands the feasibility and effectiveness of source-load collaboration in different scenarios. In addition to considering fixed energy storage resources, the method also introduces a demand response mechanism, integrates user-side flexible resources into the scheduling process, and at the same time, considers the space flexibility and time sequence adjustment capacity of mobile energy storage equipment, and builds a source-load collaborative scheduling model considering multiple types of flexible resources. Furthermore, aiming at the space-time characteristics that the mobile energy storage equipment needs to simultaneously decide the destination and charging and discharging power, the method introduces a PDQN algorithm based on a mixed action space. The application can effectively promote source-load collaboration, reduce network loss, promote new energy consumption, and improve the economic efficiency of power distribution network operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of new energy optimization scheduling technology, specifically to a distribution network source-load coordinated scheduling method that considers multiple types of flexible resources. Background Technology

[0002] With the accelerated adjustment of China's energy structure, building a new power system with new energy sources as the mainstay has become a key path to achieving carbon emission reduction. The dispatching of flexible resources in the distribution network, especially the coordinated dispatching of multiple types of flexible resources, has become an important technology for optimizing grid operation, improving the absorption capacity of new energy sources, and enhancing system economy. Flexible resources include stationary energy storage devices, demand response mechanisms, and mobile energy storage devices. They can effectively enhance the grid's regulation capacity, ensure stable system operation, and are indispensable stabilizers for the new power system. They are also globally recognized as the most effective means of supporting the development of new energy sources such as photovoltaics.

[0003] However, given the intermittency and uncertainty of photovoltaic power generation and the volatility of system load, existing dispatching methods in the literature mostly focus on traditional fixed energy storage and generation dispatching strategies. They fail to fully explore the potential of other flexible resources on the load side of the distribution network besides energy storage. In particular, the spatiotemporal differences and dispatching capabilities of flexible resources such as demand response mechanisms and mobile energy storage devices have not been effectively integrated. This results in significant limitations in optimizing resource allocation and improving grid operating efficiency of existing dispatching methods, failing to fully leverage the regulatory role of flexible resources in different scenarios. Therefore, how to optimize source-load coordinated dispatching strategies based on multiple flexible resources and improve the dispatching efficiency of the power grid system has become a crucial problem urgently needing to be solved in the field of power system dispatching. Summary of the Invention

[0004] To address the aforementioned problems, the present invention aims to provide a power distribution network source-load coordinated scheduling method that considers multiple types of flexible resources. This method can fully leverage the regulatory role of flexible resources in different scenarios, effectively promote source-load coordination, reduce network losses, promote the consumption of new energy sources, and improve the economic efficiency of power distribution network operation. The technical solution is as follows:

[0005] A distribution network source-load coordinated scheduling method considering multiple types of flexible resources includes the following steps:

[0006] Step 1: Consider multiple types of flexible resources, including stationary energy storage devices, mobile energy storage devices, and demand response resources. Conduct in-depth analysis on the unique characteristics of different devices, establish mathematical models for each type of flexible resource, and establish a collaborative scheduling model for multiple types of flexible resources in conjunction with source-side equipment including coal-fired power units and distributed wind power.

[0007] Step 2: Determine the objective function and constraints based on the multi-type flexible resource collaborative scheduling model. The objective function considers the participation of multiple types of flexible resources in the optimal scheduling, covering the demand response compensation cost and the call cost of mobile energy storage.

[0008] Step 3: Establish a parameterized deep Q-network algorithm framework based on Markov decision process to simultaneously optimize the discrete actions of mobile energy storage transfer nodes and the continuous actions of demand response power in scheduling.

[0009] Step 4: Based on the parameterized deep Q-network algorithm framework, solve for the optimal day-ahead scheduling strategy of the constructed multi-type flexible resource collaborative scheduling model.

[0010] The beneficial effects of this invention are:

[0011] 1) Based on fixed energy storage, this invention introduces a demand response mechanism, incorporates user-side flexible resources into the scheduling process, improves the system's adjustment capability under load fluctuations, and promotes the consumption of new energy sources;

[0012] 2) This invention takes into account the spatial flexibility and timing adjustment capability of mobile energy storage devices during the scheduling process, breaking through the limitations of traditional fixed energy storage and achieving wider and more flexible support for the power grid.

[0013] 3) This invention utilizes Markov decision processes to establish a parameterized deep Q-network (PDQN) algorithm, which enables simultaneous optimization of discrete actions such as mobile energy storage transfer nodes and continuous actions such as demand response power during scheduling. This avoids the accuracy loss caused by action discretization in traditional methods, thereby improving the refinement and accuracy of scheduling strategies. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the power-transportation network coupling relationship of the present invention.

[0015] Figure 2 This is a schematic diagram of the PDQN algorithm structure of the present invention.

[0016] Figure 3 This is a schematic diagram of the power distribution network source-load coordinated scheduling framework based on the PDQN algorithm of this invention.

[0017] Figure 4 This is the improved IEEE 33-node distribution network topology used in this invention.

[0018] Figure 5 This is a comparison chart of the algorithm optimization performance of the scheduling control model in this invention.

[0019] Figure 6The present invention considers the load change curves before and after demand response.

[0020] Figure 7(a) is a schematic diagram of the charging and discharging of fixed energy storage after applying the solution of the present invention.

[0021] Figure 7(b) is a schematic diagram of the charging and discharging of mobile energy storage after applying the solution of the present invention.

[0022] Figure 7(c) is a schematic diagram of the State of Charge (SOC) of stationary energy storage after applying the scheme of the present invention.

[0023] Figure 7(d) is a schematic diagram of a mobile energy storage SOC after applying the solution of the present invention.

[0024] Figure 8(a) is a schematic diagram of the mobile energy storage path at 00:00 after applying the solution of the present invention.

[0025] Figure 8(b) is a schematic diagram of the mobile energy storage path from 01:00 to 09:00 after applying the solution of the present invention.

[0026] Figure 8(c) is a schematic diagram of the mobile energy storage path from 10:00 to 17:00 after applying the solution of the present invention.

[0027] Figure 8(d) is a schematic diagram of the mobile energy storage path from 18:00 to 20:00 after applying the solution of the present invention.

[0028] Figure 8(e) is a schematic diagram of the mobile energy storage path at 21:00 after applying the solution of the present invention.

[0029] Figure 8(f) is a schematic diagram of the mobile energy storage path from 22:00 to 23:00 after applying the solution of the present invention.

[0030] Figure 9 The graph shows the change in wind power output before and after applying the solution of this invention.

[0031] Figure 10 The graph shows the output change of conventional units before and after applying the solution of this invention.

[0032] Figure 11 The graph shows the network loss changes before and after applying the solution of this invention. Detailed Implementation

[0033] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0034] Reference Figures 1 to 3 As one embodiment of the present invention, a method for coordinated scheduling of power distribution network sources and loads considering multiple types of flexible resources is provided, including:

[0035] S1: Considering multiple types of flexible resources, including stationary energy storage devices, mobile energy storage devices, and demand response resources, we conduct in-depth analysis of the unique characteristics of different devices, establish mathematical models corresponding to each flexible resource, and, together with source-side devices such as coal-fired units and distributed wind power, establish a source-load coordinated scheduling model that considers multiple types of flexible resources.

[0036] Specifically, the flexible resources studied in this invention include stationary energy storage devices, mobile energy storage devices, and demand response resources. By scheduling the above-mentioned load-side flexible resources, in conjunction with source-side equipment such as coal-fired power units and distributed wind power, a multi-type flexible resource collaborative scheduling model is established to achieve source-load coordination in the distribution network. Its operational framework is as follows: Figure 1 As shown.

[0037] Energy storage devices can formulate charging and discharging plans based on the predicted output of new energy sources, enhancing the controllability of electricity, promoting the consumption of new energy, and improving the economic efficiency of dispatch. The stationary energy storage model is shown in the following equation:

[0038] ;

[0039] ;

[0040] ;

[0041] ;

[0042] In the formula, This represents the energy storage state of the fixed energy storage system at time t. and These represent the stationary energy storage charging efficiency and discharging efficiency, respectively. and These represent the upper and lower limits of fixed energy storage, respectively; This represents the fixed energy storage charging power at time t; This represents the fixed energy storage discharge power at time t; and These represent the upper limits of fixed energy storage charging power and discharging power, respectively. This represents the length of the time period.

[0043] The mobile energy storage studied in this invention is a mobile energy storage vehicle equipped with energy storage devices. It has mobility and can perform charging and discharging operations at nodes in the power grid with charging interfaces. Its scheduling involves two aspects: power and path. On the one hand, it is necessary to optimize the charging and discharging strategy based on power grid information; on the other hand, it is necessary to calculate the optimal path to the target charging node by combining traffic network information.

[0044] In graph theory, Dijkstra's algorithm is an efficient method for finding the shortest path. It uses a greedy strategy to iteratively update node distances and calculates the shortest path for a mobile energy storage device from its current location i to its target node j. The calculation formula is as follows:

[0045] ;

[0046] In the formula, The optimal path matrix; This represents the shortest distance from node i to node j at time t; ; Let t be the total number of nodes at time t.

[0047] The mobile energy storage operation model needs to meet multi-dimensional mobility constraints. Spatially, the mobile energy storage vehicle can only move within a preset range and can only perform charging and discharging operations at nodes equipped with charging interfaces. Temporally, at any given time t, the mobile energy storage vehicle can only be in one of three states: charging, discharging, or moving. In summary, the constraints for mobile energy storage are:

[0048] ;

[0049] ;

[0050] ;

[0051] In the formula, To indicate the operating status of mobile energy storage, 1 indicates that the mobile energy storage is moving, and 0 indicates that the mobile energy storage is stationary. If connected to a charging port, it can be charged and discharged. This indicates a flag indicating the connection of mobile energy storage to the distribution network. If the mobile energy storage is connected to the grid at node i at time t, then... The value is 1 if the mobile energy storage is in an off-grid state, and 0 if the mobile energy storage is in an off-grid state. Mobile energy storage cannot be in both grid-connected and mobile states at the same time. Let h be the set of nodes in the transportation network where mobile energy storage charging ports are installed; h represents the travel time of the mobile energy storage; and T is the total duration of the operation cycle. This represents the shortest time required for mobile energy storage to move from node i to node j in the transportation network, which can be obtained from the shortest path calculated above.

[0052] Apart from route optimization, the charging and discharging cost model for mobile energy storage vehicles is the same as that for stationary energy storage vehicles.

[0053] User loads are located at nodes in the distribution network system. Based on their ability to participate in demand-side response, user loads can be categorized into rigid loads and flexible loads. The power grid can adjust flexible loads by adjusting electricity prices or signing contracts, ensuring that user electricity consumption aligns with the grid's needs. Flexible loads participating in demand response can be further divided into loads that can be reduced or transferred.

[0054] Transferable loads alter the original load curve by adjusting users' electricity consumption times. To ensure that users' normal production is not affected, when users participate in demand response as transferable loads, the total load must remain constant throughout the entire dispatch cycle. The model is as follows:

[0055] ;

[0056] ;

[0057] ;

[0058] In the formula, and These represent the load before and after the load transfer adjustment at time t, respectively. and These correspond to the load transferred in and the load transferred out based on demand response at time t, respectively. This represents the maximum transfer margin of the load.

[0059] Reduceable load refers to a portion of the load that can be flexibly reduced based on grid demand within a certain timeframe. The key characteristic of reduceable load is that it responds to grid demand by reducing equipment power or temporarily interrupting equipment power, without affecting the basic functions of the equipment. The model for reduceable load is as follows:

[0060] ;

[0061] ;

[0062] ;

[0063] ;

[0064] ;

[0065] In the formula, The power before load reduction at time t; Let t be the amount of load reduction at time t; The maximum load that can be reduced; This represents the maximum number of cuts. It is a 0-1 variable, representing the load reduction state. It is 0 when the load does not respond to the load reduction and 1 when the load responds to the load reduction. and These represent the minimum load reduction time and the maximum load reduction time, respectively.

[0066] Adjusting all flexible loads through a demand response mechanism can effectively balance electricity supply and demand, alleviate grid pressure, and support the efficient integration of renewable energy. The model for the total load after demand response adjustment is as follows:

[0067] ;

[0068] In the formula, This indicates the adjusted total load.

[0069] S2: Determine the objective function and constraints based on the multi-type flexible resource collaborative scheduling model. The objective function considers the participation of multiple types of flexible resources in the optimal scheduling, covering the demand response compensation cost and the call cost of mobile energy storage.

[0070] Specifically, this invention studies the source-load coordinated scheduling problem of distribution network in the current stage. Its optimization scheduling objective is to minimize the total operating cost of the distribution network, including wind curtailment cost, thermal power unit generation cost, demand response compensation cost, energy storage operation cost, and network loss cost. A flexible resource optimization scheduling model is established with the goal of minimizing the total cost.

[0071] ;

[0072] In the formula, Let t be the operating cost of the thermal power unit. Let t be the demand response cost; The loss cost of fixed energy storage at time t; Let be the cost of calling up mobile energy storage at time t; Let t be the network loss cost at time t; This indicates the cost of wind power.

[0073] The operating cost of a thermal power unit is shown in the following formula:

[0074] ;

[0075] In the formula, , and Let be the power generation cost coefficient of the j-th thermal power unit; Let t be the output power of the thermal power units under node j at time t.

[0076] Demand response costs consist of the costs of calling both scalable and transferable loads:

[0077] ;

[0078] In the formula, For unit call cost that can reduce load; The unit call cost for transferable load.

[0079] The losses of stationary energy storage are related to its charging and discharging power, as follows:

[0080] ;

[0081] In the formula, This represents the unit charge / discharge power loss cost of stationary energy storage.

[0082] The cost of deploying mobile energy storage consists of two parts: charging and discharging costs and relocation costs.

[0083] ;

[0084] In the formula, Cost per unit charge / discharge power loss for mobile energy storage; The cost per kilometer for mobile energy storage; The distance traveled by the mobile energy storage vehicle from time t-1 to time t; and These represent the charging power and discharging power of the mobile energy storage at time t, respectively.

[0085] Wind power cost It consists of wind curtailment cost and wind turbine operating cost, and the specific formula is as follows:

[0086] ;

[0087] In the formula, The unit power generation cost of wind turbines; Let t be the wind power generation capacity. Cost per unit of wind curtailment; This refers to the amount of wind that is abandoned.

[0088] In addition to meeting the basic constraints for equipment operation, power grid operation also needs to meet constraints such as system unit constraints, power flow balance constraints, and node voltage constraints.

[0089] The system unit constraints are as follows:

[0090] ;

[0091] ;

[0092] In the formula, The gradient rate of thermal power units; Let t be the reactive power output of the thermal power units under node j at time t. , , and These represent the minimum active power output, minimum reactive power output, maximum active power output, and maximum reactive power output of thermal power unit j, respectively. , These represent the actual and maximum wind power output at time t, respectively.

[0093] The power flow balance constraints are as follows:

[0094] ;

[0095] ;

[0096] In the formula, and These represent the active power injected into node i by all generators and energy storage devices at time t. Let t be the reactive power injected into node i at time t; and These represent the adjusted active load and reactive power of node i at time t; Let be the voltage magnitude of node i at time t; Let be the voltage phase angle difference on line ij at time t; and Let be the conductance and susceptance between nodes ij and ij, respectively.

[0097] The voltage constraints are as follows:

[0098] ;

[0099] In the formula, and These represent the lower and upper voltage limits of node i, respectively.

[0100] S3: A parameterized deep Q-network algorithm framework based on Markov decision processes is established, which realizes the simultaneous optimization of discrete actions such as mobile energy storage transfer nodes and continuous actions such as demand response power in scheduling. The algorithm diagram is shown below. Figure 2 As shown. Among them, Layer K Refers to the first K Layered neural networks, Q ( w This represents the action value network. x (θ) represents the parametric action network. Q (1) to Q ( K This represents the Q-network's estimation of the action value for different discrete actions. argmax ( Q ( i )) indicates selecting the action with the largest Q value from multiple actions.

[0101] Specifically, this invention utilizes the DRL algorithm to solve the source-load coordinated scheduling problem of a distribution network, which requires first representing the scheduling strategy solution process as a Markov decision process. In the day-ahead phase, the Markov decision process is based on the system's global state space. Action space Global reward function R, system state transition probability and reward discount factor Composition. The specific structure is as follows:

[0102] State space S: The day-ahead control target is related to the operating state of the distribution network. It is necessary to obtain the state changes of the distribution network at various time periods based on the decision results, and use the available state information as the state space for reinforcement learning. In the reinforcement learning task solved by this invention, the state information includes power generation. ,Voltage Energy storage state of charge ,load Time t and the location of the mobile energy storage The state space is represented as:

[0103] ;

[0104] Action space In the current phase, stationary energy storage, mobile energy storage, and demand response loads are considered as decision-making objects for the intelligent agent. For different types of equipment, the grid dispatch center should assess the grid operating status based on status information and rationally allocate the power output of each energy storage unit and the user's response, generating corresponding control actions. Simultaneously, as mobile energy storage is a movable device, the intelligent agent also needs to determine the node location of the mobile energy storage unit. Therefore, the action space of the intelligent agent is defined. The mixed action space, consisting of discrete and continuous actions, is shown in the following equation:

[0105] ;

[0106] in, Let t be the power of the fixed energy storage. Power for mobile energy storage; The amount of power change in response to user demand; The location of mobile energy storage for decision-making by intelligent agents.

[0107] Reward R: The scheduling result must not only include the objective function to be optimized, but also consider the penalty function term for violating the corresponding constraints. Similarly, when the agent violates the corresponding constraints, the agent is given an appropriate penalty. Simultaneously, the objective function is adaptively adjusted, transforming the problem of minimizing the total system cost into a reward maximization form of reinforcement learning. Therefore, the reward function r(t) of the agent can be established as follows:

[0108] ;

[0109] In the formula, Let be the penalty coefficient for the m-th constraint. , These are the actual value and boundary value of the m-th constraint parameter, respectively.

[0110] Based on the aforementioned Markov decision model, the flexible resource scheduling in S1 can be transformed into a DRL framework and solved through reinforcement learning. The distribution network flexible resource model of this invention includes both continuous actions such as energy storage charging and discharging, and demand response power, as well as discrete actions such as mobile energy storage transfer nodes.

[0111] However, traditional reinforcement learning methods have limitations when dealing with such mixed action spaces. To address this issue, this invention introduces the PDQN (Parameterized Deep Q-Network) algorithm. PDQN decouples discrete actions from continuous actions through a parameterization mechanism, thereby achieving simultaneous optimization of both, effectively reducing accuracy loss and improving the algorithm's optimization performance. Compared with traditional methods, this algorithm is better suited to the complex scenarios of flexible resource scheduling in distribution networks.

[0112] The PDQN structure consists of a deterministic policy network and a Q-value network. The deterministic policy network generates continuous actions based on the environment state, while the Q-value network concatenates the state and continuous actions for evaluation and selects the optimal discrete action based on the evaluation result. The algorithm's operation can be summarized as follows: the agent first obtains the state from the environment and inputs it into the policy network to generate continuous actions; then, the state and continuous actions are input together into the Q-value network to obtain the action value and select the optimal discrete action.

[0113] S4: Solve the optimal day-ahead scheduling strategy for the constructed multi-type flexible resource collaborative scheduling model based on the parameterized deep Q-network algorithm framework. This includes: setting the number of training rounds, setting one hour as the time step of the day-ahead scheduling task, and setting 24 time steps as one training round.

[0114] A framework for solving flexible resource scheduling strategies based on PDQN, such as... Figure 3 As shown, where Represents state variables, Represents action variables, Indicates a reward. Indicates taking action The state variables after that, As a termination marker, Represents network parameters, This represents the action selection probability distribution output by the policy network. Represents the probability distribution of output action selection The network parameters of the neural network, Represents the probability distribution of output action selection The network parameters of the neural network, Indicates the neural network's first... Features of hidden layers The algorithm represents the policy gradient, TD error refers to the temporal difference error, and ResNet refers to the residual network. Its update process is divided into Q-value network update and deterministic policy network update.

[0115] For the PDQN algorithm with a hybrid action space, the action space is divided into discrete actions and continuous actions, denoted as k and x, respectively. k For discrete action k, its update process depends on the fitting of the Q-value. The network update is guided by the magnitude of the Q-value, which can be expressed as:

[0116] ;

[0117] In the formula, The continuous actions are derived from the state at time t+1; and These represent the state, discrete action, and continuous action at time t, respectively. The reward at time t; For the set of action variables, represents the set of state variables; E represents the expected value.

[0118] Since the output of continuous actions is obtained from a deterministic policy network, the parameters of the deterministic policy network are used. Approximately fit the continuous action x k Then, according to the Bellman equation, the target Q value can be obtained as follows:

[0119] ;

[0120] In the formula, These are the update parameters for the Q network.

[0121] The Q-network is updated by minimizing the loss function, which is as follows:

[0122] ;

[0123] When updating a deterministic policy network, the purpose of the update is to find the optimal network parameters. ,exist If the optimal Q-value can be obtained when the condition is fixed, then the loss function for updating the deterministic policy network is defined as:

[0124] ;

[0125] In the formula, These are the update parameters at time t; This represents the total number of actions.

[0126] Reference Figures 4-6 Figures 7(a)-7(d), 8(a)-8(f) and Figures 9-11 This is the second embodiment of the present invention, which provides a distribution network source-load coordinated scheduling method that considers multiple types of flexible resources. In order to verify the beneficial effects of the present invention, a simulation experiment is conducted for scientific demonstration.

[0127] Specifically, a simulation experiment was conducted using an improved IEEE 33-bus system as the analysis object. A thermal power unit is connected to node 1, and distributed wind turbines are connected to node 9. The fixed energy storage parameters are 500 kWh / 250 kW, located at node 20. The mobile energy storage vehicle carries an energy storage capacity of 500 kWh / 250 kW, initially located at node 1. Mobile energy storage charging stations are set up at nodes 5, 9, and 23, respectively. The initial SOC of both fixed and mobile energy storage is 0.2, with upper and lower limits of 0.1 and 0.95, respectively. The grid node topology is shown below. Figure 4 As shown in Table 1, the wind power generation data comes from an actual system in a region of Jiangsu Province, China in 2022, and is multiplied by a coefficient to adapt to the system capacity. The load node information and demand response time periods participating in demand response are shown in Table 1. This embodiment selects a 29-node transportation network as the traffic environment example, and its coupling relationship with the IEEE 33-node system is shown in Table 2.

[0128] Table 1 Demand Response Period and Response Information

[0129] .

[0130] Table 2. Correspondence between transportation network nodes and power grid nodes

[0131] .

[0132] This embodiment uses PyTorch as the framework, employs Python 3.7 programming, and combines the PDQN algorithm to solve the scheduling strategy. The reinforcement learning algorithm is trained for 5500 epochs with a batch size of 128, and the Adam optimizer is used to update network parameters. To verify the effectiveness of the proposed source-load cooperative scheduling method, the following four scenarios are set up for test examples:

[0133] Scenario 1: Consider replacing mobile energy storage with fixed energy storage of equivalent configuration, without considering demand response.

[0134] Scenario 2: Consider replacing mobile energy storage with fixed energy storage of equivalent configuration, taking demand response into account.

[0135] Scenario 3: Consider stationary and mobile energy storage, but do not consider demand response.

[0136] Scenario 4: Simultaneously consider the source-load coordinated scheduling of demand response, mobile energy storage, and stationary energy storage.

[0137] To verify the advantages of the algorithm proposed in this invention, scenario 4 was selected as the test scenario in this embodiment. Its optimization performance was compared with that of the PDQN algorithm constructed from a fully connected neural network, the TD3 algorithm with continuous action discretization, and the DDPG (Deep Deterministic Policy Gradient) algorithm. The results are as follows: Figure 5 The results are shown, with the dark solid line representing the mean reward and the shaded area representing the variance of the reward. As can be seen from the figure, the proposed PDQN algorithm outperforms the comparative algorithms in terms of optimization performance. Compared to traditional reinforcement learning algorithms TD3 and DDPG, the algorithm of this invention is slower in the first 1600 training rounds. This is because TD3 and DDPG use an exploration method that adds noise to continuous actions, resulting in higher exploration efficiency and faster optimization. The algorithm of this invention needs to coordinate the greedy exploration strategy for discrete actions with the noisy exploration strategy for continuous actions, making it difficult to achieve consistency between the two exploration methods in the early stages of training. During the mid-training phase, the algorithm proposed in this invention gradually learns the coupling relationship between discrete and continuous actions (the coupling between movement and charging / discharging power in mobile energy storage), and its optimization performance gradually improves. It surpasses the DDPG and TD3 algorithms at 1100 and 2500 rounds respectively, and gradually converges at 4500 rounds, finding the optimal strategy and converging to around -9.45. In contrast, DDPG and TD3, due to their method of fitting the mixed action space using continuous action discretization, have limited optimization capabilities and gradually fall into local optima, ultimately converging only to around -9.57 and -9.55. Furthermore, due to their lack of explicit modeling of mixed actions, their stability in the later stages is poor. Therefore, the algorithm proposed in this invention outperforms traditional algorithms in both stability and convergence.

[0138] The scheduling results of Scenario 4 (the collaborative scheduling method proposed in this invention) and Scenario 1 (the traditional scheduling method) were compared and analyzed to verify the advantages of the proposed method in reducing network losses and facilitating the consumption of new energy.

[0139] Based on the multi-type flexible collaborative scheduling model proposed in this invention, a comparative analysis of load changes before and after demand response is conducted, and the results are as follows: Figure 6As shown, demand response-based incentive scheduling achieves coordination between the load side and the source side, realizing the effect of "peak shaving and valley filling." Specifically, before implementing demand response, during the midday peak period of 10:00-16:00 and the evening peak period of 21:00-22:00, wind power output reached its upper limit, increasing the pressure on the power grid for peak regulation. Conventional unit output was used to achieve power balance, but this resulted in additional network losses and operating costs. During the period of 05:00-09:00 and 23:00, the mismatch between load and wind power output led to a large amount of wind curtailment, reducing the economic efficiency of operation.

[0140] After implementing the demand response mechanism, the load transfer mechanism shifted the load during the midday and evening peak hours to periods of higher wind speeds (05:00-09:00 and 23:00), promoting the consumption of renewable energy. Simultaneously, in conjunction with load reduction adjustments, a load gap of nearly 0.2MW was achieved during peak load periods, effectively reducing the pressure on the power grid for peak shaving. By providing incentives to users through the demand response mechanism, user-side loads were guided to change their energy consumption patterns to meet the grid's adjustment needs, achieving effective coordination between power generation and load.

[0141] The energy storage charging and discharging situation after the scheduling of scenario 4 is shown in Figures 7(a)-7(d). Figures 7(a) and 7(b) correspond to the charging and discharging power of fixed energy storage and mobile energy storage, respectively. Figures 7(c) and 7(d) show the SOC changes of fixed energy storage and mobile energy storage.

[0142] As shown in Figures 7(a)-7(d), the charging and discharging of both stationary and mobile energy storage systems follow the changes in load and renewable energy output, exhibiting a convergence of actions. During the period from 00:00 to 09:00, wind power output increases, and both energy storage systems charge to absorb wind power and reach the upper limit of their State of Charge (SOC). During the period from 10:00 to 17:00, the load gradually increases, increasing grid pressure. At this time, both energy storage systems choose to discharge instead of generating thermal power to reduce operating costs. After discharging to the lowest SOC value, during the period from 18:00 to 20:00, the energy storage systems choose to charge, partly due to the increased wind power requiring renewable energy absorption, and partly due to the impending increase in load requiring energy storage to cope with the evening peak. At 21:00, during the evening peak period, renewable energy output cannot cover load demand, and the energy storage systems choose to discharge to reduce grid pressure. During the period from 22:00 to 23:00, the load decreases, and renewable energy output increases. At this time, the energy storage systems again choose to charge to reduce wind curtailment costs.

[0143] Figures 8(a)-8(f) illustrate the movement path of the mobile energy storage vehicle during a 24-hour scheduling cycle. Figures 8(a)-8(f) show the node location and planned movement path of the energy storage vehicle at each time period. As shown in the figures, at the initial scheduling time of 00:00, the mobile energy storage vehicle is at initial node 1. After scheduling begins, the mobile energy storage vehicle senses the renewable energy output and moves to node 9, remaining there until 09:00. This is because node 9 is adjacent to renewable energy nodes, and moving to this node allows for more effective local consumption of renewable energy, reducing transmission costs caused by grid losses. From 10:00 to 17:00, renewable energy output decreases, load increases, and grid pressure increases. To respond to the grid's peak-shaving needs, the mobile energy storage vehicle moves to node 5, which is closer to the load center, to discharge, replacing thermal power output and reducing transmission grid loss costs. From 18:00 to 23:00, the mobile energy storage vehicle again moves its location according to changes in renewable energy output. Between 18:00 and 20:00, wind power output increased again, so the mobile energy storage vehicle moved back to node 9 to recharge in order to reduce wind curtailment. At 21:00, the power grid experienced a peak load, and wind power output decreased. To reduce the pressure on the power grid for peak shaving and reduce grid losses, the mobile energy storage vehicle returned to the heavily loaded node 5 and performed a discharge operation at the charging station. Between 22:00 and 23:00, after the peak load, wind power output increased, but there was still redundant wind power that could not be absorbed. At this time, the mobile energy storage vehicle moved back to node 9 to recharge in order to improve the utilization rate of renewable energy.

[0144] Comparison of wind power output changes in scenarios 1 and 4 (e.g.) Figure 9 As shown. By Figure 9 As can be seen, after adopting the method proposed in this invention, wind power output increased for most of the nighttime periods during 00:00-09:00, 17:00-19:00, and 23:00. This is mainly due to the demand response mechanism and the transferability characteristics of energy storage devices, which shift peak loads to this period. After one day of scheduling, the amount of wind curtailment decreased from 8.56 MWh to 7.87 MWh, a reduction of 0.69 MWh, and the curtailment rate decreased by 2.3%, indicating more effective utilization of wind power resources.

[0145] Figure 10The changes in conventional turbine output under scenarios 1 and 4 are illustrated. To accommodate the increased wind power output and ensure effective wind power absorption, conventional turbines operate at minimum power during the early morning and at 23:00. Between 10:00-16:00 and 21:00-22:00, the output of conventional turbines in scenario 4 is lower than that in scenario 1. This is because, with the coordination of energy storage and demand response, the net load decreases, replacing some of the thermal power output. This reflects that the method proposed in this invention can effectively utilize load-side resources and reduce operating costs in conjunction with source-side demand.

[0146] Figure 11 The network loss situation under scenarios 1 and 4 was compared. As shown in the figure, the network loss in scenario 4 was significantly reduced during high-load periods (10:00-16:00 and 21:00-22:00). This is because the mobile energy storage moved to node 5, reaching a location closer to the load center to perform on-site discharge operations, shortening the power transmission path. Simultaneously, in conjunction with the demand response mechanism, some peak loads were reduced or transferred, further reducing losses. Furthermore, because the mobile energy storage moved to charge near renewable energy generation nodes, network losses were also reduced during other time periods (such as 00:00-09:00 and 23:00), resulting in a total reduction of 0.75 MWh of network loss over the complete dispatch cycle. Through the effective coordination of multiple types of flexible resources, the network loss level of the distribution network can be significantly improved.

[0147] To compare the effectiveness of the proposed method, simulation tests were conducted in four scenarios, and the economic indicators under different scenarios were calculated. The results are shown in Table 3. According to Table 3, compared with Scenario 1, which only adds a demand response mechanism, Scenario 2's total cost decreased by 179.50 yuan. Although the demand response cost increased, the demand response mechanism responded well to the grid's demand, playing a role in "peak shaving and valley filling." Scenario 3 uses mobile energy storage as a flexible dispatch resource. Compared with the independent optimization of fixed energy storage in Scenario 2, the network loss cost decreased by 380.62 yuan, and the total cost decreased by 282.76 yuan. Scenario 4 combines the good network loss reduction characteristics of mobile energy storage with the load adjustability characteristics of the demand response mechanism. The total dispatch cost decreased by 138.06 yuan compared with Scenario 3, achieving further optimization. Compared with Scenario 1, which did not add any flexible resources, the total cost decreased by 600.32 yuan. This demonstrates that the power grid source-load coordinated scheduling method proposed in this invention, which considers multiple types of flexible resources, can achieve effective coordination between the source and load sides. It designs optimal scheduling strategies for mobile energy storage, demand response loads, fixed energy storage, and power generation equipment, providing flexibility support for the power grid, increasing the absorption of new energy, reducing system network losses, lowering system operating costs, and enhancing operating efficiency.

[0148] Table 3 Economic Outcomes in Different Scenarios

[0149] .

[0150] As can be seen from the above embodiments, this invention proposes a source-load coordinated scheduling method for distribution networks that considers multiple types of flexible resources. This method introduces various flexible resources, enriching the resource types for source-load bilateral coordination, thereby expanding the feasibility and effectiveness of source-load coordination in different scenarios. In addition to considering fixed energy storage resources, a demand response mechanism is introduced to incorporate user-side flexible resources into the scheduling process; simultaneously, the spatial flexibility and temporal adjustment capabilities of mobile energy storage devices are taken into account, constructing a source-load coordinated scheduling model that considers multiple types of flexible resources. Furthermore, considering the spatiotemporal characteristics of mobile energy storage devices needing to simultaneously decide on destination and charging / discharging power, a PDQN algorithm based on a hybrid action space is introduced. The method proposed in this invention can effectively promote source-load coordination, reduce network losses, promote the consumption of new energy sources, and improve the economic efficiency of distribution network operation.

Claims

1. A power distribution network source-load collaborative scheduling method considering multiple types of flexible resources, characterized in that, The method comprises the following steps: Step 1: considering multiple types of flexible resources including fixed energy storage devices, mobile energy storage devices and demand response resources, in-depth analysis is made according to the unique characteristics of different devices, and a mathematical model corresponding to each flexible resource is established, and a multi-type flexible resource coordinated scheduling model is established in cooperation with source-side devices including coal-fired generating units and distributed wind power; Step 2: determining a target function and a constraint condition according to the multi-type flexible resource coordinated scheduling model, wherein the target function considers the participation of multiple types of flexible resources in optimized scheduling and covers demand response compensation cost and mobile energy storage calling cost; Step 3: establishing a parameterized deep Q network algorithm framework based on Markov decision process to simultaneously optimize discrete actions including mobile energy storage transfer nodes and continuous actions including demand response power in scheduling; Step 4: based on the parameterized deep Q network algorithm framework, the optimal day-ahead scheduling strategy of the multi-type flexible resource coordinated scheduling model is solved; The target function in step 2 is as follows: The target of optimized scheduling is to minimize the total operation cost of the power distribution network, and the total operation cost includes wind power abandonment cost, power generation cost of thermal power units, demand response compensation cost, energy storage operation cost and network loss cost, and a flexible resource optimized scheduling model is established with the target of minimizing the total cost: ; Ctotalis the total operation cost; Ctis the operation cost of thermal power units at time t; Ctis the demand response cost at time t; Ctis the wear cost of fixed energy storage at time t; Ctis the calling cost of mobile energy storage at time t; Ctis the corresponding network loss cost at time t; Ctis the wind power cost; Operation cost of thermal power generating unit As shown in the following formula: ; In the formula, , and is the power generation cost coefficient of the jth thermal power unit; is the output power of the thermal power unit t under the jurisdiction of node j at time t. Demand response cost Consists of the cost of invoking both curtailable and transferable loads: ; In the formula, is the unit invocation cost of the load that can be reduced; is the unit invocation cost of the load that can be transferred. Loss cost of fixed energy storage Corresponding to the charge and discharge power thereof, as follows: ; In the formula, represents the unit charge-discharge power loss cost of fixed energy storage; Call cost of mobile energy storage It consists of two parts, namely the charging and discharging cost and the migration cost: ; In the formula, is the unit charging and discharging power loss cost of mobile energy storage; is the mobile cost per kilometer of mobile energy storage; is the moving distance of the mobile energy storage vehicle at time t-1 to t; and is the charging power and discharging power of mobile energy storage at time t, respectively; Wind power cost The wind power cost is composed of the abandoned wind cost and the wind turbine operation cost, and the specific formula is: ; In the formula, is the unit power generation cost of the wind turbine generator; is the wind power generation power at time t; is the unit wind curtailment cost; is the wind curtailment amount. 2.The power grid source-load collaborative scheduling method considering multi-type flexible resources of claim 1, wherein, In step 1, the multi-type flexible resource coordinated scheduling model specifically includes: 1) fixed energy storage model: ; ; ; ; In the formula, represents the energy storage state of the fixed energy storage at the current time t; and respectively represent the charging efficiency and discharging efficiency of the fixed energy storage; and respectively represent the upper limit and lower limit of the energy of the fixed energy storage; represents the charging power of the fixed energy storage at time t; represents the discharging power of the fixed energy storage at time t; and respectively represent the upper limit of the charging power and the upper limit of the discharging power of the fixed energy storage; is the time period length; 2) mobile energy storage model: The shortest path of the mobile energy storage vehicle carrying energy storage devices is as follows: ; wherein is the optimal path matrix; denotes the shortest distance from node i to node j at time t; ; is the total number of nodes at time t; The constraints of mobile energy storage are as follows: ; ; ; In the formula, represents the running state of mobile energy storage, 1 means that the mobile energy storage is moving, 0 means that the mobile energy storage is in a static state, and if the charging port is accessed, the charging and discharging operation is performed; represents the flag bit of the mobile energy storage accessing the power distribution network, if the mobile energy storage is connected to the grid at node i at time t, then is 1, otherwise if the mobile energy storage is in an off-grid state, the value is 0, and the mobile energy storage cannot be in both grid-connected and mobile states at the same time; is the node set where the mobile energy storage charging port is arranged in the traffic network; h represents the moving time of the mobile energy storage; T is the total time length of the operation cycle; represents the shortest time for the mobile energy storage to transfer from node i to node j in the traffic network, which is obtained from the shortest path. In addition to path optimization, the charging and discharging cost model of the mobile energy storage vehicle is the same as that of the fixed energy storage model; 3) demand response resource model: Flexible loads participating in demand response are divided into transferable loads and reducible loads; The transferable load model is as follows: ; ; ; In the formula, and respectively represent the load before and after the load transfer adjustment at time t; and respectively correspond to the load transferred in and the load transferred out according to the demand response at time t; is the maximum transferable margin of the load; is the load prediction value at time t; The reducible load model is as follows: ; ; ; ; ; wherein, is the power before load shedding at time t; is the amount of load shedding at time t; is the maximum load that can be shed; is the maximum number of load shedding; is a 0-1 variable, indicating the state of load shedding, 0 if the load is not responsive to shedding, 1 if the load is responsive to shedding; and respectively represent the minimum load shedding time and the maximum load shedding time; is the power after load shedding at time t; is the load after shedding at time t; After adjusting all flexible loads through the demand response mechanism, the model of the total load after demand response adjustment is as follows: ; In the formula, represents the adjusted total load. 3.The power grid source-load collaborative scheduling method considering multi-type flexible resources of claim 2, wherein, The constraint conditions in step 2 are as follows: (1) the power grid operation needs to meet system unit constraints, power flow balance constraints and node voltage constraints; specifically as follows: 1) the system unit constraints are as follows: ; ; In the formula, is the climbing rate of the thermal power unit; is the reactive power output of the thermal power unit under the jurisdiction of node j at time t; , , and are the minimum active power, the minimum reactive power, the maximum active power and the maximum reactive power of the thermal power unit j respectively, , are the actual output and the maximum output of the wind power at time t respectively. 2) the power flow balance constraints are as follows: ; ; where, and Pi(t) and Qi(t) are the active and reactive power injected by all generators and energy storage devices at node i at time t, respectively, Qi(t) is the reactive power injected at node i at time t; and Pi(t) and Qi(t) are the adjusted active and reactive power at node i at time t, respectively; and Vi(t) and Vj(t) are the voltage magnitudes at node i and node j at time t, respectively; is the voltage phase angle difference between line ij at time t; and Gi,j and Bj,i are the conductance and susceptance between nodes i and j, respectively; 3) the voltage constraints are as follows: ; wherein and Vminand Vmaxdenote the lower and upper voltage limits of node i, respectively.

4. The power distribution network source-load collaborative scheduling method considering multiple types of flexible resources according to claim 3, characterized in that, In step 3, the Markov Decision Process consists of a system global state space , an action space , a global reward function R, system state transition probabilities , and a reward discount factor ; the detailed structure is as follows: In the solved reinforcement learning task, the state information includes the power generation , voltage , energy storage state of charge , load , time t and the location where the mobile energy storage is located , the state space is expressed as: ; The action space A of the intelligent agent is a mixed action space composed of discrete and continuous actions, as shown in the following formula: ; wherein, is the power of the fixed energy storage at time t; is the power of the mobile energy storage; is the power change amount in response to the user demand; is the location of the mobile energy storage decided by the agent. When the agent violates the corresponding constraint, the agent is given a corresponding punishment ; At the same time, the objective function is adaptively adjusted to convert the system total cost minimization problem into the reward maximization form of reinforcement learning, and the reward function r(t) of the agent is established as: ; In the formula, is the penalty coefficient of the mth limit, , are the actual value and the boundary value of the mth limit parameter, respectively. The parameterized deep Q network algorithm framework is introduced to solve the scheduling model, the parameterized deep Q network algorithm introduces a parameterization mechanism, combines the advantages of Q value network and deep deterministic policy gradient algorithm, and respectively uses a Q value network and a deterministic policy network to select discrete actions and continuous actions, wherein the deterministic policy network is used to select continuous actions, the Q value network obtains an evaluation value Q in combination with continuous actions and input states, and selects a discrete action according to the Q value; the intelligent agent obtains a state from the environment and inputs it into the deterministic policy network, and outputs a continuous action from the deterministic policy network, and inputs the continuous action and the state into the Q value network, and selects a discrete action with the maximum value as the discrete action.

5. The power distribution network source-load collaborative scheduling method considering multiple types of flexible resources according to claim 4, characterized in that, In step 4, the algorithm updating process is divided into Q value network updating and deterministic policy network updating; For the parameterized deep Q-network algorithm under the mixed action space, the action space is divided into discrete actions and continuous actions, denoted as k and x k For the discrete action k, its update process depends on the fitting of the Q value, and the network is updated through the guidance of the Q value size, and the Q value is represented as: ; wherein, is the continuous action obtained from the state at time t+1; and are the state, discrete action and continuous action at time t, respectively; is the reward at time t; is the set of action variables, is the set of state variables; E denotes the expected value; Since the output of the continuous action is obtained by the deterministic policy network, the parameters of the deterministic policy network are utilized The continuous action x is approximately fitted k The target Q value is obtained by the Bellman equation is: ; wherein is the updated parameter of the Q network; represents the policy value calculated at time t according to the state and the network parameters and action k; The update of the Q-network is by minimizing the loss function implemented, whose formula is as follows: ; When updating the deterministic policy network, the purpose of the deterministic policy network update is to find the optimal parameters , the parameters are updated , the optimal Q value is obtained at a fixed time, and the loss function of the deterministic policy network update is defined as: ; In the formula, is the updated parameter at time t; is the total number of actions.

Citation Information

Patent Citations

  • Source-grid-load-storage bi-level collaborative planning method for power distribution network

    WO2025020587A1

  • Medium-voltage power distribution network source-grid-storage double-layer planning method considering flexibility

    WO2025065759A1