Multi-time scale demand response optimization method and device based on dual Q network

By constructing a multi-timescale demand response optimization method based on dual-Q networks, the problems of response uniformity and user participation of traditional demand response strategies in complex power systems are solved, enabling flexible response to uncertain disturbances and stable and economical operation of the power grid.

CN120875427APending Publication Date: 2025-10-31RES INST OF ECONOMICS & TECH STATE GRID SHANDONG ELECTRIC POWER
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511047652.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional demand response strategies suffer from limited response dimensions, low user participation, and poor real-time adaptability when facing complex power systems. Furthermore, existing deep reinforcement learning-based methods struggle to effectively handle high-frequency uncertain disturbances and power physics constraints, resulting in control strategies lacking feasibility and interpretability.

Method used

A multi-timescale demand response optimization method based on a dual-Q network is constructed. The objective function and physical constraints are built in the initialization phase. The dual deep Q network DDQN is used for agent interaction and network training. The load scheduling is optimized by combining a multi-objective reward mechanism and physical constraints.

Benefits of technology

It achieves coordinated optimization under multiple users and multiple time scales, enhances the ability of the active distribution network to cope with uncertain disturbances, improves the utilization efficiency of renewable energy and the guarantee of power quality, and ensures the economic efficiency and stable operation of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875427A_ABST
    Figure CN120875427A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of demand response optimization, and particularly relates to a multi-time scale demand response optimization method and device based on a double Q network, and the method comprises the steps: stage 1, initialization stage: importing basic data of a power distribution network, and constructing a target function and a double deep Q network DDQN structure; stage 2, an agent interaction stage: the main evaluation network predicts Q values of all potential actions, an agent selects the actions according to a greedy strategy, the selected actions are executed in a simulation environment, and an instant reward is calculated according to a reward function; stage 3, a network training updating stage: forming a training sample set, calculating a target Q value of a selected action by using a target network, calculating a square difference between a predicted Q value and a target value, and updating network parameters; and 4, a strategy deployment stage: when the learning process is converged or reaches the maximum training round, outputting a final demand response strategy. The economical efficiency of the power grid can be improved, and stable operation of the power grid is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of demand response optimization technology, specifically relating to a multi-timescale demand response optimization method and apparatus based on dual-Q networks. Background Technology

[0002] As power systems increasingly exhibit a trend towards decentralization and digitalization, the traditional distribution network's role as a passive power transmission carrier is undergoing transformation. Modern distribution networks are gradually evolving into dynamic interactive platforms characterized by bidirectional power flow, distributed generation, flexible loads, and user behavior responses. This evolution is primarily driven by the significant increase in the integration rate of distributed energy sources (such as rooftop solar power) and the rise of new electricity consumption patterns such as electric vehicles and smart terminal loads. While these changes enhance system flexibility, they also introduce unprecedented volatility and operational uncertainties, posing new challenges to the scheduling and control of distribution networks.

[0003] Demand response (DR), as a key mechanism for improving grid flexibility and control capabilities, has become a core research direction in active distribution networks. Traditional DR strategies typically rely on simple rules such as static load shedding or time-of-use pricing, with centralized operators uniformly scheduling user groups. This approach suffers from significant shortcomings, including a single response dimension, limited user participation, and poor real-time adaptability. In recent years, optimization modeling methods such as mixed-integer programming and robust optimization have been widely introduced to improve the refinement and dispatchability of demand response. Some studies have also considered hierarchical game theory and user behavior modeling. However, most of these methods rely on perfect prediction or deterministic parameter settings, making it difficult to cope with high-frequency uncertain disturbances. Furthermore, they suffer from high computational complexity in high-dimensional, multi-constraint scenarios, limiting real-time deployment.

[0004] Meanwhile, reinforcement learning (RL) technology has attracted widespread attention in the field of energy system dispatching due to its characteristics of not requiring precise system modeling and obtaining optimal control strategies through interactive learning. In particular, deep reinforcement learning (Deep RL), by combining deep neural networks with RL algorithms, has the ability to handle high-dimensional, continuous, and incompletely observable environments, and has been applied to multiple scenarios such as energy storage control and electric vehicle charging management. Nevertheless, most existing demand response studies based on DRL are limited to isolated users or simplified systems, lacking systematic modeling of power physics constraints, user heterogeneity, and time-coupled behaviors. This results in control strategies lacking feasibility and interpretability, making it difficult to promote and apply them in actual power grids. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a multi-timescale demand response optimization method and apparatus based on dual-Q network, and to construct a demand response optimization framework based on dual deep Q network DDQN, which can improve the economy of the power grid and ensure the stable operation of the power grid.

[0006] To achieve the above objectives, this invention provides a multi-timescale demand response optimization method based on dual-Q networks, comprising the following steps: Phase 1, Initialization Phase: Import basic data of the distribution network, construct an objective function that includes multi-dimensional cost drivers such as economic benefits, photovoltaic absorption and voltage stability, set the constraints of the objective function, and set the physical constraints of the distribution network. A dual-depth Q-network (DDQN) structure is constructed, including a main evaluation network and a target network. The state space and action space are defined, network parameters are set, and then the initial operating state is set for each node in the distribution network. Phase 2, Agent Interaction Phase: Obtain the current power distribution network state information from the simulation environment, input the current state into the main evaluation network, predict the Q value of all potential actions, select actions according to the greedy strategy, execute the selected actions in the simulation environment, and calculate the immediate reward according to the reward function. Store the current state, action, reward, and next state in the experience replay pool; Phase 3, Network Training and Update Phase: Randomly sample a batch of data from the experience replay pool to form a training sample set. Use the target network to calculate the target Q value of the selected action. Calculate the squared difference between the predicted Q value and the target value according to the temporal difference loss function. Update the parameters of the main evaluation network through the backpropagation algorithm. Update the parameters of the target network using the exponential moving average of the weights of the main network according to the soft update rule. Phase 4, Strategy Deployment Phase: Based on whether the loss function or average reward has converged, or whether the training round limit has been reached, determine whether the learning process should terminate. When the learning process converges or the maximum training round is reached, the final demand response strategy is output.

[0007] As a preferred embodiment of the present invention, in stage one, the basic data of the distribution network includes the distribution network topology, node numbers and electrical parameters, user category information, baseline load, photovoltaic predicted output, and electricity price curve. The objective function constructed is: (1); In the formula, It is the demand response action coefficient; It is a power flow variable in the power grid system; It is a set of auxiliary variables; n is the index of the node, N is the total number of nodes; T represents the running cycle. This represents a specific moment; l is the index of the route, and L is the set of routes; Indicates that node n at time n The demand response action coefficient represents the proportion of load adjustment. Indicates that node n at time n The time-varying electricity purchase price; Indicates that node n at time n Electricity procured from the upper-level power grid under demand response measures; Indicates that node n at time n The load reduction penalty factor; Indicates that node n at time n Involuntary load reduction under demand response measures; Indicates that node n at time n The weighting coefficient for photovoltaic reduction penalties; Indicates that node n at time n The demand response measures led to a reduction in power generation due to local photovoltaic overcapacity; Indicates that line l is at time... Voltage deviation penalty coefficient; Indicates that line l is at time... The actual voltage; Indicates that line l is at time... Trend variables; This indicates the deviation of line l from the reference value; This indicates the allowable tolerance for line l.

[0008] As a preferred embodiment of the present invention, two penalty layers are introduced into the objective function to regulate the fairness and time smoothness of the demand response, as follows: (2); In the formula, u represents the index of the user, and U is the set of users; The weight of user u is indicated; This represents the auxiliary time index used to calculate the user's u-cycle average response; Indicates user u at time... Demand response action coefficient; , These represent user u at time [time]. , Demand response action coefficient; The smoothness sensitivity coefficient represents user u. The constraints of the objective function include: Active power balance constraints: (3); (4); In the formula, Indicates that node n at time n The final effective load; Indicates that node n at time n Baseline load; Indicates that node n at time n Involuntary load reduction; Indicates that node n at time n Reduced photovoltaic power output; , These represent the sets of outgoing and incoming routes connected to node n, respectively. , These represent the time intervals of line g. Active power flow in the inflow and outflow directions; Indicates that node n at time n Local photovoltaic power generation capacity; Apparent power constraint: (5); In the formula, , These represent the time intervals of line l. The active and reactive power components; This indicates the heat capacity of line l; Demand response action rate constraint: (6); In the formula, This represents the minimum downward adjustment limit for user u's demand response action between adjacent time periods.

[0009] As a preferred embodiment of the present invention, the physical constraints of the power distribution network include: Voltage amplitude constraint: (7); In the formula, Indicates that node n at time n The voltage amplitude; , These represent the lower and upper limits of the safe operating voltage at node n, respectively. Solar photovoltaic reduction constraints: (8); In the formula, This represents the set of nodes that have access to photovoltaic power. Load transfer constraints: (9); In the formula, This represents the set of users who can accept load time migration; This indicates the maximum load delay duration that user u can accept; Indicates user u at time The final effective load; Indicates user u at time Baseline load; Electric vehicle demand response service time constraints: (10); In the formula, Represents the set of electric vehicle users; Let be the indicator variable, representing the electric vehicle of user u at time t. Whether it is connected to a charging station, 1 indicates connected, 0 indicates not connected; Indicates the arrival time of the electric vehicle; Indicates the time the electric vehicle leaves; Constraints of energy storage devices: (11); In the formula, S represents the set of energy storage devices, and s is the index of the energy storage device; , These represent the time intervals of energy storage device s and s, respectively. , The state of charge; , These represent the charging and discharging efficiencies of the energy storage device s, respectively. , These represent the time intervals of energy storage device s and s, respectively. The charging and discharging power; Power factor constraint: (12); In the formula, , These represent the time intervals of node n. Active and reactive power; This represents the maximum power factor angle at node n; Load recovery constraints: (13); In the formula, Represents a set of residential users; This indicates the length of the recovery window for user u after load adjustment; Indicates user u at time The baseline load.

[0010] As a preferred embodiment of the present invention, the state space is defined as including time intervals. Baseline load vector Photovoltaic power generation forecast Electricity price signals Voltage measurement value ,action state vector ; The action space includes the different load response strategies that a user can execute at any given time, including increasing, maintaining, and decreasing; it also initializes the current network weight parameters. Target network parameters Synchronization factor Learning rate and time agent exploration probability The current network is the primary evaluation network; initial state It contains information on the initial load, photovoltaic output, voltage level, and corresponding electricity price period for all nodes. This state will be used as the starting point for the interaction between the agent and the environment and will be input into the main evaluation network.

[0011] As a preferred embodiment of the present invention, in the second stage, the simulation environment is a complete simulated power grid environment built by importing the topology of the distribution network, node numbers and electrical parameters, category information of various users, user baseline load, photovoltaic predicted output, and electricity price curve.

[0012] As a preferred embodiment of the present invention, the specific configuration of the DDQN structure includes: Set participation budget constraints: (14); In the formula, This represents the maximum total response budget that user u can provide throughout the entire cycle; Set a maximum total demand response constraint: (15); In the formula, Indicates the time of the power grid system The maximum allowable total demand response; Define time The state vector of the reinforcement learning agent input, i.e. : (16); In the formula, , Representing time respectively Voltage measurement value, action vector; Define time The action vector, i.e. : (17); In the formula, superscript Indicates transpose; Represent a real vector space with dimension equal to the number of users; Each element This represents the flexible load adjustment ratio of user u, i.e., the demand response action coefficient. Define the current network and the target network: (18); In the formula, Indicates the current network; Indicate the target network; For a moment The state vector; Indicates the current time The instant reward received; Discount factor; Indicate the next state The optimal action chosen by the network at this time; Indicates the expected value; This is the set of parameters for the current network; The set of parameters for the target network; The target network's soft update rules are as follows: (19); The exploration probability of an agent is defined as: (20); In the formula, To explore initial values ​​for the probability; This indicates the lower bound of the exploration probability; This refers to the decay rate parameter; Experience replay settings: (twenty one); In the formula, This represents a batch of experience samples of length K, starting from the k-th record, where i is used to index each record in the experience replay. , These are the states at steps i and i+1, respectively; In the state The following actions were taken; For the action The resulting rewards; Define the reward function: (twenty two); In the formula, Represents the time-varying electricity purchase price of node n; Indicates that node n at time n Electricity purchased from the upper-level power grid; Represents node n The load reduction penalty factor; This represents the photovoltaic reduction penalty weight coefficient for node n; Indicates the line Voltage deviation penalty coefficient; The loss function used during training is: (twenty three); In the formula, Represents the TD loss function; This represents the number of samples in the batch. This represents the state corresponding to the w-th sample. In the state The following actions were taken; This represents the target Q-value of the w-th sample; The current network parameters are updated in the following way: (twenty four); In the formula, For the TD loss function with respect to parameters The gradient; Set physical boundary constraints: (25); In the formula, , These represent the minimum and maximum allowable values ​​of the demand response action coefficient, respectively. This represents a clipping function that restricts input values ​​to a given upper and lower bound range.

[0013] As a preferred embodiment of the present invention, the reinforcement learning process of the DDQN structure is as follows: The agent selects the optimal action based on the current state and updates its Q-value using a reward signal to optimize the demand response strategy. (26); In the formula, Indicates time Standardized state vector; A mean vector representing the states of the training sample set; The standard deviation vector representing the state of the training sample set; Define sensitivity index: (27); In the formula, For sensitivity indicators; The weighting coefficient for photovoltaic reduction items; For a moment Local solar power generation was reduced due to overcapacity; For a moment The total photovoltaic power generation capacity; This is the weighting coefficient for the voltage deviation term; Define the convergence metric: (28); In the formula, This represents the convergence index; V is the size of the sliding window, i.e., how many training epochs have passed recently; E is the number of the current training epoch. , These are the average returns for the v-th and v-1th periods, respectively.

[0014] As a preferred embodiment of the present invention, the network training update phase is triggered by executing a network training update process once every certain period of time or after the data volume requirement of the experience replay pool is met.

[0015] A multi-timescale demand response optimization device based on dual-Q networks includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The method described above is implemented by executing the computer program through the processor.

[0016] The beneficial effects of this invention are: This invention constructs a novel demand response framework that integrates the physical characteristics of distribution network operation, differences in user behavior, and a learning-based control mechanism to achieve coordinated optimization across multiple users and time scales. This enhances the ability of the active distribution network to cope with uncertain disturbances and improves the utilization efficiency of renewable energy and the level of power quality assurance. The intelligent agent can optimize load dispatch based on factors including energy efficiency, photovoltaic power generation limitations, and voltage stability, thereby improving the economics of the power grid and ensuring its stable operation.

[0017] This invention enables an intelligent response mechanism for heterogeneous loads in active distribution networks that does not require precise modeling. It solves the problems of traditional demand response strategies, such as the single response, limited user participation, and poor real-time adaptability in complex environments, and significantly improves the ability of active distribution networks to cope with uncertain disturbances.

[0018] This invention introduces a multi-objective reward mechanism that comprehensively considers operating costs, photovoltaic absorption, and voltage stability, achieving a precise balance between economic benefits, renewable energy absorption, and voltage stability. This not only enhances the guidance and convergence of the reinforcement learning process but also improves the interpretability and control quality of the learned strategy, ensuring the high efficiency and stability of the power grid operation. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the principle of this invention; Figure 2 This is a flowchart of the distribution network demand response scheduling process optimized by dual-depth Q network in an embodiment of the present invention. Detailed Implementation

[0020] The embodiments of the present invention will be further described below with reference to the accompanying drawings: Example 1: As Figure 1 As shown, the multi-timescale demand response optimization method based on dual-Q networks includes the following steps: Phase 1, Initialization Phase: Import basic data of the distribution network, construct an objective function that includes multi-dimensional cost drivers such as economic benefits, photovoltaic absorption and voltage stability, set the constraints of the objective function, and set the physical constraints of the distribution network. A dual-depth Q-network (DDQN) structure is constructed, including a main evaluation network and a target network. The state space and action space are defined, network parameters are set, and then the initial operating state is set for each node in the distribution network. Phase 2, Agent Interaction Phase: Obtain the current power distribution network state information from the simulation environment, input the current state into the main evaluation network, predict the Q value of all potential actions, select actions according to the greedy strategy, execute the selected actions in the simulation environment, and calculate the immediate reward according to the reward function. Store the current state, action, reward, and next state in the experience replay pool; Phase 3, Network Training and Update Phase: Randomly sample a batch of data from the experience replay pool to form a training sample set. Use the target network to calculate the target Q value of the selected action. Calculate the squared difference between the predicted Q value and the target value according to the temporal difference loss function. Update the parameters of the main evaluation network through the backpropagation algorithm. Update the parameters of the target network using the exponential moving average of the weights of the main network according to the soft update rule. Phase 4, Strategy Deployment Phase: Based on whether the loss function or average reward has converged, or whether the training round limit has been reached, determine whether the learning process should terminate. When the learning process converges or the maximum training round is reached, the final demand response strategy is output.

[0021] In Phase 1, the basic data of the distribution network includes the distribution network topology, node numbers and electrical parameters, user category information, baseline load, photovoltaic power generation forecast, and electricity price curve. The objective function constructed is: (1); In the formula, It is the demand response action coefficient; It is a power flow variable in the power grid system; It is a set of auxiliary variables (used to represent intermediate calculation results or virtual control signals); n is the node index, N is the total number of nodes; T represents the running cycle. This represents a specific moment; l is the index of the route, and L is the set of routes; Indicates that node n at time n The demand response action coefficient represents the proportion of load adjustment. Indicates that node n at time n The time-varying electricity purchase price; Indicates that node n at time n Electricity procured from the upper-level power grid under demand response measures; Indicates that node n at time n The load reduction penalty factor; Indicates that node n at time n Involuntary load reduction under demand response measures; Indicates that node n at time n The weighting coefficient for photovoltaic reduction penalties; Indicates that node n at time n The demand response measures led to a reduction in power generation due to local photovoltaic overcapacity; Indicates that line l is at time... Voltage deviation penalty coefficient; Indicates that line l is at time... The actual voltage; Indicates that line l is at time... Trend variables; This indicates the deviation of line l from the reference value; This indicates the allowable tolerance for line l.

[0022] The first term sums up the cumulative cost over the operating cycle, while the nested terms capture the hierarchical cost of all bus nodes. The second term adds a secondary voltage deviation penalty term for each line. Overall, the objective function reflects a precise balance between economic benefits, renewable energy consumption, and voltage stability.

[0023] Two penalty layers are introduced into the objective function to regulate the fairness and time smoothness of the demand response, as follows: (2); In the formula, u represents the index of the user, and U is the set of users; The weight of user u is indicated; This represents the auxiliary time index used to calculate the user's u-cycle average response; Indicates user u at time... Demand response action coefficient; , These represent user u at time [time]. , Demand response action coefficient; The smoothness sensitivity coefficient represents user u. Penalize user behavior that deviates from its historical average, i.e., compare the average action over the planning period with the specific actions per hour. This is scaled using user-specific weights. This ensures that no participant bears a disproportionate burden during demand response, thus promoting fairness among heterogeneous user groups. The second term reinforces inertia by penalizing abrupt jumps between consecutive actions through a quadratic term of time difference. The user's time smoothness sensitivity coefficient modulates sensitivity to abrupt changes, promoting temporal stability of behavior, which is crucial for practical deployment in residential and commercial load scenarios. Together, these components make the demand response strategy both socially acceptable and operationally feasible.

[0024] The constraints of the objective function include: Active power balance constraints: (3); (4); In the formula, Indicates that node n at time n The final effective load (the total amount of electrical energy actually consumed in the power grid within a specific time period). Indicates that node n at time n Baseline load; Indicates that node n at time n Involuntary load reduction; Indicates that node n at time n Reduced photovoltaic power output; , These represent the sets of outgoing and incoming routes connected to node n, respectively. , These represent the time intervals of line g. Active power flow in the inflow and outflow directions (based on topological direction, following Kirchhoff's laws). Indicates that node n at time n The local photovoltaic power generation capacity; Equation (3) lays the foundation for node balance constraints and links demand response decisions with power flow variables. Equation (4) is the active power balance constraint for node n, ensuring that net inflow minus outflow equals net demand.

[0025] Apparent power constraint: (5); In the formula, , These represent the time intervals of line l. The active and reactive power components; This represents the heat capacity of line l; the apparent power of each line l is limited to its heat capacity. Within this scope, ensure that no line overload occurs under demand-response-based power flow redistribution.

[0026] Demand response action rate constraint: (6); In the formula, This represents the minimum downward limit for user demand response actions between adjacent time periods. This constraint limits the maximum rate of decrease in user demand response actions between adjacent time periods.

[0027] During demand response, the physical constraints of the power grid (such as power balance and voltage limitations) must be considered to ensure that the strategy selected by the agent meets the stability and security requirements of the power grid. The physical constraints of the distribution network include: Voltage amplitude constraint: (7); In the formula, Indicates that node n at time n The voltage amplitude; , These represent the lower and upper limits of the safe operation of the voltage at node n, respectively; voltage amplitude. It is essential to keep each node and time step within a safe range. This is to ensure operational integrity and protect voltage-sensitive equipment. This is a standard but critical condition for the feasibility of a power distribution system.

[0028] Solar photovoltaic reduction constraints: (8); In the formula, This represents the set of nodes with photovoltaic (PV) access; this constraint limits PV cutoffs to the available power generation range, prohibits negative cutoffs (i.e., excess consumption), reflects physical availability constraints, and prevents infeasible demand response scheduling.

[0029] Load transfer constraints: (9); In the formula, This represents the set of users who can accept load time migration; Indicates the maximum load delay duration that user u can accept (shift window); Indicates user u at time The final effective load; Indicates user u at time The baseline load is established; industrial load transfer constraints ensure that all reduced electricity consumption is compensated within a tolerance window. This method achieves load reshaping without net energy loss, while satisfying thermal or industrial process constraints.

[0030] Electric vehicle demand response service time constraints: (10); In the formula, Represents the set of electric vehicle users; Let be the indicator variable, representing the electric vehicle of user u at time t. Whether it is connected to a charging station, 1 indicates connected, 0 indicates not connected; Indicates the arrival time of the electric vehicle. This indicates the time when the electric vehicle leaves and the available charging time window, constraining the actual time period during which the electric vehicle participates in demand response services and avoiding unreasonable scheduling arrangements.

[0031] Constraints of energy storage devices: (11); In the formula, S represents the set of energy storage devices, and s is the index of the energy storage device; , These represent the time intervals of energy storage device s and s, respectively. , The state of charge; , These represent the charging and discharging efficiencies of the energy storage device s (typically less than 1). , These represent the time intervals of energy storage device s and s, respectively. The charging and discharging power; this dynamic constraint across time periods is key to time-coupled demand response decisions.

[0032] Power factor constraint: (12); In the formula, , These represent the time intervals of node n. Active and reactive power; This represents the maximum power factor angle at node n; the power factor constraint limits the angle between the actual power and the apparent power to a certain value. Within a given range, this control maintains voltage support while reducing redundant reactive power flow. This control is particularly important when regulating active / reactive power flow in demand response.

[0033] Load recovery constraints: (13); In the formula, Represents a set of residential users; This indicates the length of the recovery window for user u after load adjustment; Indicates user u at time The baseline load. This simulates activities such as device charging or space heating, which can only be delayed and not eliminated.

[0034] The state space is defined as including time intervals. Baseline load vector Photovoltaic power generation forecast Electricity price signals Voltage measurement value ,action state vector ; The action space includes the different load response strategies that a user can execute at any given time, including increasing, maintaining, and decreasing; it also initializes the current network weight parameters. Target network parameters Synchronization factor Learning rate and time agent exploration probability The current network is the primary evaluation network; initial state It contains information on the initial load, photovoltaic output, voltage level, and corresponding electricity price period for all nodes. This state will be used as the starting point for the interaction between the agent and the environment and will be input into the main evaluation network.

[0035] In Phase Two, the simulation environment is a complete simulated power grid environment built by importing the distribution network topology, node numbers and electrical parameters, category information of various users, user baseline load, photovoltaic predicted output, and electricity price curves.

[0036] Considering objectives such as energy efficiency, photovoltaic power generation absorption, and voltage stability, corresponding reward signals are calculated to guide the learning process of the agent. The specific settings of the DDQN structure include: Set participation budget constraints: (14); In the formula, This represents the maximum total response budget that user u can provide throughout the entire cycle. This value is calibrated individually for each user based on historical contracts, fatigue levels, or equipment capabilities. Set a maximum total demand response constraint: (15); In the formula, Indicates the time of the power grid system The maximum allowable total demand response; system-level protection mechanism: at any given time step, the total contribution of aggregated demand response must not exceed the upper limit, ensuring that the grid operator can still control the overall demand response limit under forecast uncertainty.

[0037] Define time The state vector of the reinforcement learning agent input, i.e. : (16); In the formula, , Representing time respectively The voltage measurement value and action vector encapsulate both the current system conditions and the time context that depends on memory.

[0038] Define time The action vector, i.e. : (17); In the formula, superscript Indicates transpose; Represent a real vector space with dimension equal to the number of users; Each element This represents the flexible load adjustment ratio of user u, i.e., the demand response action coefficient. The space is continuous and multidimensional, allowing for fine-grained control over the response behavior to the needs of all user types.

[0039] Define the current network and the target network: (18); In the formula, This represents the current network (current Q-network), used to estimate the value under the current policy. Indicates the current state of the network. Next action Value estimation; Indicates the target network (target Q network); For a moment The state vector; Indicates the current time The instant reward received; As a discount factor, it weighs future rewards; Indicate the next state The optimal action chosen by the network at this time; Indicates the expected value; This is the set of parameters for the current network; The set of parameters for the target network; The target update rule for a dual-depth Q network is as follows: state-action pairs Q-value estimation employs a dual-network guided mechanism. The network selects actions via argmax, while the target network... Evaluate the selected action. This decoupling method solves the overestimation bias problem in traditional Q-learning.

[0040] The target network's soft update rules are as follows: (19); The target network's soft update rule ensures learning stability. Target parameters. Updated using the exponential moving average of the main network weights, by a synchronization factor. (Usually a small constant such as 0.005 is used to control the update magnitude.)

[0041] The exploration probability of an agent is defined as: (20); In the formula, To explore initial values ​​for the probability; This indicates the lower bound of the exploration probability; This is the decay rate parameter, controlling the rate of exponential decrease. A greedy strategy ensures thorough exploration in the early stages, gradually shifting towards exploitation in the later stages; Experience replay settings: (twenty one); In the formula, This represents a batch of experience samples of length K, starting from the k-th record, where i is used to index each record in the experience replay. , These are the states at steps i and i+1, respectively; In the state The following actions were taken; For the action The reward derived from mini-batch experience tuples sampled from the experience replay cache is used for stochastic gradient descent. Each tuple contains the current state, action, reward, and next state. Mini-batch training reduces inter-sample correlation and improves convergence stability.

[0042] Define the reward function: (twenty two); In the formula, Represents the time-varying electricity purchase price of node n; Indicates that node n at time n Electricity purchased from the upper-level power grid; Represents node n The load reduction penalty factor; This represents the photovoltaic reduction penalty weight coefficient for node n; Indicates the line The voltage deviation penalty coefficient is used to measure the cost of voltage deviation from the reference value. The reward function penalizes system operating cost factors: electricity purchase cost, load shedding, photovoltaic curtailment, and voltage deviation. Each penalty item is scaled by a time-varying or line-specific weight coefficient, enabling the agent to learn system-friendly scheduling strategies. The loss function used during training is: (twenty three); In the formula, Represents the TD loss function; This represents the number of samples in the batch. This represents the state corresponding to the w-th sample. In the state The following actions were taken; The target Q-value represents the w-th sample; the temporal difference loss function measures the predicted Q-value and the target value calculated by equation (18). The squared difference between the two. Minimizing this loss through gradient descent is the core training objective of Dual DQN.

[0043] The current network parameters are updated in the following way: (twenty four); In the formula, For the TD loss function with respect to parameters The gradient of temporal difference loss guides how to adjust weights. To minimize prediction error over time.

[0044] Set physical boundary constraints: (25); In the formula, , These represent the minimum and maximum allowable values ​​of the demand response action coefficient, respectively. This represents a clipping function that restricts input values ​​to a given upper and lower bound range.

[0045] The reinforcement learning process of the DDQN structure is as follows: The agent selects the optimal action based on the current state and updates the Q value through reward signals to optimize the demand response strategy, gradually learning how to minimize costs and maintain grid stability in the power grid. (26); In the formula, Indicates time Standardized state vector; A mean vector representing the states of the training sample set; The standard deviation vector represents the state of the training sample set; the state input passes through the training set statistics. Normalization is a standard preprocessing step that accelerates convergence and stabilizes neural network training at different data scales.

[0046] Define sensitivity index: (27); In the formula, For sensitivity indicators; The weighting coefficient for photovoltaic reduction items; For a moment Local solar power generation was reduced due to overcapacity; For a moment The total photovoltaic power generation capacity; The weighting coefficient for the voltage deviation term; the sensitivity index reflects the system's responsiveness to DR actions, balancing the contributions of photovoltaic reduction and voltage deviation, and can be used to adjust reward weights or optimize system target priorities.

[0047] Define the convergence metric: (28); In the formula, This represents the convergence metric, which measures the change in reward over the most recent rounds; V is the size of the sliding window, i.e., how many training epochs have passed recently; E is the number of the current training epoch. , These are the average returns for the v-th and v-1th periods, respectively. Average return over K rounds Define a sliding convergence metric, which is typically used to detect training plateaus or trigger an early stopping mechanism when improvements stall.

[0048] The network training update phase is triggered by executing the network training update process once every certain period of time or after the data volume requirement of the experience replay pool is met.

[0049] like Figure 2 As shown, the complete process of this embodiment is as follows: S1. First, import the required basic data, including the distribution network topology, node numbers and electrical parameters, user category information, user baseline load, photovoltaic power output forecast, and electricity price curves. Using this information, build a complete simulated power grid environment, ready to enter the intelligent agent scheduling simulation phase. S2. Construct the overall structure of Double DQN, including a main evaluation network and a target network, which have the same structure but independent parameters. The state space is defined as including time steps. Baseline load vector Photovoltaic power generation forecast Electricity price signals Voltage measurement value ,action state vector The action space includes the different load response strategies that a user can execute at any given time (such as increasing, maintaining, or decreasing); it also initializes the current network weight parameters. Target network parameters Synchronization factor Learning rate Exploring Probability Key parameters; S3. Set the initial state before the simulation begins. This includes information on the initial load, photovoltaic output, voltage level, and corresponding electricity price periods for all nodes. This state will serve as the starting point for the agent's interaction with the environment, inputting it into the evaluation network. S4. The reinforcement learning algorithm based on dual deep Q network advances along the time axis hour by hour. At each time step, the agent interacts with the environment and learns the policy. S5. Extract the current time from the simulation environment. status It is fed into the main evaluation network as input; S6. The main evaluation network evaluates all possible actions in the current state. Perform Q-value prediction to obtain ; S7, based on Strategy for selecting the current action: based on probability Choose random actions to increase exploration; S8, based on probability Choose the action with the maximum Q value. ; S9. Perform power flow analysis to verify whether the actions lead to violations of hard constraints such as voltage overruns and current overloads. Based on the current operation results, calculate the immediate reward using equation (22). ; S10. During the interaction process, every certain number of time steps or after the data volume requirement of the experience pool is met, a network training update process is executed once: a batch of data is randomly sampled from the experience replay pool to form a training sample set. Execute steps S5 through S8, then use the target network to calculate the target value for that action. Calculate the Temporal Difference (TD) loss function The parameters of the main evaluation network are updated using a backpropagation algorithm (such as gradient descent). This makes it approach the target network, and finally the soft update of the target network is achieved through equation (19); S11. If the loss function or average reward converges, or the number of training rounds reaches the upper limit, then terminate the learning process and output the final policy.

[0050] This embodiment constructs a unified control framework for multiple users and time periods based on a dual deep Q network. For heterogeneous loads such as residential, commercial users and electric vehicles in active distribution networks, it proposes an intelligent response mechanism that does not require precise modeling, enabling flexible and stable load regulation strategy learning even in complex and uncertain environments.

[0051] A mathematical modeling architecture that integrates power grid physical constraints and user behavior constraints is constructed. Operational constraints such as voltage amplitude limits, node power balance equations, upper and lower boundaries of load response, and available time windows are introduced to ensure that the output strategy of the agent not only meets the physical operation requirements of the system, but also satisfies the adjustability and acceptability of the user side.

[0052] The design incorporates a multi-objective reward mechanism that jointly considers operating costs, photovoltaic power consumption, and voltage stability. This achieves multi-objective coordination among operating costs, photovoltaic curtailment, and voltage stability, enhancing the guidance and convergence of the reinforcement learning process while improving the interpretability and control quality of the learned strategies.

[0053] Systematic simulation verification was conducted on the IEEE 33-bus system with dynamic photovoltaic output and heterogeneous user structure. The results show that the framework proposed in this embodiment exhibits good strategy convergence, scalability and operation performance under different prediction errors, user participation rates and system configurations. It provides a feasible path and technical foundation for further promotion to intelligent power distribution systems with distributed control, multi-agent collaboration and integration with market mechanisms.

[0054] Example 2: A multi-timescale demand response optimization device based on dual-Q network, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the method in Example 1 is implemented by the processor executing the computer program.

Claims

1. A multi-time-scale demand response optimization method based on dual-Q networks, characterized in that... Includes the following stages: Phase 1, Initialization Phase: Import basic data of the distribution network, construct an objective function that includes multi-dimensional cost drivers such as economic benefits, photovoltaic absorption and voltage stability, set the constraints of the objective function, and set the physical constraints of the distribution network. A dual-depth Q-network (DDQN) structure is constructed, including a main evaluation network and a target network. The state space and action space are defined, network parameters are set, and then the initial operating state is set for each node in the distribution network. Phase 2, Agent Interaction Phase: Obtain the current power distribution network state information from the simulation environment, input the current state into the main evaluation network, predict the Q value of all potential actions, select actions according to the greedy strategy, execute the selected actions in the simulation environment, and calculate the immediate reward according to the reward function. Store the current state, action, reward, and next state in the experience replay pool; Phase 3, Network Training and Update Phase: Randomly sample a batch of data from the experience replay pool to form a training sample set. Use the target network to calculate the target Q value of the selected action. Calculate the squared difference between the predicted Q value and the target value according to the temporal difference loss function. Update the parameters of the main evaluation network through the backpropagation algorithm. Update the parameters of the target network using the exponential moving average of the weights of the main network according to the soft update rule. Phase 4, Strategy Deployment Phase: Based on whether the loss function or average reward has converged, or whether the training round limit has been reached, determine whether the learning process should terminate. When the learning process converges or the maximum training round is reached, the final demand response strategy is output.

2. The multi-timescale demand response optimization method based on dual-Q networks according to claim 1, characterized in that, In Phase One, the basic data of the distribution network includes the distribution network topology, node numbers and electrical parameters, user category information, baseline load, photovoltaic power generation forecast, and electricity price curve. The objective function constructed is: (1); In the formula, It is the demand response action coefficient; It is a power flow variable in the power grid system; It is a set of auxiliary variables; n is the index of the node, N is the total number of nodes; T represents the running cycle. Indicates one of the moments in time; l is the index of the line, and L is the set of lines; Indicates that node n at time n The demand response action coefficient represents the proportion of load adjustment. Indicates that node n at time n The time-varying electricity purchase price; Indicates that node n at time n Electricity procured from the upper-level power grid under demand response measures; Indicates that node n at time n The load reduction penalty factor; Indicates that node n at time n Involuntary load reduction under demand response measures; Indicates that node n at time n The weighting coefficient for photovoltaic reduction penalties; Indicates that node n at time n The demand response measures led to a reduction in power generation due to local photovoltaic overcapacity; Indicates that line l is at time... Voltage deviation penalty coefficient; Indicates that line l is at time... The actual voltage; Indicates that line l is at time... Trend variables; This indicates the deviation of line l from the reference value; This indicates the allowable tolerance for line l.

3. The multi-timescale demand response optimization method based on dual-Q networks according to claim 2, characterized in that, Two penalty layers are introduced into the objective function to regulate the fairness and time smoothness of the demand response, as follows: (2); In the formula, u represents the index of the user, and U is the set of users; The weight of user u is indicated; This represents the auxiliary time index used to calculate the user's u-cycle average response; Indicates user u at time Demand response action coefficient; , These represent user u at time [time]. , Demand response action coefficient; The smoothness sensitivity coefficient represents user u. The constraints of the objective function include: Active power balance constraints: (3); (4); In the formula, Indicates that node n at time n The final effective load; Indicates that node n at time n Baseline load; Indicates that node n at time n Involuntary load reduction; Indicates that node n at time n Reduced photovoltaic power output; , These represent the sets of outgoing and incoming routes connected to node n, respectively. , These represent the time intervals of line g. Active power flow in the inflow and outflow directions; Indicates that node n at time n Local photovoltaic power generation capacity; Apparent power constraint: (5); In the formula, , These represent the time intervals of line l. The active and reactive power components; This indicates the heat capacity of line l; Demand response action rate constraint: (6); In the formula, This represents the minimum downward adjustment limit for user u's demand response action between adjacent time periods.

4. The multi-timescale demand response optimization method based on dual-Q networks according to claim 3, characterized in that, The physical constraints of the power distribution network include: Voltage amplitude constraint: (7); In the formula, Indicates that node n at time n The voltage amplitude; , These represent the lower and upper limits of the safe operating voltage at node n, respectively. Solar photovoltaic reduction constraints: (8); In the formula, This represents the set of nodes that have access to photovoltaic power. Load transfer constraints: (9); In the formula, This represents the set of users who can accept load time migration; This indicates the maximum load delay duration that user u can accept; Indicates user u at time The final effective load; Indicates user u at time Baseline load; Electric vehicle demand response service time constraints: (10); In the formula, Represents the set of electric vehicle users; As an indicator variable, it represents the electric vehicle of user u at time t. Whether it is connected to a charging station, 1 indicates connected, 0 indicates not connected; Indicates the arrival time of the electric vehicle; Indicates the time the electric vehicle leaves; Constraints of energy storage devices: (11); In the formula, S represents the set of energy storage devices, and s is the index of the energy storage device; , These represent the time intervals of energy storage device s and s, respectively. , The state of charge; , These represent the charging and discharging efficiencies of the energy storage device s, respectively. , These represent the time intervals of energy storage device s and s, respectively. The charging and discharging power; Power factor constraint: (12); In the formula, , These represent the time intervals of node n. Active and reactive power; This represents the maximum power factor angle at node n; Load recovery constraints: (13); In the formula, Represents a set of residential users; This indicates the length of the recovery window for user u after load adjustment; Indicates user u at time The baseline load.

5. The multi-timescale demand response optimization method based on dual-Q networks according to claim 4, characterized in that, The state space is defined as including time intervals. Baseline load vector Photovoltaic power generation forecast Electricity price signals Voltage measurement value ,action state vector ; The action space includes the different load response strategies that a user can execute at any given time, including increasing, maintaining, and decreasing; it also initializes the current network weight parameters. Target network parameters Synchronization factor Learning rate and time agent exploration probability The current network is the primary evaluation network; initial state It contains information on the initial load, photovoltaic output, voltage level, and corresponding electricity price period for all nodes. This state will be used as the starting point for the interaction between the agent and the environment and will be input into the main evaluation network.

6. The multi-timescale demand response optimization method based on dual-Q networks according to claim 1, characterized in that, In Phase Two, the simulation environment consists of a complete simulated power grid environment built by importing the distribution network topology, node numbers and electrical parameters, user category information, user baseline load, photovoltaic predicted output, and electricity price curves.

7. The multi-timescale demand response optimization method based on dual-Q networks according to claim 5, characterized in that, The specific settings of the DDQN structure include: Set participation budget constraints: (14); In the formula, This represents the maximum total response budget that user u can provide throughout the entire cycle; Set a maximum total demand response constraint: (15); In the formula, Indicates the time of the power grid system The maximum allowable total demand response; Define time The state vector of the reinforcement learning agent input, i.e. : (16); In the formula, , Representing time respectively Voltage measurement value, action vector; Define time The action vector, i.e. : (17); In the formula, superscript Indicates transpose; Represent a real vector space with dimension equal to the number of users; Each element This represents the flexible load adjustment ratio of user u, i.e., the demand response action coefficient. Define the current network and the target network: (18); In the formula, Indicates the current network; Indicate the target network; For a moment The state vector; Indicates the current time The instant reward received; Discount factor; Indicate the next state The optimal action chosen by the network at this time; Indicates the expected value; This is the set of parameters for the current network; The set of parameters for the target network; The target network's soft update rules are as follows: (19); The exploration probability of an agent is defined as: (20); In the formula, To explore initial values ​​for the probability; This indicates the lower bound of the exploration probability; This refers to the decay rate parameter; Experience replay settings: (21); In the formula, This represents a batch of experience samples of length K, starting from the k-th record, where i is used to index each record in the experience replay. , These are the states at steps i and i+1, respectively; In the state The following actions were taken; For the action The resulting rewards; Define the reward function: (22); In the formula, Represents the time-varying electricity purchase price of node n; Indicates that node n at time n Electricity purchased from the upper-level power grid; Represents node n The load reduction penalty factor; This represents the photovoltaic reduction penalty weight coefficient for node n; Indicates the line Voltage deviation penalty coefficient; The loss function used during training is: (23); In the formula, Represents the TD loss function; This represents the number of samples in the batch. This represents the state corresponding to the w-th sample. In the state The following actions were taken; This represents the target Q-value of the w-th sample; The current network parameters are updated in the following way: (24); In the formula, For the TD loss function with respect to parameters The gradient; Set physical boundary constraints: (25); In the formula, , These represent the minimum and maximum allowable values ​​of the demand response action coefficient, respectively. This represents a clipping function that restricts input values ​​to a given upper and lower bound range.

8. The multi-timescale demand response optimization method based on dual-Q networks according to claim 7, characterized in that, The reinforcement learning process of the DDQN structure is as follows: The agent selects the optimal action based on the current state and updates its Q-value using a reward signal to optimize the demand response strategy. (26); In the formula, Indicates time Standardized state vector; A mean vector representing the states of the training sample set; The standard deviation vector representing the state of the training sample set; Define sensitivity index: (27); In the formula, For sensitivity indicators; The weighting coefficient for photovoltaic reduction items; For a moment Local solar power generation was reduced due to overcapacity; For a moment The total photovoltaic power generation capacity; This is the weighting coefficient for the voltage deviation term; Define the convergence metric: (28); In the formula, This represents the convergence index; V is the size of the sliding window, i.e., how many training epochs have passed recently; E is the number of the current training epoch. , These are the average returns for the v-th and v-1th periods, respectively.

9. The multi-timescale demand response optimization method based on dual-Q networks according to claim 1, characterized in that, The network training update phase is triggered by executing the network training update process once every certain period of time or after the data volume requirement of the experience replay pool is met.

10. A multi-timescale demand response optimization device based on a dual-Q network, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the computer program to implement the method described in any one of claims 1-9.

Citation Information

Cited By

  • Power distribution network dynamic planning investment decision-making method and system based on deep double-Q network

    CN121073262A

  • Display method, electronic device, readable storage medium and program product

    CN121833099A