Distribution-micro-user real-time power interaction method considering user response uncertainty

By employing Markov decision processes and deep reinforcement learning within an Actor-Critic framework in the collaborative optimization of distribution networks and microgrids, combined with online learning algorithms based on multi-armed slot machine theory, the problem of user response uncertainty was solved, thereby improving the response accuracy of microgrids and the reliability and economy of distribution systems.

CN121939418APending Publication Date: 2026-04-28HEBEI UNIV OF TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEBEI UNIV OF TECH
Filing Date
2025-11-29
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address user response uncertainties in the coordinated optimization of distribution networks and microgrids, leading to increased decision-making complexity and risks during real-time operation.

Method used

A distribution network dispatching model is constructed using Markov decision process and Actor-Critic framework. Combining deep reinforcement learning and multi-armed slot machine theory, a real-time power interaction method between distribution, micro, and users is built by optimizing user response through online learning algorithms.

Benefits of technology

It enables dynamic optimization of user response uncertainty, improves the response accuracy of microgrids and the reliability and economy of power distribution systems, and reduces the dependence of traditional optimization methods on precise mathematical models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121939418A_ABST
    Figure CN121939418A_ABST
Patent Text Reader

Abstract

The invention relates to a power distribution-micro-user real-time power interaction method considering user response uncertainty. The method comprises the following steps: step 1, describing a power distribution network dispatching dynamic state through a Markov decision process; based on an Actor-Critic framework, constructing a micro-configuration interaction model; training an intelligent agent by using historical experience data, and generating a scheduling instruction issued to the micro-grid group in combination with a day-ahead scheduling result and a real-time power state; step 2, based on the micro-grid group scheduling instruction issued by the power distribution network in the step 1, performing parallel execution by each micro-grid, constructing a micro-grid user interaction model according to a multi-arm long machine theory, and solving the model by adopting an online learning algorithm fused with a confidence interval strategy, dynamically selecting a user combination execution instruction capable of responding to the micro-grid group scheduling instruction issued by the power distribution network in the step 1, so as to minimize the response deviation of the micro-grid; and sending a response instruction to the user, collecting actual response data of the user to continuously update the online learning algorithm parameters, summarizing the actual response data, and uploading the summarized actual response data to the power distribution network, thereby completing power distribution-micro-user real-time power interaction. According to the method, the problem that the response of the micro-grid is uncertain during micro-distribution optimization operation can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of collaborative optimization technology of distribution network and microgrid, and relates to a real-time power interaction method between distribution, microgrid and user, especially a real-time power interaction method between distribution, microgrid and user that considers the uncertainty of user response. Background Technology

[0002] With the large-scale grid connection of distributed energy resources, distribution networks face challenges such as the mismatch between source and load distribution in time and space, and difficulties in energy absorption, caused by the strong randomness of distributed energy output. To ensure energy absorption and the safe and stable operation of the power system, more flexible resources are needed to participate in dispatching to support system flexibility. Microgrids, as an effective carrier for the aggregation of distributed energy resources, can significantly reduce the impact of distributed energy resources on the distribution network and are an important means to improve the operating economy of the distribution system and the absorption of new energy.

[0003] Coordinated optimization of distribution networks and microgrids is a crucial means to fully leverage the flexible regulation potential of microgrids. Existing research mainly employs centralized optimization, hierarchical optimization, and distributed optimization methods to achieve distribution-microgrid coordinated optimization. Centralized optimization constructs a global optimization model encompassing the distribution network and all microgrids, which is then solved uniformly by a central controller to achieve optimal overall system operation. Hierarchical optimization decomposes the coordinated optimization problem into multiple levels. The outer layer of the distribution network optimizes its internal units and the power allocated to the connection points of each microgrid, while the inner layer of each microgrid optimizes the scheduling of its internal distributed energy resources based on the power allocated to the connection points. This method relies on an accurate simplified model and ignores the uncertainties within the microgrid. Distributed optimization treats each microgrid as an independent decision-making entity, utilizing distributed algorithms such as the alternating direction multiplier method and the augmented Lagrange relaxation method to achieve coordinated operation of the distribution network and multiple microgrids through limited information interaction.

[0004] The aforementioned methods primarily address day-ahead coordination and optimization scenarios, failing to adequately consider uncertainties in real-time operation. During intraday real-time operation, the distribution network faces power deficits due to forecasting deviations or unforeseen events, requiring rapid response based on flexible resources. Furthermore, due to the time-varying nature of user response intentions within microgrids and disturbances from external environmental factors, user response behavior exhibits significant randomness and uncertainty, leading to deviations in the microgrid's response to the distribution network. This increases the decision-making complexity and operational risks of the distribution network.

[0005] To address the aforementioned issues, this invention proposes a real-time power interaction method between the distribution system and the micro-user that considers the uncertainty of user response.

[0006] A search revealed no prior art patents that are identical or similar to this invention. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention proposes a real-time power interaction method between distribution, microgrid, and users that considers user response uncertainty, which can solve the problem of uncertain microgrid response during distribution-microgrid optimized operation.

[0008] The above-mentioned objective of this invention is achieved through the following technical solution: A real-time power interaction method for distribution-micro-user that considers user response uncertainty includes the following steps: Step 1: Characterize the dynamics of distribution network dispatching through Markov decision process; construct a distribution-microgrid interaction model based on the Actor-Critic framework; train the agent using historical experience data, and generate dispatching instructions to be issued to the microgrid group by combining the day-ahead dispatching results and real-time power status. Step 2: Based on the microgrid group dispatch instructions issued by the distribution network in Step 1, each microgrid executes them in parallel. A microgrid user interaction model is constructed according to the multi-armed slot machine theory. An online learning algorithm with fusion upper confidence interval strategy is used to solve the model. User combinations that can respond to the microgrid group dispatch instructions issued by the distribution network in Step 1 are dynamically selected to execute the instructions in order to minimize the microgrid response deviation. Response instructions are sent to users, and the actual response data of users is collected to continue updating the parameters of the online learning algorithm. The actual response data is summarized and uploaded to the distribution network, thereby completing the real-time power interaction between distribution, microgrids, and users.

[0009] The specific steps of step 1 include: (1) The real-time scheduling problem of the distribution network is modeled as a Markov decision process (MDP), and an MDP model that can describe the real-time scheduling decision problem of the distribution network is established. (2) Based on the MDP model established in step (1) of step 1, construct an agent containing an Actor and a Critic network, and train it to learn the optimal strategy for solving the MDP problem. (3) Based on the agent trained in step (2) of step 1, deploy it in the MDP model constructed in step (1) of step 1 for online application. The distribution network layer inputs the real-time collected system status and day-ahead scheduling results into the agent. The agent outputs a set of power scheduling instructions that take into account the microgrid response capability, considering the current power deficit and system status, and sends them to the microgrid group.

[0010] Furthermore, the specific steps of step 1 (1) include: For distribution networks: ① State: Contains the state of each node. Active power during operation before the time period reactive power Each node at Power deficit during the period and the actual response power of each microgrid. , is represented as: (1) In the formula, Active power The set, reactive power The set, Power deficit The set, For true response power A set; ②Action: The set of actions is: (2) In the formula For the response objectives of each microgrid The set of elements must satisfy the following constraints: (3) In the formula For microgrids Maximum response power.

[0011] ③Profit: The revenue of the distribution network is: (4) An action-value function is needed. To estimate its profit. This function can measure the profit in a given state. Under the following strategy, the distribution network Select Action The effectiveness, The higher the action, the better. Therefore, the agent's optimal strategy... It can be represented as: (5) In the formula This is the action space.

[0012] ④ State transition probability: The state transition probability in the MDP model is a pattern implicit in historical data, rather than being explicitly defined.

[0013] Moreover, the specific steps of step 1, step (2) include: ① Network initialization: Initialize the Actor network (parameters) respectively. ) and Critic network (parameters) The Actor network is responsible for generating action instructions based on the input state; the Critic network is responsible for evaluating the long-term expected value of the state-action pair.

[0014] ② Offline training and interaction: The agent interacts extensively with the environment through simulation or historical data. Each interaction generates experience data, including state, action, reward, and next state, which is stored in an experience replay buffer.

[0015] ③ Network parameter update: During training, data is sampled from the buffer and the network is updated sequentially. Updating the Critic network: updating parameters by minimizing the temporal difference error. The goal is to make the Critic network's evaluation of state-action values ​​more accurate. Update parameters. The direction is to minimize equation (6). Compared to The gradient is shown by equation (7). The update is given by equation (8).

[0016] (6) (7) (8) In the formula, for Target approximation value Let it be the expected function; This is the gradient calculation function; This represents the learning rate of the Critic network.

[0017] Updating the Actor network: The update objective of the Actor network depends on the computational results of the Critic network. (Updated...) for: (9) (10) In the formula, is the learning rate of the Actor network.

[0018] ④ Introduce adaptive learning rate: Introduce discount factor and To update the learning rate, based on the first... The next learning session The loss value used to determine the goodness of fit of a function to the true value is used to judge its quality. (11) Substituting the discount factor into equation (8) yields the new updated equation: (12) Furthermore, the specific steps of step 2 include: (1) The user aggregation problem within the microgrid is modeled as a combined online multi-armed slot machine problem, and a multi-armed slot machine model of microgrid-user interaction is established; (2) Based on the MAB model established in step (1) of step 2, the mean-variance upper confidence interval algorithm is used to make decisions and determine the set of responding users; (3) Based on the user set determined in step (2) of step 2, update the user response and parameters. The microgrid summarizes the actual response quantities of all selected users to calculate the total actual response power of the microgrid and uploads it to the distribution network, thereby completing the real-time power interaction between distribution, microgrid and user.

[0019] Moreover, the specific steps of step 2 (1) include: ① Put each user Viewed as an arm of a slot machine, the microgrid operator is seen as a player in a casino, whose task is to adjust the power distribution target issued by the distribution network in the e-th aggregation event. Select a suitable user set Send instructions.

[0020] ② Select user set After that, each user included Feedback will be provided, showing the actual power adjustment amount for the selected user. This corresponds to the "profit" in the MAB problem. .

[0021] ③ Each user They will follow different probability distribution models These distributions are unknown to microgrid operators.

[0022] The player's goal is to find the optimal arm and maximize their expected total gain over all rounds, as shown in equation (13). (13) In the formula Let be the expected function. It refers to the number of game rounds. Is it a selection arm? The generated random rewards.

[0023] Microgrid operators aim to select the optimal set of users. Sending commands to make the actual power adjustment as close as possible to the target. Reduce the deviation between the two The expected value is shown in equation (14).

[0024] (14) Moreover, the specific method of step (2) of step 2 is as follows: Calculating the index value: To balance exploring unknown users with utilizing known high-value users, and considering the stability of user responses, the MV-UCB algorithm is adopted. This algorithm integrates the mean-variance model into the MAB model. Based on the traditional model that only considers the mean, a variance term is introduced to measure risk, resulting in the following model: (15) (16) (17) In the formula, This represents the index value used to sort each user during the e-th aggregation. This represents the mean-variance model. These are the hyperparameters of the algorithm. For the confidence level, Indicates the number of times a user is selected. This represents the response value when the user is selected for the rth time.

[0025] Calculate an index value for each user This value consists of three parts: the sample mean of users' historical response volume, the variance term used to measure risk, and the upper confidence bound term to encourage exploration.

[0026] User sorting and selection: Sort all users from highest to lowest according to their calculated index values. Then, select users in this order until the sum of the contracted capacities of the selected users can satisfy the dispatch instructions issued by the distribution network, thus determining the final set of users to be executed.

[0027] Based on the index calculated using equation (15), each user is sorted, so that... Users with higher values ​​are prioritized, and then the response demand is determined based on the distribution network's published requirements. Select user set ,satisfy: (18) In the formula This represents the set of users selected during the e-th round of aggregation. Indicates according to After sorting, the j-th The index of the user it represents.

[0028] Moreover, the specific method of step 2 (3) is as follows: microgrids to selected user sets Specific response instructions are sent. Upon receiving the instructions, users independently weigh their own electricity consumption preferences, determine the actual response power, and feed it back to the microgrid. The microgrid collects the actual response quantities of its internal users and uses this data to update the historical statistical data of the responding users, including parameters such as the number of times they were selected, the mean and variance of the response quantity. Through the above process, the actual response power of this round of aggregation is obtained, and online learning of the user response characteristic model is completed. Through continuous learning, the evaluation of user response characteristics is gradually updated. The microgrid summarizes the actual response quantities fed back by all selected users to calculate the total actual response power of the microgrid and uploads it to the distribution network, thereby completing the real-time power interaction between distribution, microgrid, and users.

[0029] The advantages and beneficial effects of this invention are as follows: 1. This invention proposes a multi-level interaction method for distribution network and microgrid users that considers the uncertainty of user response. (1) For the problem of power deficit compensation in real-time interaction, a distribution network-microgrid interaction model is established based on deep reinforcement learning (DRL). Power dispatch instructions are dynamically generated through the Actor-Critic network structure to achieve compensation for real-time power deficit. (2) For the problem of dependence of traditional optimization methods on the accurate simplified model of microgrid, the optimal strategy is learned from historical operating data using data-driven deep reinforcement learning method, which solves the problem of dependence on the accurate mathematical model of microgrid. (3) For the problem of user response uncertainty, a microgrid-user interaction model is constructed based on the theory of multi-armed bandit (MAB) and online learning algorithm. User selection is dynamically optimized through the mean-variance confidence interval strategy to achieve online learning of user response uncertainty and improve the response accuracy of microgrid.

[0030] 2. This invention proposes a real-time interaction method for distribution network-microgrid-user based on deep reinforcement learning and online learning. Through the deep reinforcement learning-based distribution network-microgrid interaction model established in step 1, a Markov decision process is used to accurately characterize the dynamics of distribution network scheduling. In step 1(1), a complete real-time decision framework is constructed by defining a state space containing power deficit and the actual response power of each microgrid, as well as an action space for issuing scheduling instructions to the microgrid group. In step 1(2), through offline training and online application of the Actor-Critic network, the agent can generate optimal scheduling instructions under uncertain microgrid responses, effectively addressing the power deficit problem in real-time interaction and improving the reliability and economy of the distribution system.

[0031] 3. This invention proposes a distribution-micro interaction model based on DRL. The deep reinforcement learning method used in step 1 is a data-driven decision-making mechanism. In the network parameter update process of step (2) of step 1, it is entirely based on historical operating data, learning the optimal scheduling strategy from the data without relying on the precise mathematical model of the microgrid, thus breaking through the dependence of traditional optimization methods on the precise mathematical model of the system.

[0032] 4. This invention proposes a microgrid-user interaction model based on MAB theory, which effectively addresses the problem of user response uncertainty. The microgrid-user interaction model based on multi-armed slot machine theory constructed in step 2 evaluates user response characteristics through an online learning algorithm. In the cyclical execution of sub-step (2) and step (3) of step 2, the microgrid continuously optimizes the user selection strategy through the process of issuing instructions, collecting responses, and updating the model, realizing online learning of user response uncertainty and improving the response accuracy of the microgrid. Attached Figure Description

[0033] Figure 1 This is a diagram of the multi-level interactive architecture of the matching-micro-user of the present invention; Figure 2 This is a flowchart of the DRL-based distribution network decision-making process of the present invention; Figure 3 This is a diagram of the microgrid aggregation model based on the mean-variance upper confidence interval algorithm of the present invention; Figure 4 This is an improved IEEE 33-node topology diagram according to the present invention; Figure 5 This is a comparison chart of polymerization deviations in this invention; Figure 6 This is a graph showing the rejection rate of users selected in each round of the present invention. Figure 7 This is a response deviation diagram for different schemes of the present invention. Detailed Implementation

[0034] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings: A real-time power interaction method for distribution-microgrid-user that considers user response uncertainty decomposes the distribution network and microgrid interaction architecture into three levels: distribution, microgrid, and user. The architecture is constructed from two aspects: distribution-microgrid interaction and microgrid-user interaction, including the following steps: Step 1: Characterize the dynamics of distribution network dispatching through Markov decision process; construct a distribution-microgrid interaction model based on the Actor-Critic framework; train the agent using historical experience data, and generate dispatching instructions to be issued to the microgrid group by combining the day-ahead dispatching results and real-time power status. The specific steps of step 1 include: (1) The real-time scheduling problem of the distribution network is modeled as a Markov Decision Process (MDP), and an MDP model that can describe the real-time scheduling decision problem of the distribution network is established. The specific steps of step 1 (1) include: In the process of decision optimization based on deep reinforcement learning (DRL) in power distribution networks, it is necessary to determine the state set, action set, profit, and state transition probability of the MDP. For distribution networks: ① State: Contains the state of each node. Active power during operation before the time period reactive power Each node at Power deficit during the period and the actual response power of each microgrid. This is represented as: (1) In the formula, Active power The set, reactive power The set, Power deficit The set, For true response power The set, ②Action: When a power deficit occurs in the distribution network, it is necessary to consider both the uncertainty of the microgrid response and the need to reduce network losses after compensation operation. The distribution network rationally allocates the response targets of each microgrid based on the operating status of each node; therefore, the action set is as follows: (2) In the formula For the response objectives of each microgrid The set of elements must satisfy the following constraints: (3) In the formula For microgrids Maximum response power.

[0035] ③Profit: After each microgrid receives and responds to the command, the microgrid responds first, and the remaining power deficit is compensated by the transmission network. Therefore, the distribution network's benefit is: (4) Before issuing instructions, the distribution network cannot accurately calculate profits because it cannot know the actual response status of each microgrid. Therefore, an action-value function is needed. To estimate its profit. This function can measure the profit in a given state. Under the following strategy, the distribution network Select Action The effectiveness, The higher the action, the better. Therefore, the agent's optimal strategy... It can be represented as: (5) In the formula This is the action space.

[0036] ④ State transition probability: The model is a process of learning the optimal decision action based on historical actions and actual reward responses under certain conditions. Therefore, the state transition probability in the MDP model is an implicit pattern in historical data, rather than an explicit definition.

[0037] (2) Based on the MDP model established in step (1) of step 1, construct an agent containing an Actor and a Critic network, and train it to learn the optimal strategy for solving the MDP problem. The specific steps of step 1, step (2) include: First, the Actor and Critic networks are initialized. Then, empirical data is collected through interaction with the environment and stored in an experience replay buffer. During training, network parameters are updated by sampling data from the buffer: first, the Critic network is updated to improve value assessment accuracy by minimizing temporal difference errors; then, the Actor network is updated, and policy performance is improved through policy gradients based on the Critic network's evaluation results; a discount factor is introduced to dynamically adjust the learning rate to accelerate convergence. Through this process, a trained deep reinforcement learning agent is obtained, capable of generating power scheduling commands for the microgrid group based on real-time states. ① Network initialization: Initialize the Actor network (parameters) respectively. ) and Critic network (parameters) The Actor network is responsible for generating action instructions based on the input state; the Critic network is responsible for evaluating the long-term expected value of the state-action pair.

[0038] ② Offline training and interaction: The agent interacts extensively with the environment through simulation or historical data. Each interaction generates experience data, including state, action, reward, and next state, which is stored in an experience replay buffer.

[0039] ③ Network parameter update: During training, data is sampled from the buffer and the network is updated sequentially. Updating the Critic network: updating parameters by minimizing the temporal difference error. The goal is to make the Critic network's evaluation of state-action values ​​more accurate. Update parameters. The direction is to minimize equation (6). Compared to The gradient is shown by equation (7). The update is given by equation (8).

[0040] (6) (7) (8) In the formula, for Target approximation value Let it be the expected function; This is the gradient calculation function; This represents the learning rate of the Critic network.

[0041] Updating the Actor network: The update objective of the Actor network depends on the computational results of the Critic network. (Updated...) for: (9) (10) In the formula, is the learning rate of the Actor network.

[0042] ④ Introducing an adaptive learning rate: To improve the convergence speed, a discount factor is introduced. and To update the learning rate, based on the first... The next learning session The loss value used to determine the goodness of fit of a function to the true value is used to judge its quality. (11) Substituting the discount factor into equation (8) yields the new updated equation: (12) (3) Based on the agent trained in step (2) of step 1, deploy it in the MDP model constructed in step (1) of step 1 for online application. The distribution network layer inputs the real-time collected system status and day-ahead scheduling results into the agent. The agent outputs a set of power scheduling instructions that take into account the microgrid response capability, considering the current power deficit and system status, and sends them to the microgrid group.

[0043] Step 2: Based on the microgrid group dispatch instructions issued by the distribution network in Step 1, each microgrid executes them in parallel. A microgrid user interaction model is constructed according to the multi-armed slot machine theory. An online learning algorithm with fusion upper confidence interval strategy is used to solve the model. User combinations that can respond to the microgrid group dispatch instructions issued by the distribution network in Step 1 are dynamically selected to execute the instructions, so as to minimize the overall microgrid response deviation. Response instructions are sent to users, and the actual response data of users is collected to update the parameters of the online learning algorithm. The actual response data is summarized and uploaded to the distribution network, thereby completing the real-time power interaction between distribution, microgrids, and users.

[0044] The specific steps of step 2 include: (1) The user aggregation problem within the microgrid is modeled as a combined online multi-armed slot machine problem, and a multi-armed slot machine model of microgrid-user interaction is established; Each user is viewed as an arm of a slot machine, with the microgrid operator as the player. Each aggregation task is defined as selecting a suitable set of users to send instructions based on the power regulation target issued by the distribution network, clearly defining the optimization objective to make the actual power regulation as close as possible to the instruction value, and minimizing response deviation. User aggregation is transformed into a solvable online learning problem model. For the scenario where users within microgrid i participate in demand response, in the e-th aggregation, the distribution network needs microgrid i to provide... When responding, the system operator selects some users to send instructions, and then each user adjusts their load to form a considerable power aggregate, which meets the grid's dispatching needs.

[0045] The specific steps of step 2, step (1) include: ① Put each user Viewed as an arm of a slot machine, the microgrid operator is seen as a player in a casino, whose task is to adjust the power distribution target issued by the distribution network in the e-th aggregation event. Select a suitable user set Send instructions.

[0046] ② Select user set After that, each user included Feedback will be provided, showing the actual power adjustment amount for the selected user. This corresponds to the "profit" in the MAB problem. .

[0047] ③ Each user They will follow different probability distribution models These distributions are unknown to microgrid operators.

[0048] The player's goal is to find the optimal arm and maximize their expected total gain over all rounds, as shown in equation (13). (13) In the formula Let be the expected function. It refers to the number of game rounds. Is it a selection arm? The generated random rewards.

[0049] Microgrid operators aim to select the optimal set of users. Sending commands to make the actual power adjustment as close as possible to the target. Reduce the deviation between the two The expected value is shown in equation (14).

[0050] (14) (2) Based on the MAB model established in step (1) of step 2, the mean-variance upper confidence bound (MV-UCB) algorithm is used to make decisions and determine the user set; The specific method for step (2) of step 2 is as follows: First, an index value is calculated for each user. This index value consists of three parts: the sample mean of historical response volumes, a variance term measuring risk, and an upper confidence bound term. Then, all users are sorted from highest to lowest based on their calculated index values. Finally, users are selected sequentially according to this sorting order until the sum of the contracted capacities of the selected users satisfies the dispatch command issued by the distribution network, thus determining the final set of users for execution. This process generates a user aggregation scheme that theoretically offers the highest response accuracy and stability for the current dispatch command. Calculating the index value: To balance exploring unknown users with utilizing known high-value users, and considering the stability of user responses, the MV-UCB algorithm is adopted. This algorithm integrates the mean-variance model into the MAB model. Based on the traditional model that only considers the mean, a variance term is introduced to measure risk, resulting in the following model: (15) (16) (17) In the formula, This represents the index value used to sort each user during the e-th aggregation. This represents the mean-variance model. These are the hyperparameters of the algorithm. For the confidence level, Indicates the number of times a user is selected. This represents the response value when the user is selected for the rth time.

[0051] Calculate an index value for each user This value consists of three parts: the sample mean of users' historical response volume, the variance term used to measure risk, and the upper confidence bound term to encourage exploration.

[0052] User sorting and selection: Sort all users from highest to lowest according to their calculated index values. Then, select users in this order until the sum of the contracted capacities of the selected users can satisfy the dispatch instructions issued by the distribution network, thus determining the final set of users to be executed.

[0053] Based on the index calculated using equation (15), each user is sorted, so that... Users with higher values ​​are prioritized, and then the response demand is determined based on the distribution network's published requirements. Select user set ,satisfy: (18) In the formula This represents the set of users selected during the e-th round of aggregation. Indicates according to After sorting, the j-th The index of the user it represents.

[0054] (3) Based on the user set determined in step (2) of step 2, update the user response and parameters, and finally obtain the optimal response user combination that can respond to the distribution network command. The microgrid summarizes the actual response quantities of all selected users to calculate the total actual response power of the microgrid and uploads it to the distribution network.

[0055] The specific method for step 2, step (3) is as follows: microgrids to selected user sets Specific response instructions are sent. Upon receiving the instructions, users independently weigh their own electricity consumption preferences, determine the actual response power, and feed it back to the microgrid. The microgrid collects the actual response quantities from its internal users and uses this data to update the historical statistical data of the responding users, including parameters such as the number of times they were selected, the mean and variance of the response quantity. Through the above process, the actual response power of this round of aggregation is obtained, and online learning of the user response characteristic model is completed. Through continuous learning, the evaluation of user response characteristics is gradually updated to obtain the optimal combination of responding users capable of responding to distribution network instructions. The microgrid summarizes the actual response quantities fed back by all selected users to calculate the total actual response power of the microgrid and uploads it to the distribution network.

[0056] The working principle of this invention is: In the distribution network, deep reinforcement learning is used to generate decisions, modeling the real-time scheduling problem as a Markov decision process and solving it through an Actor-Critic framework. The trained agent can directly output the optimal power scheduling command based on the real-time system state, thus transforming the complex optimization problem into efficient policy execution without relying on the precise mathematical model of the microgrid. In the microgrid, online learning algorithms are used to optimize user selection, modeling the user aggregation problem as a multi-armed slot machine model. Faced with scheduling commands issued by the distribution network, the microgrid uses a mean-variance upper confidence interval algorithm to dynamically select user combinations with high response willingness and good stability, thereby decomposing power commands into user actions and updating the understanding of user response characteristics in each interaction. The system achieves continuous optimization through closed-loop feedback: the microgrid aggregates and uploads the actual user response data to the distribution network. This data serves as empirical data for the distribution network layer DRL agent, used for continuous policy training and improvement; it also constitutes historical experience used to drive new rounds of scheduling decisions.

[0057] Example 1: This invention decomposes the interaction architecture between the distribution network and the microgrid into three levels: distribution, microgrid, and user. Figure 1 A multi-level interactive architecture diagram of distribution-microgrid-user is constructed, focusing on two aspects: distribution-microgrid interaction and microgrid-user interaction. First, the distribution network layer sends power dispatch commands to the microgrid group; each microgrid sends demand response commands to its internal users; subsequently, users adjust their electricity load according to the commands and feed back their actual response power to their respective microgrids; the microgrids then transmit the collected aggregated response data back to the distribution network layer, achieving a closed-loop real-time interaction from global dispatch to local response and then to global feedback. For the distribution-microgrid collaborative optimization problem, a multi-agent distribution-microgrid interaction model is proposed. Historical experience data is used to train the agents, and dispatch commands are issued to the microgrids based on day-ahead dispatch results and real-time power status. Secondly, considering the uncertainty of user response, a microgrid-user interaction model is constructed and solved using a multi-armed slot machine model and an online learning algorithm incorporating upper confidence interval strategies. This yields the optimal user combination capable of responding to dispatch commands, and the actual user response information is transmitted to the distribution network, achieving coordinated operation between the distribution network and the microgrid.

[0058] The distribution-micro interaction model employs Markov decision processes to model the real-time control problem of the distribution network. To solve the decision-making problem in the continuous action space, Actor networks and Critic networks are used, where the Actor network (with parameters of...) The Critic network (with parameters) is responsible for generating decision-making strategies. This is used to evaluate the merits of the strategy.

[0059] In the process of decision optimization based on DRL in distribution networks, it is necessary to determine the state set (State), action set (Action), profit, and state transition probabilities of the MDP. For distribution networks, we have: (1) State: contains the state of each node. Active power during operation before the time period reactive power Each node at Power deficit during the period and the actual response power of each microgrid. This is represented as: (1) In the formula, Active power The set, reactive power The set, Power deficit The set, For true response power The set, (2) Action: When a power deficit occurs in the distribution network, the distribution network must consider the uncertainty of the microgrid response on the one hand, and reduce network losses after compensation operation on the other. The distribution network needs to rationally allocate the response targets of each microgrid based on the operating status of each node; therefore, the action set is the response targets allocated to each microgrid: (2) In the formula For the response objectives of each microgrid The set of responses must satisfy the following constraints: (3) In the formula For microgrids Maximum response power.

[0060] (3) Profit: After each microgrid receives and responds to the command, the microgrid first performs load response, and then the transmission network compensates for the remaining power deficit. Therefore, the distribution network's profit is the negative of this part of the cost. (4) Before issuing instructions, the distribution network cannot accurately calculate profits because it cannot know the actual response status of each microgrid. Therefore, an action-value function is needed. To estimate its profit. This function can measure the profit in a given state. Under the following strategy, the distribution network Select Action The effectiveness, The higher the movement, the better. Therefore, the optimal strategy is... It can be represented as: (5) In the formula This is the action space.

[0061] (4) State transition probability: The model is a process of learning the optimal decision action based on historical actions and actual benefit responses under certain conditions. Therefore, the state transition probability in the MDP model is a pattern implicit in historical data, rather than an explicit definition. The goal of the model is to learn the relationship between state-action-benefit through historical data and optimize the decision strategy.

[0062] In the Critic network, update parameters. The direction is to minimize equation (6). Compared to The gradient is shown by equation (7). The update is given by equation (8).

[0063] (6) (7) (8) In the formula, for Target approximation value Let it be the expected function; This is the gradient calculation function; This represents the learning rate of the Critic network.

[0064] The update objective of the Actor network depends on the computational results of the Critic network. The updated... for: (9) (10) In the formula, is the learning rate of the Actor network.

[0065] To improve the convergence speed, the following is introduced: and The learning rate is updated using the discount factor, based on the first... The next learning session The loss value used to determine the goodness of fit of a function to the true value is used to judge its quality. (11) Substituting it into equation (8), we get the new updated equation: (12) See the DRL-based distribution network decision-making flowchart. Figure 2 The distribution network allocates response targets to each microgrid based on the Actor-Critic algorithm. Each microgrid issues instructions to its internal users to respond to demand based on the response targets. The distribution network then optimizes power flow based on the actual response results of each microgrid and calculates the power response cost, generating experience in the process. The data is stored in an experience pool. During subsequent training, the distribution network uses an experience replay mechanism to sample historical experience from the experience pool to optimize the strategy. The Actor network and Critic network are continuously updated during training, gradually learning the optimal strategy for power deficits under various operating conditions to achieve better power response target allocation, thereby reducing the distribution network's reserve cost and reducing the response pressure on the transmission network.

[0066] This invention enables online learning of user response uncertainty in microgrids, employs a multi-armed slot machine model to characterize load response characteristics, and proposes an online learning algorithm based on upper confidence intervals to construct a microgrid-user interaction model through dynamic trade-offs.

[0067] The microgrid user interaction model is modeled based on the multi-armed slot machine theory, as described in the following details: (1) Put each user Viewed as an arm of a slot machine, the microgrid operator is seen as the player, whose task is to adjust the power distribution target issued by the distribution network during the e-th load adjustment event. Select a suitable user set Send instructions.

[0068] (2) Select user set After that, each user included Feedback will be provided, specifically the actual load adjustment amount for the selected user. This corresponds to the "profit" in the MAB problem. .

[0069] (3) Each user They will follow different probability distribution models These distributions are unknown to microgrid operators.

[0070] Furthermore, the optimization objectives of resident load aggregation and the MAB problem are different: the player's goal is to find the optimal arm and maximize their expected total revenue over all rounds, as shown in Equation (13). (13) In the formula Let be the expected function. It refers to the number of game rounds. Is it a selection arm? The generated random rewards.

[0071] The system operator is there to select the optimal set of users. Sending commands to make the actual power adjustment as close as possible to the target. To reduce the deviation between the two The expected value is shown in equation (14).

[0072] (14) This paper proposes a method for solving the multi-armed slot machine problem that considers the upper confidence bound (UCB). The basic process of the UCB algorithm is as follows: (1) Confidence interval calculation: based on the arm Historical average Calculate its confidence interval centered at [center]. The width of this range reflects the uncertainty of returns; the wider the range, the less we know about the arm.

[0073] (2) Index value calculation: At the e-th decision, for each arm Calculate an index value The index value is used to assess its expected returns. The formula for calculating the index value is shown in equation (15), where... These are the hyperparameters of the algorithm. Indicates arm The number of times it was selected.

[0074] (15) (3) Arm selection: based on index value The algorithm selects the arm with the highest priority for exploration or development based on its size. In this way, the UCB algorithm achieves an effective balance between exploration (trying unknown arms) and development (selecting known high-yield arms), thus gradually approaching the optimal strategy.

[0075] To account for risk in decision-making, the Mean-Variance Upper Confidence Bound (MV-UCB) algorithm is adopted. This algorithm incorporates the mean-variance model into the MAB model, adding a variance term to the traditional mean-based model to obtain the following model: (16) (17) (18) In the formula, This represents the index value used to sort each user during the e-th aggregation. This represents the mean-variance model. These are the hyperparameters of the algorithm. For the confidence level, Indicates the number of times a user is selected. Let represent the response value when the user is selected for the rth time. Based on the index calculated by equation (16), sort the users so that... Users with higher values ​​are prioritized, and then the response demand is determined based on the distribution network's published requirements. Select user groups satisfy: (19) In the formula This represents the set of users selected during the e-th round of aggregation. Indicates according to After sorting, the j-th The index representing the user. The algorithm framework is as follows:

[0076] Based on this algorithm, the microgrid user interaction model is obtained. Figure 3 . Figure 3 This paper demonstrates an aggregation model for microgrid layer interaction with users based on the MV-UCB algorithm. The workflow is as follows: After receiving dispatch instructions from the distribution network, the microgrid calculates a comprehensive index value for each user and selects users according to this index until the power requirements are met. Then, instructions are sent to the selected users, and their actual response data is collected to update the user model parameters online. Finally, the aggregated total response power is reported to the distribution network, minimizing response deviation by dynamically optimizing the user selection strategy.

[0077] The accuracy and effectiveness of the present invention will be verified through specific calculation examples below.

[0078] Example Analysis: Taking an improved IEEE 33-node network as an example, the effectiveness of the proposed multi-level interactive response method of matching-micro-user is verified. The improved IEEE 33-node topology diagram is shown below. Figure 4 . Figure 4The photovoltaic (PV) systems are installed at nodes 2, 7, and 11, with capacities of 300 kW, 600 kW, and 400 kW, respectively. Microgrid 1 is equipped with a 1200 kW PV system, a 500 kW wind turbine, a 1000 kW thermal energy storage boiler, and a 500 kW absorption chiller. Its cooling, heating, and electrical loads are 120 kW, 120 kW, and 650 kW, respectively. Microgrid 2 has the same configuration as microgrid 1. Microgrid 3 is equipped with a 400 kW wind turbine, with a load of 90% of that of microgrid 1, and the remaining configuration is the same as microgrid 1. Microgrids 1 and 2 have a combined energy storage capacity of 500 kWh, while microgrid 3 has a combined energy storage capacity of 400 kWh, with a charge / discharge rate of 0.5 for all energy storage systems.

[0079] Consider a microgrid containing n=200 residential users, assuming each user has a contracted power regulation capacity of 0.5 kW. Set TMAB = 500 aggregation events, with each distribution network issuing a power response demand of 40 kW to the microgrid. Use a Beta distribution to randomly generate the user load response probabilities within the microgrid, satisfying the average value of the overall response. To verify the improvement in power response accuracy of the proposed online learning algorithm for microgrid aggregation users, a traditional random selection (RS) aggregation method was used as a control. RS aggregation only considers the average response. and the average probability of participating in the response Given a response target, the microgrid randomly selects users from all users. Each user sends a response command; the response result is shown below. Figure 5 . Figure 5 Each point in the table represents the response result. The left sequence represents the microgrid's response result in 500 aggregate events using the MV-UCB algorithm, and the right sequence represents the microgrid's response result in 500 aggregate events using the RS algorithm. The MV-UCB algorithm shows a smaller fluctuation range in response results compared to the RS method, indicating that the MV-UCB algorithm can improve the response accuracy of the microgrid compared to the traditional RS algorithm.

[0080] The percentage of users who rejected the response after it was sent is shown in the image. Figure 6 , Figure 6 The upper curve represents the rejection rate of users selected by the RS algorithm, while the lower curve represents the rejection rate of users selected by the MV-UCB algorithm. After approximately 100 online learning iterations, the user group that should be selected for the 40 kW response target was basically determined. The MV-UCB algorithm can accurately identify users with a higher willingness to respond, and the rejection rate decreases, but it still fluctuates around 5%, which is due to the inherent uncertainty of residents' responses.

[0081] Based on the day-ahead dispatch data, considering disturbances within 95% of each time period, Monte Carlo sampling was used to generate intraday power deficit data, resulting in 4800 sets of data for training. In the initial training phase, the agent performs random actions to accumulate experience. As training progresses, the agent gradually matures and can ultimately provide more accurate guidance to the distribution network. To verify the effectiveness of the algorithm, the following experimental scenario was set: Following the day-ahead dispatch plan, the distribution network experiences a 75kW power deficit in photovoltaic power during the 60th time period of the day, while the load increases by 70kW, resulting in a total power deficit of 150kW in the distribution network. Different methods were used to generate intraday dispatch plans at this time: Solution 1: The method proposed in this invention.

[0082] Option 2: Microgrids participate in power response, but the uncertainty of the distribution network is not considered.

[0083] Option 3: Power balancing is achieved solely by the transmission network.

[0084] The balancing power provided by each microgrid and transmission network under the three schemes is shown in the figure. Figure 7 The bar chart, from bottom to top, represents the power provided by the transmission network, and the power provided by microgrids 1, 2, and 3. All three schemes can compensate for the real-time power deficit in the distribution network. Due to network losses, the increased power from all three schemes exceeds the power demand within the distribution network. Schemes 1 and 2 utilize microgrids in power response, reducing dependence on the transmission network. Furthermore, considering the uncertainty of microgrid response, Scheme 1 rationally allocates the response targets of each microgrid, improving the response target of microgrid 2 and enhancing the overall response accuracy of the distribution system, thereby reducing the reliance on the transmission network for power balancing.

[0085] It should be emphasized that the embodiments described in this invention are illustrative rather than limiting. Therefore, this invention includes, but is not limited to, the embodiments described in the specific implementation. Any other implementations derived by those skilled in the art based on the technical solutions of this invention are also within the scope of protection of this invention.

Claims

1. A real-time power interaction method for distribution-micro-user considering user response uncertainty, characterized in that: Includes the following steps: Step 1: Characterize the dynamics of distribution network dispatch using Markov decision processes; A microgrid interaction model is constructed based on the Actor-Critic framework; the agent is trained using historical experience data, and scheduling instructions to be issued to the microgrid group are generated by combining the day-ahead scheduling results and real-time power status. Step 2: Based on the microgrid group dispatch instructions issued by the distribution network in Step 1, each microgrid executes them in parallel. A microgrid user interaction model is constructed according to the multi-armed slot machine theory. An online learning algorithm with fusion upper confidence interval strategy is used to solve the model. User combinations that can respond to the microgrid group dispatch instructions issued by the distribution network in Step 1 are dynamically selected to execute the instructions in order to minimize the microgrid response deviation. Response instructions are sent to users, and the actual response data of users is collected to continue updating the parameters of the online learning algorithm. The actual response data is summarized and uploaded to the distribution network, thereby completing the real-time power interaction between distribution, microgrids, and users.

2. The real-time power interaction method for distribution-micro-user considering user response uncertainty according to claim 1, characterized in that: The specific steps of step 1 include: (1) The real-time scheduling problem of the distribution network is modeled as a Markov decision process (MDP), and an MDP model that can describe the real-time scheduling decision problem of the distribution network is established. (2) Based on the MDP model established in step (1) of step 1, construct an agent containing an Actor and a Critic network, and train it to learn the optimal strategy for solving the MDP problem. (3) Based on the agent trained in step (2) of step 1, deploy it in the MDP model constructed in step (1) of step 1 for online application; the distribution network layer inputs the real-time collected system status and day-ahead scheduling results into the agent, and the agent outputs a set of power scheduling instructions that take into account the microgrid response capability for the current power deficit and system status and sends them to the microgrid group.

3. The real-time power interaction method for distribution-micro-user considering user response uncertainty according to claim 2, characterized in that: The specific steps of step 1 (1) include: For distribution networks: ① State: Contains the state of each node. Active power during operation before the time period reactive power Each node at Power deficit during the period and the actual response power of each microgrid. , is represented as: (1) In the formula, Active power The set, reactive power The set, Power deficit The set, For true response power A set; ②Action: The set of actions is: (2) In the formula For the response objectives of each microgrid The set of elements must satisfy the following constraints: (3) In the formula For microgrids Maximum response power; ③Profit: The revenue of the distribution network is: (4) An action-value function is needed. To estimate its profit; this function can measure the profit in a given state. Under the following strategy, the distribution network Select Action The effectiveness, The higher the action, the better; therefore, the agent's optimal strategy It can be represented as: (5) In the formula For action space; ④ State transition probability: The state transition probability in the MDP model is a pattern implicit in historical data, rather than being explicitly defined.

4. The real-time power interaction method for distribution-micro-user considering user response uncertainty according to claim 2, characterized in that: The specific steps of step 1, step (2) include: ① Network initialization: Initialize the Actor network (parameters) respectively. ) and Critic network (parameters) The Actor network is responsible for generating action instructions based on the input state; the Critic network is responsible for evaluating the long-term expected value of the state-action pair. ② Offline training and interaction: The agent interacts extensively with the environment through simulation or historical data; each interaction generates experience data, including state, action, reward and next state, which is stored in the experience replay buffer; ③ Network parameter update: During training, data is sampled from the buffer and the network is updated sequentially. Updating the Critic network: updating parameters by minimizing the temporal difference error. The goal is to make the Critic network's evaluation of state-action value more accurate; and to update the parameters. The direction is to minimize equation (6); Compared to The gradient is shown by equation (7); The update is given by equation (8); (6) (7) (8) In the formula, for Target approximation value Let it be the expected function; This is the gradient calculation function; The learning rate for the Critic network; Updating the Actor network: The update objective of the Actor network depends on the computational results of the Critic network; the updated... for: (9) (10) In the formula, is the learning rate of the Actor network; ④ Introduce adaptive learning rate: Introduce discount factor and To update the learning rate, based on the first... The next learning session The loss value used to determine the goodness of fit of a function to the true value is used to judge its quality. (11) Substituting the discount factor into equation (8) yields the new updated equation: (12)。 5. The real-time power interaction method for distribution-micro-user considering user response uncertainty according to claim 1, characterized in that: The specific steps of step 2 include: (1) The user aggregation problem within the microgrid is modeled as a combined online multi-armed slot machine problem, and a multi-armed slot machine model of microgrid-user interaction is established; (2) Based on the MAB model established in step (1) of step 2, the mean-variance upper confidence interval algorithm is used to make decisions and determine the set of responding users; (3) Based on the user set determined in step (2) of step 2, update the user response and parameters. The microgrid summarizes the actual response quantities of all selected users to calculate the total actual response power of the microgrid and uploads it to the distribution network, thereby completing the real-time power interaction between distribution, microgrid and user.

6. A real-time power interaction method for distribution-micro-user considering user response uncertainty according to claim 5, characterized in that: The specific steps of step 2, step (1) include: ① Put each user Viewed as an arm of a slot machine, the microgrid operator is seen as a player in a casino, whose task is to adjust the power distribution target issued by the distribution network in the e-th aggregation event. Select a suitable user set Send instructions; ② Select user set After that, each user included Feedback will be provided, showing the actual power adjustment amount for the selected user. This corresponds to the "profit" in the MAB problem. ; ③ Each user They will follow different probability distribution models These distributions are unknown to microgrid operators; The player's goal is to find the optimal arm and maximize their expected total gain over all rounds, as shown in equation (13). (13) In the formula Let be the expected function. It refers to the number of game rounds. Is it a selection arm? The generated random rewards; Microgrid operators aim to select the optimal set of users. Sending commands to make the actual power adjustment as close as possible to the target. Reduce the deviation between the two The expected value is shown in equation (14); (14)。 7. A real-time power interaction method for distribution-micro-user considering user response uncertainty according to claim 5, characterized in that: The specific method for step (2) of step 2 is as follows: Calculating the index value: To balance exploring unknown users with utilizing known high-value users, and considering the stability of user responses, the MV-UCB algorithm is adopted. This algorithm integrates the mean-variance model into the MAB model. Based on the traditional model that only considers the mean, a variance term is introduced to measure risk, resulting in the following model: (15) (16) (17) In the formula, This represents the index value used to sort each user during the e-th aggregation. This represents the mean-variance model. These are the hyperparameters of the algorithm. For the confidence level, Indicates the number of times a user is selected. This represents the response value when the user is selected for the rth time; Calculate an index value for each user This value consists of three parts: the sample mean of users' historical response volume, the variance term used to measure risk, and the upper confidence bound term to encourage exploration; User sorting and selection: Sort all users from high to low according to the calculated index value; then, select users in this order until the sum of the contracted capacity of the selected users can meet the dispatch instructions issued by the distribution network, thereby determining the final set of users to be executed. Based on the index calculated using equation (15), each user is sorted, so that... Users with higher values ​​are prioritized, and then the response demand is determined based on the distribution network's published requirements. Select user set ,satisfy: (18) In the formula This represents the set of users selected during the e-th round of aggregation. Indicates according to After sorting, the j-th The index of the user it represents.

8. A real-time power interaction method for distribution-micro-user considering user response uncertainty according to claim 5, characterized in that: The specific method for step 2, step (3) is as follows: microgrids to selected user sets Specific response instructions are sent. Upon receiving the instructions, users independently weigh their own electricity consumption preferences, determine the actual response power, and feed it back to the microgrid. The microgrid collects the actual response quantities of its internal users and uses this data to update the historical statistical data of the responding users, including parameters such as the number of times they were selected, the mean and variance of the response quantity. Through the above process, the actual response power of this round of aggregation is obtained, and online learning of the user response characteristic model is completed. Through continuous learning, the evaluation of user response characteristics is gradually updated. The microgrid summarizes the actual response quantities fed back by all selected users to calculate the total actual response power of the microgrid and uploads it to the distribution network, thereby completing the real-time power interaction between distribution, microgrid, and users.