Large-scale resident distributed resource distribution network voltage regulation method based on DWM-MADRL

CN122801310APending Publication Date: 2026-09-22JINCHENG POWER SUPPLY COMPANY OF STATE GRID SHANXI ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610940132.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]当前主流调压技术存在明显局限:集中式优化方法计算复杂度高、实时性差、高度依赖全局通信,难以适配居民侧资源点多面广、动态变化快的场景;基于固定规则的本地控制策略协同性不足,易出现调控冲突,高渗透场景下调压效果大幅衰减,均无法有效应对居民侧资源高随机、强时空耦合的特性

Benefits of technology

本申请所提方法,解决了传统MADRL依赖海量真实交互数据、训练存在运行风险及场景迁移能力弱的问题,大幅提升了样本效率与策略鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122801310A_ABST
    Figure CN122801310A_ABST
Patent Text Reader

Abstract

The application provides a large-scale resident distributed resource distribution network voltage regulation method based on DWM-MADRL, comprising: constructing a distribution network cooperative voltage regulation target optimization function; constructing a centralized training and decentralized execution MADRL strategy, the strategy network comprising a centralized double Critic network, a decentralized Actor network and a target network; constructing a double-layer world model DWM as a virtual training environment, composed of an upper-layer distribution network power flow dynamic model and a lower-layer resident equipment dynamic model, and the two layers are closed-loop interactive; performing three-stage progressive training, guided by the instant reward defined by the target optimization function, and jointly optimizing the DWM and the MADRL strategy; and deploying the converged decentralized Actor network to each district intelligent agent, collecting local state in real time and outputting optimal voltage regulation instructions. The method solves the problems of traditional MADRL, such as dependence on massive real interaction data, training risk and weak scene migration ability, and greatly improves the sample efficiency and strategy robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of distribution network operation control and artificial intelligence optimization technology, and in particular to a voltage regulation method for large-scale residential distributed resource distribution networks based on DWM-MADRL. Background Technology

[0002] With the goal of increasing new energy sources, the penetration rate of distributed photovoltaic, energy storage, and flexible electricity consumption resources on the residential side in the distribution network has increased significantly. This has completely changed the traditional one-way rigid operation mode of the distribution network, causing problems such as severe voltage fluctuations at nodes and frequent over-limit occurrences at the end of feeders. This seriously threatens the safe and stable operation of the distribution network, and there is an urgent need for efficient and coordinated voltage regulation technology that is adapted to the characteristics of residential side resources.

[0003] Current mainstream voltage regulation technologies have obvious limitations: centralized optimization methods have high computational complexity, poor real-time performance, and are highly dependent on global communication, making them difficult to adapt to scenarios where there are many and widely distributed resources on the residential side and rapid dynamic changes; local control strategies based on fixed rules lack coordination and are prone to regulation conflicts, and the voltage regulation effect is greatly reduced in high-penetration scenarios. All of these cannot effectively cope with the characteristics of highly random and strongly spatiotemporally coupled residential resources.

[0004] While Multi-Agent Deep Reinforcement Learning (MADRL) offers a new path for coordinated voltage regulation, its training requires massive trial-and-error interactions with real distribution networks. This not only results in high training costs and low sample efficiency but also poses safety risks such as system malfunctions and equipment damage, making engineering implementation difficult. Existing solutions that introduce world models to construct virtual training environments still suffer from three core drawbacks: First, they do not embed physical constraints such as power flow and equipment operation, making the model output prone to deviating from actual operating conditions and lacking credibility. Second, they employ a single, homogeneous structure, failing to accurately characterize the heterogeneous characteristics of power flow dynamics and residential equipment behavior, thus limiting model accuracy and generalization ability. Third, their training strategies are simplistic, lacking efficient pre-training and fine-tuning mechanisms, resulting in weak scene transfer and adaptation capabilities. Summary of the Invention

[0005] This application provides a voltage regulation method for large-scale residential distributed resource distribution networks based on DWM-MADRL. To solve the above-mentioned technical problems, this application adopts the following technical methods: This application provides a voltage regulation method for large-scale residential distributed resource distribution networks based on DWM-MADRL, including: Construct the objective optimization function for coordinated voltage regulation in the distribution network; A multi-agent deep reinforcement learning strategy based on MADRL is constructed, and a centralized training and distributed execution architecture is adopted. The policy network corresponding to the multi-agent deep reinforcement learning strategy includes a centralized dual Critic network, a distributed Actor network, a target Actor network, and a target Critic network. A two-layer world model (DWM) is constructed as a virtual training environment. The two-layer world model (DWM) includes an upper-layer power distribution network dynamic model and a lower-layer residential equipment dynamic model, and a closed-loop data interaction is formed between the two models. A three-stage progressive training process is performed, guided by the immediate reward signal defined by the objective optimization function, to jointly optimize the two-layer world model (DWM) and the multi-agent deep reinforcement learning strategy. The distributed Actor network, after training and convergence, is deployed to each transformer area agent to collect local status data in real time and output optimal voltage regulation control commands.

[0006] Optionally, the constraints of the objective optimization function include the following: Power balance constraints, safe operation constraints, voltage limits, adjustable photovoltaic operation constraints, battery operation constraints, thermal storage tank operation constraints, heat pump and electric heating equipment constraints, and cold and heat energy supply and demand balance constraints.

[0007] Optionally, the upper-level power distribution network dynamic model includes an HGAT-GRU encoder, a Transformer dynamic model, and a task output layer, while the lower-level residential equipment dynamic model includes an MLP-GRU encoder and a lightweight Transformer dynamic model.

[0008] Optionally, the three-stage progressive training includes: The training process consists of a pre-training phase based on historical data, a fine-tuning phase based on real interactive data, and a collaborative training phase that alternates between virtual and real data. The fine-tuning phase focuses on optimizing the individual parameters of the dynamic model of the lower-level resident equipment, while the collaborative training phase simultaneously updates the two-layer world model (DWM) and the policy network.

[0009] Optionally, the process of constructing the instant reward signal includes: The problem of coordinated voltage regulation in the distribution network is modeled as a multi-agent Markov decision process, in which each residential transformer area corresponds to an agent. The state space includes the energy storage status of the equipment, photovoltaic output and voltage information, and the action space includes the energy storage device and the control commands for energy storage and photovoltaic. The instantaneous reward signal of the multi-agent Markov decision process is defined as the negative value of the objective optimization function.

[0010] Optionally, the pre-training phase based on historical data includes: Randomly sample continuous time-series data of a preset duration from the historical running dataset and split it into an upper-layer power flow dataset and a lower-layer device dataset; Initialize the globally shared parameters of the two-layer world model (DWM) and the personalized parameters of the lower-layer resident equipment dynamic model, while freezing the relevant parameters of the policy network; Input the upper-layer power flow dataset into the upper-layer distribution network power flow dynamic model and output the first prediction instant reward, the first prediction task continuous flag, and the first prediction next time upper-layer coding state; input the lower-layer equipment dataset into the lower-layer residential equipment dynamic model and output the first prediction next time lower-layer coding state. Based on the first predicted instant reward and the first real instant reward obtained from the historical running dataset, a first reward prediction loss is constructed, and based on the first reward prediction loss, combined with the first reconstruction loss, the first dynamic loss, the first continuous label loss and the first representation loss, the first total loss function value is obtained by weighting and summing them according to the preset first weight. Based on the total loss function value, parameter backpropagation and update are performed: the gradient of the first total loss with respect to all parameters of the two-layer world model (DWM) is calculated using the backpropagation algorithm, and the shared parameters of the two-layer world model (DWM) are updated using the adaptive moment estimation optimizer; the personalized parameters of the lower-layer resident equipment dynamic model are initially initialized and do not participate in the gradient update for the time being. If the first total loss no longer decreases after a preset number of iterations, or reaches the preset number of pre-training iterations, then this stage ends, the basic general two-layer world model (DWM) is output, and the next training stage, based on model fine-tuning using real interaction data, begins; otherwise, the next iteration begins.

[0011] Optionally, the fine-tuning stage based on real interaction data includes: From the real interaction dataset collected from the target distribution network, time-series data containing "state-action-next state-reward" is randomly sampled for a preset duration; it is then split into upper-layer power flow real interaction dataset and lower-layer device real interaction dataset. Load the basic general two-layer world model (DWM) parameters output from the pre-training phase, freeze the core structural parameters of the upper-layer power flow dynamic model, and retain only the update permissions for the output layer and shared parameters; Input the upper-layer power flow real interaction dataset into the upper-layer power distribution network dynamic model, and output the second prediction prediction instant reward signal, the second prediction task continuous flag, and the upper-layer coding state of the second prediction next moment; input the lower-layer equipment real interaction dataset into the lower-layer residential equipment dynamic model, and output the lower-layer coding state of the second prediction next moment. Based on the second predicted instant reward signal and the second real instant reward obtained from the real interaction dataset, a second reward prediction loss is constructed; The second dynamic loss, the second continuous label loss, and the second representation loss are weighted and summed according to a preset first weight to obtain the first intermediate value; The lower-level device state reconstruction loss in the second reconstruction loss is combined with the second reward prediction loss, and the second intermediate value is obtained by weighting and summing them according to a preset second weight; the second weight is greater than the first weight. The upper-layer power flow reconstruction loss in the second reconstruction loss is weighted by a preset third weight to obtain a third intermediate value; the third weight is less than the first weight. Add the first intermediate value, the second intermediate value, and the third intermediate value to obtain the second total loss function value; The loss gradient is calculated by backpropagation, and the personalized parameters of the lower-level residential equipment dynamic model are updated first. The shared parameters of the two-layer world model (DWM) are fine-tuned using a learning rate lower than that in the pre-training stage, and the output layer parameters of the upper-level distribution network power flow dynamic model are updated synchronously. If the mean absolute error of the reward prediction is lower than the preset threshold or the preset number of fine-tuning iterations is reached, this stage ends; the target power grid dedicated world two-layer model (DWM) is output, and collaborative training with alternating virtual and real data is performed; otherwise, real data is resampled, the above fine-tuning stage is repeated, and the next iteration begins.

[0012] Optionally, a single iteration of the collaborative training phase, which alternates between virtual and real data, includes reinforcement learning policy updates and two-layer world model (DWM) updates: Training initialization: Build a virtual-real dual experience pool, initialize the distributed Actor network, centralized dual Critic network, target Actor network and target Critic network. The distributed Actor network adopts the same parameter sharing mechanism as the lower-level residential equipment dynamic model, and loads the fine-tuned target power grid dedicated world dual-layer model DWM. Reinforcement learning strategy update: Each distributed Actor network outputs voltage regulation actions based on local observation status, and engages in virtual interaction with the fine-tuned two-layer world model (DWM) to generate virtual experience data of "state-action-reward-next state" and store it in the virtual experience pool. Experience data of a preset batch size is sampled from the virtual experience pool. Using the MATD3 algorithm and a dual-Q network structure, the current Bellman error Q value is calculated based on the global state and the joint actions of all agents. Based on the current Bellman error Q value, all parameters of the centralized dual-Critic network are updated. Determine if the preset number of attempts has been reached; If so, then a value assessment is performed on the current centralized dual-Critic network to obtain the value assessment result corresponding to the current centralized dual-Critic network. Based on the value assessment results, the global shared parameters of the distributed Actor network and the individual parameters of each Actor in each substation area are updated through deterministic policy gradient. Based on the updated distributed Actor network parameters and the current centralized dual Critic network, the target Actor network parameters and the target Critic network parameters are updated using an exponential moving average method. Update the strategy to determine whether the preset number of steps has been completed; If so, output the updated reinforcement learning policy; proceed to the two-layer world model (DWM) update step; Two-layer world model DWM update: The reinforcement learning strategy obtained from the current training update is executed in a real power distribution network for a preset number of steps to collect real operation data and store it in the real experience pool. Sample empirical data of a preset batch size from the real experience pool; split it into upper-layer power flow empirical data and lower-layer equipment empirical data; The upper-level power flow experience data is input into the upper-level distribution network power flow dynamic model, and the third prediction instant reward signal, the third prediction task continuity flag, and the upper-level coding status of the third prediction next moment are output; the lower-level equipment experience data is input into the lower-level residential equipment dynamic model, and the lower-level coding status of the third prediction next moment is output. Based on the third predicted instant reward signal and the third real instant reward obtained from the real experience pool, a third reward prediction loss is constructed. The third reconstruction loss, the third continuous label loss, and the third representation loss are weighted and summed according to the preset first weight to obtain the fourth intermediate value; The third reward prediction loss and the third dynamic loss are weighted and summed according to a preset fourth weight to obtain a fifth intermediate value; the fourth weight is greater than the third weight. The fourth and fifth intermediate values ​​are added together to obtain the third total loss function value; Freeze all reinforcement learning-related parameters, calculate the loss gradient through backpropagation, and update all shared and personalized parameters of the two-layer world model (DWM) with a small learning rate. If the cumulative reward value remains stable above the preset threshold for a preset number of consecutive iterations, and the voltage exceedance rate is lower than the safety standard, then training ends; save the parameters of the converged distributed Actor network and the two-layer world model for real-time coordinated voltage regulation deployment in the distribution network; otherwise, return to the reinforcement learning policy update step and begin the next complete iteration.

[0013] This application has the following beneficial effects: The method proposed in this application solves the problems of traditional MADRL relying on massive amounts of real interaction data, training having operational risks, and weak scene transferability, and significantly improves sample efficiency and policy robustness. Attached Figure Description

[0014] Figure 1A flowchart illustrating the voltage regulation method for a large-scale residential distributed resource distribution network based on DWM-MADRL provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the upper-level power flow dynamic model of the distribution network provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the dynamic model of the lower-level residential equipment provided in the embodiments of this application; Figure 4 shows different scale distribution network bus diagrams provided in the embodiments of this application; Figure 4(a) is a 33-bus network diagram; Figure 4(b) is a 141-bus network diagram; Figure 4(c) is a 322-bus network diagram; Figure 5 shows the dynamic trend of the four core evaluation indicators provided in the embodiments of this application during the 33-bus network training process. Figure 5(a) shows the dynamic trend of the cumulative reward value during the 33-bus network training process; Figure 5(b) shows the dynamic trend of the node average voltage during the 33-bus network training process; Figure 5(c) shows the dynamic trend of the voltage limit violation rate during the 33-bus network training process; Figure 5(d) shows the dynamic trend of the line active power loss during the 33-bus network training process. Figure 6 shows the dynamic trend of the four core evaluation indicators provided in the embodiments of this application during the 141-bus network training process. Figure 6(a) shows the dynamic trend of the cumulative reward value during the 141-bus network training process; Figure 6(b) shows the dynamic trend of the node average voltage during the 141-bus network training process; Figure 6(c) shows the dynamic trend of the voltage over-limit rate during the 141-bus network training process; Figure 6(d) shows the dynamic trend of the line active power loss during the 141-bus network training process. Figure 7 shows the dynamic trend of the four core evaluation indicators provided in the embodiments of this application during the 322-bus network training process. Figure 7(a) shows the dynamic trend of the cumulative reward value during the 322-bus network training process; Figure 7(b) shows the dynamic trend of the node average voltage during the 322-bus network training process; Figure 7(c) shows the dynamic trend of the voltage over-limit rate during the 322-bus network training process; Figure 7(d) shows the dynamic trend of the line active power loss during the 322-bus network training process. Figure 8 is a performance comparison chart showing detailed numerical values ​​of the quantitative evaluation of the various methods provided in the embodiments of this application; Figure 8(a) is a performance comparison chart of the cumulative reward value of each method; Figure 8(b) is a performance comparison chart of the average node voltage of each method; Figure 8(c) is a performance comparison chart of the voltage over-limit rate of each method; Figure 8(d) is a performance comparison chart of the line active power loss of each method. Detailed Implementation

[0015] To facilitate understanding by those skilled in the art, the present application will be further described below in conjunction with embodiments and accompanying drawings. The content mentioned in the embodiments is not intended to limit the present application.

[0016] To solve the above technical problems, such as Figure 1 As shown, this application proposes a voltage regulation method for large-scale residential distributed resource distribution networks based on DWM-MADRL, including: Step S101: Construct the objective optimization function for coordinated voltage regulation of the distribution network; This step establishes a comprehensive optimization objective that considers both voltage quality and operational economy. It incorporates voltage regulation effect and grid losses into the objective function to achieve multi-objective collaborative optimization. The expression of the objective optimization function is as follows: (1) (2) (3) In the formula, This indicates the voltage deviation in the distribution network. This represents the voltage deviation cost factor; Indicates the line loss status of the distribution network; N represents the line loss coefficient; T represents the number of nodes; and T represents the control period. This represents the voltage amplitude at node i during time period t; Indicates the voltage reference value; Let be the current value of line ij during time period t; Let be the resistance value of line ij.

[0017] After clarifying the optimization objectives, in order to ensure that all control strategies meet the actual operating rules of the power grid and equipment, a complete constraint system is established from three levels: distribution network operation, equipment operation, and user energy consumption.

[0018] Distribution network power balance constraints: Power balance constraints ensure that the distribution network meets the supply and demand matching of active and reactive power at any given time. This is a fundamental physical constraint for stable grid operation, and its expression is: (4) In the formula, and These represent the active power and reactive power injected into the photovoltaic system at node i, respectively. and These represent the active power and reactive power of the load at node i, respectively. and Let be the branch conductance and susceptance between nodes i and j, respectively. The voltage between nodes i and j and The phase angle difference.

[0019] Constraints for safe operation of power distribution networks: Safe operation constraints limit voltage, current, and tie-line power to safe ranges to prevent equipment overload and voltage exceeding limits. The expression is: (5) In the formula, and These are the upper and lower limits of the voltage amplitude at node i, respectively; This represents the upper limit of the current amplitude in branch ij; and These are the upper and lower limits of the exchange power between the regional distribution network and the main grid interconnection line, respectively. This represents the maximum power exchanged between the regional distribution network and the main grid at time t.

[0020] Controllable equipment operation constraints: Based on power and safety constraints, independent operating constraints are set for various distributed resources on the residential side to ensure that the equipment works reliably and does not exceed the rated range.

[0021] Adjustable photovoltaic operating constraints: (6) (7) In the formula, This represents the adjustable active power of the photovoltaic system during time period t. This represents the adjustable reactive power of the photovoltaic system during time period t. This indicates the maximum power generation of the photovoltaic system; This indicates the maximum apparent power of the adjustable photovoltaic system.

[0022] Battery operating constraints: (8) In the formula, and These represent the battery's discharge and charging power, respectively. Indicates the state of charge of the battery; and These represent the battery's maximum charging power and maximum discharging power, respectively. and These represent the minimum and maximum states of charge of the battery, respectively.

[0023] Thermal storage tank operating constraints: (9) In the formula, This indicates the storage status of heat and cold energy in the thermal storage tank; This indicates the output cooling and heating capacity of the thermal storage tank; This indicates the cooling and heating capacity input to the thermal storage tank; and These are the upper and lower limits of the heat storage capacity of the thermal storage tank; and These represent the upper limits of heat or cold stored and released by the thermal storage tank during time period t.

[0024] Constraints of heat pumps and electric heating equipment: (10) In the formula, and These represent the power of the heat pump and the electric heating equipment, respectively. and These represent the maximum operating power of the heat pump and the electric heating equipment, respectively.

[0025] Constraints on the balance between supply and demand of heating and cooling energy: In addition to electrical constraints, to ensure a normal energy experience for residents, a cooling and heating energy balance constraint is set. This constraint ensures that residents' cooling and heating load demands are stably met, improving user comfort. The expression is: (11) (12) In the formula, and These represent the cooling capacity provided to residents by the heat pump and the heating capacity provided to residents by the electric heating equipment, respectively. and These represent the residents' cooling load and heating load demands, respectively.

[0026] Step S102: Construct a multi-agent deep reinforcement learning policy based on MADRL, using a centralized training and distributed execution architecture. The policy network of the multi-agent deep reinforcement learning policy includes a centralized dual Critic network, a distributed Actor network, a target Actor network, and a target Critic network. After the goal and constraint system is established, this application transforms the collaborative voltage regulation problem into a standard multi-agent decision-making problem and uses reinforcement learning to achieve autonomous optimization.

[0027] The above-mentioned distribution network collaborative voltage regulation problem is modeled as a multi-agent Markov decision process, where each residential transformer substation corresponds to an agent, which can be represented as a tuple. .in, For the joint state of the agents, This represents the state of agent n during time interval t; For joint actions of intelligent agents, This represents the action chosen by agent n during time period t; P is the state transition probability of the agent. For joint rewards for intelligent agents, Let represent the reward obtained by agent n in time period γ; γ is the reward discount factor. The core function of this section is to transform the objective function into a solvable problem of maximizing cumulative reward in reinforcement learning. That is, by designing a reward function that completely corresponds to the objective function, the reinforcement learning agent can maximize the cumulative reward in a way that is equivalent to minimizing the objective function.

[0028] The state space includes the energy storage status of the device, photovoltaic output, and voltage information. The state of the nth agent in time period t is represented as follows: (13) In the formula, and This represents the energy storage status of the thermal storage tanks for hot and cold water for all residents in residential area n during time period t. This indicates the battery status of all residents in residential area n during time period t; This represents the photovoltaic power generation capacity of all residents in residential area n during time period t; This represents the relative voltage of node i.

[0029] The action space includes control commands for thermal storage devices, energy storage, and photovoltaics. The action of the nth agent in time period t is represented as: (14) In the formula, and These represent the cooling and heating capacity stored or released by all residential thermal storage tanks in residential area n during time period t, respectively. This represents the charging and discharging power of all residential batteries in residential area n during time period t; This represents the reactive power output of all residential photovoltaic systems in residential area n during time period t.

[0030] The training objective of MADRL is to find the optimal policy. To maximize the expected cumulative reward, the immediate reward signal is defined as the negative value of the objective optimization function. (15) In the formula, Representation Strategy The resulting trajectory; T is the number of steps per round. This is an immediate reward signal.

[0031] Step S103: Construct a two-layer world model (DWM) as a virtual training environment. The two-layer world model (DWM) includes an upper-layer power distribution network dynamic model and a lower-layer residential equipment dynamic model, and a closed-loop data interaction is formed between the two models. To avoid over-reliance on massive amounts of power grid data for training, this application introduces a two-layer world model (DWM) as a virtual training environment and designs an efficient training process. The core function of the two-layer world model (DWM) is to construct a high-fidelity virtual environment that can accurately predict changes in the objective function under arbitrary adjustment actions, thereby replacing the real power grid as the training data for reinforcement learning. Its prediction accuracy directly determines whether the policy trained in the virtual environment can minimize the objective function in the real environment.

[0032] The two-layer world model (DWM) adopts a decoupled design between the upper and lower layers. The upper layer, the power flow dynamic model, focuses on the power flow dynamics of the power distribution network and uses a heterogeneous graph attention network-gated cyclic unit (HGAT-GRU composite structure) to extract topological and temporal features. The lower layer, the residential equipment dynamic model, focuses on the dynamics of residential equipment and uses a multilayer perceptron-gated cyclic unit (MLP-GRU) composite structure to process equipment state information. Both layers embed corresponding physical constraints to improve prediction accuracy.

[0033] Since the data processing logic in the two-layer world model (DWM) is basically the same in the later training stages (i.e., pre-training stage, fine-tuning stage, and alternating co-training stage), to avoid repeating the same process in different stages, the data processing procedure is described uniformly below: like Figure 2 As shown, the upper-level power distribution network dynamic model includes an HGAT-GRU encoder, a Transformer dynamic model, and a task output layer. The data processing process is explained in detail based on the above structure: Inputs to the upper-level distribution network power flow dynamic model: distribution network power flow status (node ​​voltage amplitude / phase angle, branch active / reactive power flow), topology adjacency matrix, and joint voltage regulation actions of all transformer substations.

[0034] After preprocessing, the above input is fed into the HGAT-GRU encoder for topological and temporal feature extraction. The specific process is as follows: First, a three-level multi-head attention calculation is performed on the above input: The node and edge features in the power flow state of the distribution network are used to learn the type interaction weights of node types and electrical attributes through multi-head type attention, and the attribute interaction weights of node types and electrical attributes through multi-head attribute attention. These attribute interaction weights and type exchange weights are then weighted and aggregated with the topological adjacency matrix to generate a global topological feature sequence.

[0035] The global topological feature sequence is then temporally encoded: it is input into a GRU encoder to capture the temporal evolution of power flow dynamics and outputs the upper-level encoded state. This upper-level encoded state is then input into a symmetric GRU decoder to output the reconstructed original power flow state, which is used to calculate the reconstruction loss.

[0036] The upper-layer encoded state is input into the Transformer dynamic model for state transition prediction. The specific process is as follows: Positional encoding is added to the upper-level encoding state. The input is a two-layer stacked Transformer Encoder, which undergoes multi-head self-attention, layer normalization, and MLP processing to generate high-order temporal dependency features. These high-order temporal dependency features are then input into a specially designed MLP (Power Flow Physical Layer). Power balance constraints, safe operation constraints, and controllable equipment operation constraints (including adjustable photovoltaic operation constraints, battery operation constraints, thermal storage tank operation constraints, heat pump and electric heating equipment constraints, and cold and heat energy supply and demand balance constraints), power balance equations, voltage safety constraints, and branch current constraints are embedded as hard constraints into the network output, forcibly generating upper-level hidden states that conform to physical laws.

[0037] The hidden state of the upper layer is input into the task output layer, and through a single independent fully connected layer, the following are output: the predicted task continuity flag (to determine whether the scheduling cycle has terminated), the predicted instant reward value (a comprehensive index of voltage deviation and line loss), and the predicted upper layer encoding state at the next moment.

[0038] like Figure 3 As shown, the dynamic model of the lower-level residential equipment includes an MLP-GRU encoder and a lightweight Transformer dynamic model. Based on the above structure, the specific data processing procedure is explained in detail: The local status and corresponding voltage regulation actions of each residential transformer substation are input into the MLP-GRU encoder for equipment feature extraction. Specifically, the input parameters are shared across the MLP layer, with all substations sharing common MLP parameters. This learns the common input-output characteristics of residential equipment while retaining unique parameters for each substation, adapting to local equipment differences and user behavior. This results in a multi-device-dimensional equipment feature tensor. This feature tensor is then input into the GRU encoder to capture the temporal fluctuations in load and equipment status, outputting the lower-level encoded state. Finally, the lower-level encoded state is input into the symmetrical GRU decoder to output the reconstructed original equipment state.

[0039] The lower-level encoded state is input into the lightweight Transformer dynamic model for equipment state prediction. That is, the lower-level encoded state is added to the position encoding, input into a layer 1 lightweight Transformer Encoder, processed by a standard Transformer block, and output through the MLP (Residential Equipment Layer). The battery charging and discharging constraints, thermal storage tank capacity constraints, temperature control equipment power constraints, and cold and heat energy balance constraints are embedded, and the lower-level hidden state is output.

[0040] The hidden state of the lower layer is then passed through a fully connected layer to output the predicted encoded state of the lower layer at the next time step.

[0041] The upper and lower layer models operate synchronously based on the unified dispatch time of the distribution network: After the lower-level residential equipment dynamic model completes its local equipment status prediction, it uploads the total injected power of the transformer area to the upper-level distribution network power flow dynamic model. The upper-level distribution network power flow dynamic model, based on the global power flow and the injected power of all transformer areas, completes global power flow prediction and then downloads the voltage of each transformer area node to the corresponding lower-level residential equipment dynamic model. The lower-level residential equipment dynamic model corrects its local prediction results based on the received voltage information, forming a closed-loop interaction. All interaction interfaces are differentiable and support end-to-end joint training.

[0042] The training of the Two-Layer World Model (DWM) aims to minimize the total loss function, which is a weighted sum of five components: reconstruction loss, reward loss, continuous labeling loss, dynamic loss, and representation loss. All loss terms ultimately aim to make the DWM a high-precision, physically reliable simulator of the objective function. By minimizing the weighted sum of these five losses, the DWM model can accurately simulate the "action-state-objective function" mapping relationship in a real power distribution network. This supports agents in safely and efficiently exploring optimal voltage regulation strategies in a virtual environment, ultimately achieving the comprehensive minimization of voltage deviation and line losses in a real power grid.

[0043] The total loss function is expressed as follows: (16) In the formula, Batch size; The sequence length; and These are the weights used to balance the various losses in the hyperparameter.

[0044] The total loss function is composed of the following components: Reconstruction loss: Reconstruction loss This is used to constrain the model's ability to accurately reproduce the full-dimensional operating state of the distribution network.

[0045] (17) In the formula, and These represent the reconstructed and actual values ​​of the power flow state of the upper-level distribution network, respectively. and These represent the reconstructed and actual values ​​of the equipment status of lower-level residents, respectively. The loss is measured by mean squared error to ensure that the potential space retains key information from the original observations.

[0046] Reward prediction loss: Predicting reward loss This is used to ensure that the reward and punishment signals output by the virtual environment are completely matched with the voltage regulation and optimization targets of the real power grid.

[0047] (18) In the formula, and These represent the predicted instant reward value from the upper layer and the actual instant reward value, respectively.

[0048] Continuous label loss: Continuous label loss It is used to standardize the model's accurate judgment of the timing logic of continuous dispatching of the distribution network.

[0049] (19) In the formula: and These represent the termination flags of the upper-level prediction and the actual outcome, respectively. A binary cross-entropy loss is employed to ensure accurate prediction of whether the task has terminated.

[0050] Dynamic loss: Dynamic loss It is used to measure the fitting accuracy of the core constraint model to the state transition law of the power grid under voltage regulation.

[0051] (20) In the formula, This represents the encoding status output by the upper-layer encoder. The encoding state predicted by the upper-level distribution network power flow dynamic model; This refers to the encoding state output by the lower-level encoder. The coded state is predicted by the dynamic model of the equipment for lower-level residents.

[0052] Indicates loss: Indicates loss Used to maintain the temporal smoothness of the extracted runtime features to ensure the stability of subsequent reinforcement learning training.

[0053] (twenty one) In the formula, This represents the encoding state output by the upper-layer encoder at the previous moment. This represents the encoding state output by the lower-level encoder at the previous moment.

[0054] The aforementioned loss functions were used in all three subsequent training phases. The main difference between the phases lies in the need to adjust the weight coefficients of each loss function according to the training objective; therefore, this will not be elaborated upon here. Furthermore, the adjectives such as "Class I" and "Class II" added before the function names are only used to distinguish the loss functions in different phases and do not change the essential content of the functions themselves—these functions are completely equivalent to the aforementioned definitions in mathematical form and core calculation.

[0055] Step S104: Perform three-stage progressive training, guided by the immediate reward signal defined by the objective optimization function, to jointly optimize the two-layer world model (DWM) and the multi-agent deep reinforcement learning strategy; The three-stage progressive training here includes a pre-training stage based on historical data, a fine-tuning stage based on real interaction data, and a collaborative training stage that alternates between virtual and real data. The fine-tuning stage focuses on optimizing the individual parameters of the dynamic model of the lower-level resident equipment, while the collaborative training stage simultaneously updates the two-layer world model (DWM) and the policy network. The specific training process is as follows: Pre-training phase based on historical data: The core objective of this training phase is to enable the two-layer world model (DWM) to learn the general operating rules of the distribution network and residential equipment groups. Training is performed only on the two-layer world model (DWM), while the reinforcement learning policy network remains frozen throughout the process. The base two-layer world model (DWM) output from this phase will serve as the initial model for the next phase of fine-tuning.

[0056] Continuous time-series data of a preset duration are randomly sampled from the historical operation dataset and split into upper-level power flow dataset (including node voltage amplitude and phase angle, active and reactive power flow of branches, and injected power of each transformer area) and lower-level equipment dataset (including distributed photovoltaic output, battery state of charge, thermal storage tank energy storage status, heating and cooling loads, and power of temperature control equipment) according to the input requirements of the two-layer world model (DWM).

[0057] Next, parameter initialization and freezing are performed based on the aforementioned historical runtime dataset. This includes initializing the globally shared parameters of the two-layer world model. With lower-level personalized parameters At the same time, all reinforcement learning-related parameters (including distributed Actor networks, centralized dual Critic networks, and corresponding target networks) are frozen to prevent non-converged policies from interfering with the world model's learning of fundamental physical laws.

[0058] Input the upper-layer power flow dataset into the upper-layer power distribution network dynamic model and output the first prediction instant reward, the first prediction task continuous flag, and the first prediction next time upper-layer coding state; input the lower-layer equipment dataset into the lower-layer residential equipment dynamic model and output the first prediction next time lower-layer coding state. Based on the first predicted immediate reward and the first real immediate reward obtained from the historical running dataset, a first reward prediction loss is constructed. Based on the first reward prediction loss, combined with the first reconstruction loss, the first dynamic loss, the first continuous label loss and the first representation loss, the first total loss function value is obtained by weighting and summing them according to the preset first weights. This quantifies the deviation between the model prediction result and the real system and provides the gradient direction for parameter updates.

[0059] Parameter Update: Based on the total loss function value, backpropagation and updating of parameters are performed. The gradient of the first total loss with respect to all parameters of the two-level world model (DWM) is calculated using the backpropagation algorithm, and the Adam adaptive moment estimator optimizer is used to update all shared parameters of the two-level world model (DWM). ; Personalized parameters only for the lower-level world model Initialize the model by assigning values, but do not participate in gradient updates yet. Optimize the model parameters through gradient descent, allowing the two-layer world model (DWM) to gradually fit the general operating rules.

[0060] Iteration Termination Judgment: Based on the updated model parameters, perform an iteration judgment. If the first total loss no longer decreases after a preset number of consecutive iterations, or reaches the preset number of pre-training iterations, then this stage ends; this stage outputs the basic general two-layer world model (DWM) and enters the next training stage for model fine-tuning based on real interaction data; otherwise, resample historical data and repeat the above process until the basic two-layer world model (DWM) converges and has basic environmental dynamic prediction capabilities.

[0061] Fine-tuning phase based on real interaction data: The core objective of this training phase is to calibrate the world model using limited real-world interaction data from the target distribution network, eliminating dynamic discrepancies between the virtual environment and the actual system, with a focus on optimizing the personalized parameters of the dynamic model for lower-level residential equipment. The target power grid-specific two-layer world model (DWM) output in this phase will serve as the virtual environment for the next phase of collaborative training.

[0062] Real data sampling: Randomly sample time-series data of "state-action-next state-reward" for a preset duration from the real interaction dataset collected from the target distribution network. This data is then split into upper-level power flow real interaction dataset and lower-level device real interaction dataset to obtain individual characteristic data of the target distribution network.

[0063] Model Loading and Parameter Freezing: Based on the aforementioned real-world interactive dataset, the basic two-layer world model (DWM) output from the pre-training phase is loaded and its parameters are frozen. The core structural parameters of the upper-layer distribution network power flow dynamic model (HGAT-GRU encoder, Transformer dynamic model main layer) are frozen, retaining only all shared parameters of the upper-layer distribution network power flow dynamic model output layer and the two-layer world model (DWM). and lower-level personalized parameters With the update permissions granted, the reinforcement learning parameters remain completely frozen, preserving general physical laws while avoiding overfitting.

[0064] Forward Propagation and Bias Calculation: The real-world interaction dataset is input into the configured model, and forward propagation and bias calculation of the two-layer world DWM model are performed. The forward propagation process of the two-layer world DWM model, consistent with the pre-training phase, is executed, inputting the real state and action sequence, outputting the prediction results, and calculating the prediction bias of the pre-trained model in the target power grid. This identifies the differences between the basic general world model and the target real system, providing a basis for subsequent parameter adjustments. Specifically, the real-world interaction dataset of the upper-layer power flow is input into the upper-layer distribution network power flow dynamic model, outputting the second prediction instant reward signal, the second prediction task continuity flag, and the upper-layer coded state at the next time step of the second prediction; the real-world interaction dataset of the lower-layer equipment is input into the lower-layer residential equipment dynamic model, outputting the lower-layer coded state at the next time step of the second prediction.

[0065] Dynamic adjustment and calculation of loss weights: The second dynamic loss, the second continuous flag loss, and the second representation loss are weighted and summed according to a preset first weight to obtain a first intermediate value; the lower-level equipment state reconstruction loss and the second reward prediction loss in the second reconstruction loss are weighted and summed according to a preset second weight to obtain a second intermediate value; the second weight is greater than the first weight; the upper-level power flow reconstruction loss in the second reconstruction loss is weighted according to a preset third weight to obtain a third intermediate value; the third weight is less than the first weight; the first intermediate value, the second intermediate value, and the third intermediate value are added together to obtain the second total loss function value. In other words, increasing the weights of the lower-level residential equipment state reconstruction loss and the upper-level reward prediction loss, and decreasing the weight of the upper-level distribution network power flow reconstruction loss, guides the model to prioritize learning the actual operating characteristics of the distributed equipment in the target area, emphasizes the fitting accuracy to the voltage regulation optimization target, and outputs the second total loss function value after weight adjustment.

[0066] Parameter Differentiation Update: Based on the adjusted weighted second total loss function value, a parameter differentiation update is performed. This involves calculating the loss gradient through backpropagation, prioritizing the updating of the individual parameters of the lower-level resident equipment dynamic model, while updating the shared parameters of the two-layer world model (DWM). Fine-tuning is performed using a learning rate lower than that used in the pre-training phase; the degree of fine-tuning is generally small. The output layer parameters of the upper-level distribution network power flow dynamic model are updated synchronously. The dynamic deviation between the virtual environment and the real system is accurately calibrated, enabling the two-layer world model (DWM) to accurately simulate the operating characteristics of the target distribution network.

[0067] Iteration termination judgment: Based on the calibrated model parameters, an iteration termination judgment is performed. If the mean absolute error of the reward prediction is lower than a preset threshold or the preset number of fine-tuning iterations is reached, this stage ends, the target power grid-specific world model is output, and the third stage of alternating virtual and real data collaborative training begins; otherwise, real data is resampled and the above process is repeated to enter the next iteration until the prediction accuracy of the world model for the target power grid meets the requirements of reinforcement learning training.

[0068] A collaborative training phase involving alternating virtual and real data: This stage is the final optimization phase of the overall strategy. It adopts a hybrid approach of virtual simulation data and real environment data to simultaneously complete the alternating iterative updates of the two-layer world model (DWM) and the multi-agent reinforcement learning strategy. A single complete iteration consists of two steps: reinforcement learning strategy update and two-layer world model (DWM) update. This process is repeated until training converges.

[0069] Training Initialization: A virtual-real dual experience pool is built, and a distributed Actor network, a centralized dual Critic network, a target Actor network, and a target Critic network are initialized. The distributed Actor network adopts a parameter sharing mechanism consistent with the dynamic model of the underlying residential equipment (global shared parameters + localized parameters). The finely tuned target power grid-specific world model is loaded. A complete collaborative training framework is built, using the dual-layer world model (DWM) as the virtual training environment, providing a safe and efficient interactive platform for reinforcement learning.

[0070] Reinforcement learning strategy update: Each distribution area's decentralized Actor network outputs voltage regulation actions based on local observations, and virtually interacts with the fine-tuned two-layer world model (DWM) to generate virtual experience data of "state-action-reward-next state" and store it in a virtual experience pool; providing massive, safe, and low-cost training samples. This solves the problem of traditional MADRL's high dependence on real data.

[0071] Experience data of a preset batch size is sampled from the virtual experience pool. A dual-Q network structure based on the MATD3 algorithm is used to calculate the current Bellman error Q value based on the global state and the joint actions of all agents. Based on the current Bellman error Q value, all parameters of the centralized dual-Critic network are updated. This improves the value assessment accuracy of the centralized Critic network, eliminates the overestimation bias of Q value, and provides accurate collaborative guidance for the Actor network.

[0072] Determine if the preset number of attempts has been reached; If so, a value assessment is performed on the current centralized dual-crit network to obtain the corresponding value assessment result. Based on this value assessment result, the globally shared parameters of the distributed actor network and the personalized parameters of each actor in each distribution area are updated through a deterministic policy gradient. This ensures global collaborative optimization while adapting to the differences in equipment characteristics and user behavior in different distribution areas.

[0073] Based on the updated parameters of the distributed Actor network and the current centralized dual Critic network, the parameters of the target Actor network and the target Critic network are updated using an exponential moving average method. This maintains the stability of the learning objective and avoids oscillations during the training process.

[0074] Update the strategy to determine whether the preset number of steps has been completed; If yes, output the updated reinforcement learning policy and proceed to the two-layer world model (DWM) update step; otherwise, continue with the reinforcement learning policy update process.

[0075] Two-layer world model DWM update: Real-world interaction and experience storage: The reinforcement learning policy obtained from the current training is executed in a real power distribution network for a preset number of steps, real-world operation data is collected and stored in a real-world experience pool, and the latest system dynamic data is obtained for continuous model calibration.

[0076] Experience sampling and forward propagation: Experience data of a predetermined batch size is sampled from the real experience pool; this data is then split into upper-level power flow experience data and lower-level equipment experience data. The upper-level power flow experience data is input into the upper-level distribution network power flow dynamic model, outputting a third-prediction instant reward signal, a third-prediction task continuity flag, and the upper-level coding state at the next moment of the third prediction. The lower-level equipment experience data is input into the lower-level residential equipment dynamic model, outputting an evaluation of the lower-level coding state at the next moment of the third prediction. The consistency between the current two-layer world model (DWM) and the real system is assessed, and parameters requiring further calibration are identified.

[0077] Loss Calculation and Update: Based on the third predicted immediate reward signal and the third real immediate reward obtained from the real experience pool, a third reward prediction loss is constructed. The third reconstruction loss, the third continuous label loss, and the third representation loss are weighted and summed according to a preset first weight to obtain a fourth intermediate value. The third reward prediction loss and the third dynamic loss are weighted and summed according to a preset fourth weight to obtain a fifth intermediate value; the fourth weight is greater than the third weight. The fourth intermediate value and the fifth intermediate value are added to obtain the third total loss function value. The standard loss weight configuration of the pre-training stage is restored, and the weights of the reward loss and dynamic loss are appropriately increased to enhance the model's prediction accuracy for voltage regulation effects and state transition patterns.

[0078] Freeze all reinforcement learning-related parameters, calculate the loss gradient through backpropagation, and update all shared parameters of the two-layer world model (DWM) with a small learning rate. With personalized parameters ; Overall training termination judgment: Based on the updated world model and reinforcement learning strategy, if the cumulative reward value remains stable above the preset threshold for a preset number of consecutive iterations and the voltage exceedance rate is lower than the safety standard, then training ends, and the parameters of the distributed Actor network and the two-layer world model (DWM) are saved for real-time coordinated voltage regulation deployment in the distribution network. Otherwise, the process returns to the reinforcement learning strategy update step and begins the next complete iteration.

[0079] Step S105: Deploy the trained and converged distributed Actor network to each transformer area agent, collect local status in real time, and output voltage regulation control commands.

[0080] After multiple rounds of iterative training, the distributed Actor networks adopted by the agents in each distribution area have gradually converged to a stable strategy, meeting the corresponding deployment conditions. At this point, the trained Actor networks can be deployed to the terminal devices of the agents in each distribution area. In actual operation, each agent in each distribution area uses local sensors (such as voltage transformers, current transformers, power monitoring modules, etc.) to collect key operating statuses of its area in real time, including feeder voltage, load current, distributed power output, and reactive power compensation device status. The agent uses the local observation data at the current moment as input and passes it to the Actor network running in the embedded or edge computing environment. After forward computation, the Actor network directly outputs the optimal voltage regulation control command for the distribution area, such as the tap position of the on-load tap changer, capacitor bank switching commands, and reference values ​​of the static var generator (SVG). The entire decision-making process does not require real-time communication between distribution areas and relies entirely on local information, thereby reducing communication bandwidth dependence and improving response speed. In this way, each agent in each distribution area can dynamically and adaptively adjust the voltage according to local conditions, achieving rapid and coordinated control of the voltage level of the entire distribution network.

[0081] Simulation verification: In the specific simulation verification stage, performance tests were conducted on three different scale standard distribution network test cases: IEEE 33 nodes (Figure 4(a)), IEEE 141 nodes (Figure 4(b)), and IEEE 322 nodes (Figure 4(c)). 32, 140, and 321 independent control agents were set up respectively to cover small-scale, medium-scale, and large-scale distribution network operation scenarios. This scheme adopts a distributed partitioning method where each agent corresponds to a single residential transformer substation. It relies on a zoned local control mode to achieve hierarchical scheduling of flexible resources across the entire region, fully conforming to the engineering layout of actual distribution network substation management.

[0082] To comprehensively and objectively demonstrate the technical advantages of the proposed method (DWM-MATD3 distributed voltage control scheme), six mainstream, cutting-edge continuous control reinforcement learning algorithms were selected as comparative benchmarks during the experiment. These included MATD3, MADDPG, SQDDPG, IPPO, MAPPO, and IDDPG methods. All comparative tests maintained completely consistent grid topology, load time-series curves, distributed energy output characteristics, equipment operating parameters, and constraint boundary conditions. Only the core control algorithm was replaced to ensure that the experimental variables were singular and the comparison results were rigorously referential.

[0083] All simulation experiments were conducted using the PyTorch deep learning framework to complete model building, training iteration, and policy deployment. The experimental hardware platform information was as follows: CPU 13th Gen Intel(R) Core(TM) i7-13700K, GPU NVDIAGeForce RTX 4070.

[0084] The experiment used two sets of historical data from actual power distribution networks spanning two years, with a data sampling interval of 3 minutes. The first set of data was used for the pre-training process of the two-layer world model, while the second set of data was used uniformly for model fine-tuning, collaborative training, and training of various comparative algorithms. The dataset includes the active power of photovoltaic power plants, as well as the cooling load, heating load, electrical load, rooftop photovoltaic active power, and total reactive power of residential transformer substations. To meet the safety operation specifications of the power grid, the experiment limited the node voltage fluctuation range to ±5% of the rated voltage. The multi-agent reinforcement learning parameters and the hyperparameters of the two-layer world model involved are summarized in Tables 1 and 2, respectively.

[0085] Table 1 MATD3 Parameter Settings ; Table 2 World Model Parameter Settings ; Figure 4 visually illustrates the network topology of IEEE 33, 141, and 322 nodes. Each node is identified by a numbered circle, clearly demonstrating the topological connections of distribution networks of different scales. During the testing phase, three sets of different random seeds were used to conduct repeated comparative experiments to avoid interference from random factors and ensure the objectivity, stability, and reliability of the test results of the proposed method and the comparison algorithm.

[0086] To comprehensively evaluate the integrated control performance of the proposed method and various comparative algorithms in different scale distribution network scenarios, this experiment selected four core evaluation indicators for quantitative analysis. Among them, the cumulative reward value (Rewad) is used to comprehensively reflect the global optimization effect and overall control adaptability of the control strategy; the average node voltage is used to quantify the system's ability to maintain a stable voltage level throughout the entire period; the voltage violation rate is a core safety evaluation indicator, directly characterizing the probability that the node voltage exceeds the safe operating range during the control process; and the lines loss is used to evaluate the economic level of distribution network operation.

[0087] The dynamic trends of the above four indicators during model training are plotted in Figures 5, 6, and 7, respectively. The data for each curve in the figures are the average values ​​of the results from three fine-tuning and co-training experiments. For clarity, the curves are presented in Table 3 and Figure 8. Detailed numerical values ​​for the quantitative evaluation of each method's indicators are provided.

[0088] Table 3 Performance Comparison of Different Methods ; As can be seen from the test results in Figures 5, 6, 7, 8 and Table 3, the method proposed in this application demonstrates significant comprehensive performance advantages in all four core evaluation indicators.

[0089] In the IEEE 33-node distribution network test scenario, compared with other comparative algorithms, the method proposed in this application consistently maintains the lowest node voltage exceedance rate and line active power loss throughout the entire test cycle. Simultaneously, its cumulative reward value (Rewad) is closest to zero, fully demonstrating that the method can accurately track and efficiently achieve the preset voltage regulation optimization target. While the MATD3 and MADDPG algorithms possess some cooperative control effects, their ability to accurately suppress voltage exceedances still lags significantly behind the method proposed in this application. The SQDDPG algorithm, due to its soft policy update mechanism, has relatively high algorithm complexity and a significantly slower convergence speed than the method proposed in this application, only outperforming the IPPO and MAPPO algorithms in overall performance. The IPPO algorithm adopts a mode where each agent's policy is optimized independently, incorporating the actions of other agents into the environmental disturbance category, neglecting the cooperative control relationship between multiple agents, resulting in insufficient policy stability and poor overall control performance. As an improvement on IPPO, the MAPPO algorithm adopts a centralized training and distributed execution architecture, which can introduce global runtime information optimization strategies during the training phase. Therefore, its performance is better than that of the IPPO algorithm, but it still cannot reach the comprehensive control level of the method proposed in this application.

[0090] As the scale of distribution network nodes continues to increase, expanding from medium-scale networks with IEEE 141 nodes to large-scale distribution networks with IEEE 322 nodes, the method proposed in this application demonstrates increasingly prominent advantages in regulation performance.

[0091] Taking the cumulative reward value (Rewad) as an example, in the IEEE 322-node large-scale distribution network test scenario, the six comparative algorithms were generally low in overall reward value due to the sharp increase in environmental complexity brought about by the expansion of network scale. Among them, the reward values ​​of MAPPO and IPPO algorithms further decreased compared to the 33-node and 141-node test scenarios, fully demonstrating that the policy learning ability of these two algorithms severely degrades as the system scale increases. Although the overall performance of MATD3, MADDPG, IDDPG, and SQDDPG algorithms is better than MAPPO and IPPO, their reward values ​​also decline significantly with the increase in network scale. In stark contrast, the method proposed in this application maintains a high reward level throughout the entire training cycle.

[0092] The test results of the average voltage index show that with the continuous expansion of the distribution network node scale, the fluctuation range of node voltage under various comparison algorithms has been significantly aggravated. However, the method proposed in this application can ensure that the node voltage of the entire network always remains stable around the rated reference voltage.

[0093] In terms of core safety indicators such as voltage over-limit rate and economic indicators such as line active power loss, the method proposed in this application also demonstrates excellent scenario adaptability and operational robustness. Even though the MATD3 algorithm, after fine parameter tuning, shows suboptimal performance in the comparison scheme, its overall control effect is still significantly weaker than the method proposed in this application, which has the integrated capability of "prediction-planning". This fully highlights the architectural innovation advantage of the method proposed in this application, which uses the two-layer world model (DWM) as a decision-making prior constraint and training acceleration module.

[0094] In summary, the method proposed in this application is the first to deeply integrate a two-layer world model with the MATD3 algorithm, which uses a centralized training and distributed execution architecture, to build a virtual environment-driven collaborative voltage regulation system for power distribution networks. At the same time, it designs a three-stage progressive training process for collaborative training and combines a multi-agent parameter sharing mechanism to solve the pain points of traditional MADRL voltage regulation methods, such as reliance on massive amounts of real interaction data, operational risks during training, and weak scenario transfer capabilities. This significantly improves sample utilization efficiency and policy robustness.

[0095] In some embodiments, this application also provides a computer system including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0096] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding processes in the methods described above in the embodiments of this application; for brevity, further details are omitted here.

[0097] The above embodiments are preferred implementations of this application. In addition, this application can be implemented in other ways. Any obvious substitutions without departing from the concept of this technical solution are within the protection scope of this application.

[0098] To facilitate understanding by those skilled in the art of the improvements made by this application compared to the prior art, some of the accompanying drawings and descriptions have been simplified, and for clarity, some other elements have been omitted from this application. Those skilled in the art should realize that these omitted elements may also constitute the content of this application.

Claims

1. A voltage regulation method for large-scale residential distributed resource distribution networks based on DWM-MADRL, characterized in that, include: Construct the objective optimization function for coordinated voltage regulation in the distribution network; A multi-agent deep reinforcement learning strategy based on MADRL is constructed, and a centralized training and distributed execution architecture is adopted. The policy networks corresponding to the multi-agent deep reinforcement learning strategy include a centralized dual-Critic network, a decentralized Actor network, a target Actor network, and a target Critic network. A two-layer world model (DWM) is constructed as a virtual training environment. The two-layer world model (DWM) includes an upper-layer power distribution network dynamic model and a lower-layer residential equipment dynamic model, and a closed-loop data interaction is formed between the two models. A three-stage progressive training process is performed, guided by the immediate reward signal defined by the objective optimization function, to jointly optimize the two-layer world model (DWM) and the multi-agent deep reinforcement learning strategy. The distributed Actor network, after training and convergence, is deployed to each transformer area agent to collect local status data in real time and output optimal voltage regulation control commands.

2. The method according to claim 1, characterized in that, The constraints of the objective optimization function include the following: Power balance constraints, safe operation constraints, voltage limits, adjustable photovoltaic operation constraints, battery operation constraints, thermal storage tank operation constraints, heat pump and electric heating equipment constraints, and cold and heat energy supply and demand balance constraints.

3. The method according to claim 2, characterized in that, The upper-level power distribution network dynamic model includes an HGAT-GRU encoder, a Transformer dynamic model, and a task output layer, while the lower-level residential equipment dynamic model includes an MLP-GRU encoder and a lightweight Transformer dynamic model.

4. The method according to claim 3, characterized in that, The three-stage progressive training includes: The training process consists of a pre-training phase based on historical data, a fine-tuning phase based on real interaction data, and a collaborative training phase that alternates between virtual and real data. The fine-tuning phase optimizes the individual parameters of the dynamic model of the lower-level resident equipment, while the collaborative training phase synchronously updates the two-layer world model (DWM) and the policy network.

5. The method according to claim 4, characterized in that, The process of constructing the instant reward signal includes: The problem of coordinated voltage regulation in the distribution network is modeled as a multi-agent Markov decision process, in which each residential transformer area corresponds to an agent. The state space includes the energy storage status of the equipment, photovoltaic output and voltage information, and the action space includes the energy storage device and the control commands for energy storage and photovoltaic. The instantaneous reward signal of the multi-agent Markov decision process is defined as the negative value of the objective optimization function.

6. The method according to claim 5, characterized in that, The pre-training phase based on historical data includes: Randomly sample continuous time-series data of a preset duration from the historical running dataset and split it into an upper-layer power flow dataset and a lower-layer device dataset; Initialize the globally shared parameters of the two-layer world model (DWM) and the personalized parameters of the lower-layer resident equipment dynamic model, while freezing the relevant parameters of the policy network; Input the upper-layer power flow dataset into the upper-layer distribution network power flow dynamic model and output the first prediction instant reward, the first prediction task continuous flag, and the first prediction next time upper-layer coding state; input the lower-layer equipment dataset into the lower-layer residential equipment dynamic model and output the first prediction next time lower-layer coding state. Based on the first predicted instant reward and the first real instant reward obtained from the historical running dataset, a first reward prediction loss is constructed, and based on the first reward prediction loss, combined with the first reconstruction loss, the first dynamic loss, the first continuous label loss and the first representation loss, the first total loss is obtained by weighting and summing them according to the preset first weight. The gradient of the first total loss with respect to all parameters of the two-layer world model (DWM) is calculated using the backpropagation algorithm. The shared parameters of the two-layer world model (DWM) are updated using the adaptive moment estimation optimizer. The personalized parameters of the lower-layer resident equipment dynamic model are initially initialized and do not participate in the gradient update for the time being. If the first total loss no longer decreases after a preset number of iterations, or reaches the preset number of pre-training iterations, then this stage ends, the basic general two-layer world model (DWM) is output, and the next training stage, based on model fine-tuning using real interaction data, begins; otherwise, the next iteration begins.

7. The method according to claim 6, characterized in that, The fine-tuning phase based on real interaction data includes: From the real interaction dataset collected from the target distribution network, time-series data containing "state-action-next state-reward" for a preset duration are randomly sampled; this data is then split into upper-layer power flow real interaction dataset and lower-layer device real interaction dataset. Load the basic general two-layer world model (DWM) parameters output from the pre-training phase, freeze the core structural parameters of the upper-layer power flow dynamic model, and retain only the update permissions for the output layer and shared parameters; Input the upper-layer power flow real interaction dataset into the upper-layer power distribution network dynamic model, and output the second prediction prediction instant reward signal, the second prediction task continuous flag, and the upper-layer coding state of the second prediction next moment; input the lower-layer equipment real interaction dataset into the lower-layer residential equipment dynamic model, and output the lower-layer coding state of the second prediction next moment. Based on the second predicted instant reward signal and the second real instant reward obtained from the real interaction dataset, a second reward prediction loss is constructed; The second dynamic loss, the second continuous label loss, and the second representation loss are weighted and summed according to a preset first weight to obtain the first intermediate value; The lower-level device state reconstruction loss in the second reconstruction loss is combined with the second reward prediction loss, and the second intermediate value is obtained by weighting and summing them according to a preset second weight; the second weight is greater than the first weight. The upper-layer power flow reconstruction loss in the second reconstruction loss is weighted by a preset third weight to obtain a third intermediate value; the third weight is less than the first weight. Add the first intermediate value, the second intermediate value, and the third intermediate value to obtain the second total loss function value; The loss gradient is calculated by backpropagation, and the personalized parameters of the lower-level residential equipment dynamic model are updated first. The shared parameters of the two-layer world model (DWM) are fine-tuned using a learning rate lower than that in the pre-training stage, and the output layer parameters of the upper-level distribution network power flow dynamic model are updated synchronously. If the mean absolute error of the reward prediction is lower than the preset threshold or the preset number of fine-tuning iterations is reached, this stage ends; the target power grid dedicated world two-layer model (DWM) is output, and collaborative training with alternating virtual and real data is performed; otherwise, real data is resampled, the above fine-tuning stage is repeated, and the next iteration begins.

8. The method according to claim 7, characterized in that, A single iteration of the collaborative training phase, which alternates between virtual and real data, includes reinforcement learning policy updates and two-layer world model (DWM) updates: Training initialization: Build a virtual-real dual experience pool, initialize the distributed Actor network, centralized dual Critic network, target Actor network and target Critic network. The distributed Actor network adopts the same parameter sharing mechanism as the lower-level residential equipment dynamic model, and loads the fine-tuned target power grid dedicated world dual-layer model DWM. Reinforcement learning strategy update: Each distributed Actor network outputs voltage regulation actions based on local observation status, and engages in virtual interaction with the fine-tuned two-layer world model (DWM) to generate virtual experience data of "state-action-reward-next state" and store it in the virtual experience pool. Experience data of a preset batch size is sampled from the virtual experience pool. Using the MATD3 algorithm and a dual-Q network structure, the current Bellman error Q value is calculated based on the global state and the joint actions of all agents. Based on the current Bellman error Q value, all parameters of the centralized dual-Critic network are updated. Determine if the preset number of attempts has been reached; If so, then a value assessment is performed on the current centralized dual-Critic network to obtain the value assessment result corresponding to the current centralized dual-Critic network. Based on the value assessment results, the global shared parameters of the distributed Actor network and the individual parameters of each Actor in each substation area are updated through deterministic policy gradient. Based on the updated distributed Actor network parameters and the current centralized dual Critic network, the target Actor network parameters and the target Critic network parameters are updated using an exponential moving average method. Update the strategy to determine whether the preset number of steps has been completed; If so, output the updated reinforcement learning policy; Enter the two-layer world model (DWM) update steps; Two-layer world model DWM update: The reinforcement learning strategy obtained from the current training update is executed in a real power distribution network for a preset number of steps to collect real operation data and store it in the real experience pool. Sample empirical data of a preset batch size from a real experience pool; It is broken down into upper-level power flow experience data and lower-level equipment experience data; The upper-level power flow experience data is input into the upper-level distribution network power flow dynamic model, and the third prediction instant reward signal, the third prediction task continuity flag and the upper-level coding status of the third prediction next moment are output. The empirical data of the lower-level equipment is input into the dynamic model of the lower-level resident equipment, and the third prediction of the lower-level coded state at the next moment is output. Based on the third predicted instant reward signal and the third real instant reward obtained from the real experience pool, a third reward prediction loss is constructed. The third reconstruction loss, the third continuous label loss, and the third representation loss are weighted and summed according to the preset first weight to obtain the fourth intermediate value; The third reward prediction loss and the third dynamic loss are weighted and summed according to the preset fourth weight to obtain the fifth intermediate value; The fourth weight is greater than the third weight; The fourth and fifth intermediate values ​​are added together to obtain the third total loss function value; Freeze all reinforcement learning-related parameters, calculate the loss gradient through backpropagation, and update all shared and personalized parameters of the two-layer world model (DWM) with a small learning rate. If the cumulative reward value remains stable above the preset threshold for a preset number of consecutive iterations, and the voltage exceedance rate is lower than the safety standard, then training ends; save the parameters of the converged distributed Actor network and the two-layer world model for real-time coordinated voltage regulation deployment in the distribution network; otherwise, return to the reinforcement learning policy update step and begin the next complete iteration.