Large-scale resident distributed resource collaborative power distribution network voltage regulation method

By using a heterogeneous Transformer network architecture and a distributed centralized training framework, the problems of slow response of traditional voltage regulation devices and low efficiency of MADRL algorithm are solved, realizing efficient collaborative voltage regulation of large-scale distributed residential resources, protecting privacy and meeting real-time decision-making requirements.

CN121906519APending Publication Date: 2026-04-21JINCHENG POWER SUPPLY COMPANY OF STATE GRID SHANXI ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINCHENG POWER SUPPLY COMPANY OF STATE GRID SHANXI ELECTRIC POWER
Filing Date
2025-11-20
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Traditional voltage regulation devices in power distribution networks have long response times and poor continuous voltage regulation capabilities, resulting in insufficient utilization of distributed resources. Furthermore, the existing MADRL algorithm has low learning efficiency and unstable training results in large-scale multi-agent control environments.

Method used

The Actor-Critic network, which adopts a heterogeneous Transformer network architecture and combines distributed and centralized training frameworks, uses residential buildings as the smallest control unit and achieves coordinated pressure regulation of distributed resources in residents through standardized action space and confidence gating mechanism.

Benefits of technology

It improves the utilization efficiency of distributed resources, protects residents' privacy, shortens decision-making time, meets the needs of real-time regulation, and enhances the stability and accuracy of learning and training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121906519A_ABST
    Figure CN121906519A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power distribution network voltage regulation, and discloses a large-scale resident distributed resource collaborative power distribution network voltage regulation method, which is characterized in that an Actor network and a Critic network are respectively replaced by a heterogeneous Transformer, and time-space information of a power distribution network is considered, so that a building-level power distribution network voltage regulation framework based on an ST-Het-TMATD3 algorithm is designed; the pressure regulating process comprises the following steps: setting a standardized action space and confidence; taking the time sequence as the input of the Actor network, and outputting a corresponding action value; taking the action value and the global state sequence as the input of the Critic network, and outputting an evaluation value in combination with the geographic information of the intelligent agent; and based on the action value and the evaluation value, combining a reward function to continuously carry out cooperative voltage regulation. According to the method, a novel heterogeneous mode is adopted, a Transform is used for replacing an Actor network and a Critic network, a novel sequence input method is designed, the method is suitable for a large-scale agent control environment, and the problems that resident distributed resources are not fully utilized, a model voltage regulation method is complex, and multi-agent training efficiency is low are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voltage regulation technology in distribution networks, specifically a voltage regulation method for large-scale residential distributed resource coordination in distribution networks. Background Technology

[0002] In recent years, large-scale distributed photovoltaic (PV) grid connection has brought a series of challenges to the distribution network, including voltage amplitude fluctuations and voltage quality degradation. To address these issues, traditional distribution networks typically employ voltage regulators. However, traditional voltage regulators suffer from long response times, poor continuous voltage regulation capabilities, and the potential for frequent operation to reduce equipment reliability and increase maintenance costs. Compared to traditional voltage regulators, residential distributed resources are characterized by their large scale, fast response speed, and flexible control. Therefore, residential distributed resources are gradually becoming an important resource for enhancing the flexible regulation capabilities of the power system.

[0003] In summary, there is an urgent need for a new technical solution for large-scale collaborative voltage regulation of distribution networks for residential distributed resources, in order to carry out collaborative voltage regulation of distribution networks for residential distributed resources. Summary of the Invention

[0004] The purpose of this application is to provide a distribution network voltage regulation method for large-scale residential distributed resource coordination, so as to solve the inefficiency problem of distributed resources as scheduling objects in the prior art.

[0005] To achieve the above objectives, this application provides a method for voltage regulation in a large-scale residential distributed resource coordination distribution network, the method comprising: A pre-constructed distribution network voltage regulation framework includes residential distributed resources, multiple agents, and the distribution network environment; wherein, the residential distributed resources are based on residential buildings as the smallest control unit; Within the aforementioned distribution network voltage regulation framework, distributed resource-coordinated distribution network voltage regulation is performed based on a decentralized Actor network and a centralized Critic network; wherein, both the Actor network and the Critic network adopt a heterogeneous Transformer network architecture; Standardize the action space and confidence level for heterogeneous distributed residential resources; The local state information and historical data of each agent are combined to obtain a time series. The time series is used as the input of the corresponding Actor network to output the corresponding action value. The action values ​​and global state sequences of the multi-agent network are used as inputs to the corresponding Critic network, and combined with the geographic information of the agents, the corresponding evaluation values ​​are output. Based on the action value and the evaluation value, after training the model centrally using the Critic network in conjunction with the reward function, the coordinated voltage regulation is continuously performed using the distributed execution method of the Actor network; wherein, the reward function includes at least a voltage deviation term, which is used to characterize the voltage deviation of the distribution network.

[0006] Preferably, the residential loads in the distributed residential resources correspond to heterogeneous loads, which include at least adjustable photovoltaics, batteries, and thermal storage tanks.

[0007] Preferably, the reward function further includes a residential electricity purchase cost term, which represents the residential electricity purchase cost; the mathematical expression of the reward function is: in, This refers to the immediate rewards that the environment provides to each agent at each time interval. The willingness to regulate is measured by the cost of electricity purchased by residents after the intelligent agent makes scheduling decisions. For the voltage of each node, and These are the preset weight parameters.

[0008] Preferably, the reward function is used to solve a preset objective function, which aims to minimize the combined cost of electricity purchase for each resident and the voltage deviation of the distribution network; the mathematical expression of the objective function is: in, Indicates residents The cost of purchasing electricity, This represents the electricity purchase cost coefficient. This indicates the voltage deviation in the distribution network. This represents the voltage deviation cost coefficient; and , .

[0009] Preferably, the objective function is subject to preset constraints, including power balance constraints of the distribution network, safe operation constraints of the distribution network, operation constraints of controllable equipment, and heat and cold energy balance constraints of the residential buildings. The mathematical expression for the power balance constraint of the distribution network is: in, For nodes The active power injected by the photovoltaic system For nodes The reactive power injected by the photovoltaic system For nodes Active power of load For nodes Reactive power of load For nodes With nodes The branch conductance between For nodes With nodes The susceptance between them For nodes With nodes voltage between and voltage The phase angle difference; The mathematical expression for the safety operation constraints of the distribution network is: in, for Time period nodes Voltage amplitude, For nodes The upper limit of voltage amplitude, For nodes The lower limit of voltage amplitude for Time-of-day branch The current value, branch road The upper limit of the current amplitude, This represents the maximum exchange power between the regional distribution network and the main grid interconnection line. This is the upper limit of the maximum exchange power between the regional distribution network and the main grid interconnection line. This is the lower limit of the maximum power exchanged between the regional distribution network and the main grid tie line. The purpose of the power exchange constraint with the main grid tie line is to avoid the negative impact of large fluctuations in the power of the regional distribution network on the upstream transmission network and lines. The mathematical expression for the thermal energy balance constraint of the residential building is: in, This indicates the cooling capacity provided by the heat pump to the residents. This indicates the amount of heat provided to residents by the electric heating equipment. This indicates the residents' cooling load demand. This indicates the residents' heat load demand.

[0010] Preferably, the operational constraints of the controllable equipment include at least the operational constraints on adjustable photovoltaic systems, batteries, static var compensators, thermal storage tanks, and heat pumps and electric heating equipment. The mathematical expression for the operational constraints of the adjustable photovoltaic system is: in, express The active power of photovoltaics can be adjusted during different time periods. This indicates the maximum power generation capacity of the photovoltaic system. This indicates the maximum apparent power of the adjustable photovoltaic system; The mathematical expression for the operating constraints of the battery is: in, This indicates the battery's discharge power. This indicates the charging power of the battery. Indicates the state of charge of the battery. This indicates the maximum charging power of the battery. This indicates the maximum discharge power of the battery. This represents the minimum state of charge of the battery. This represents the battery's maximum state of charge. The mathematical expression for the operating constraints of the static var compensator is: in, for Time-of-use static var compensator The upper limit of the adjustable power range, for Time-of-use static var compensator The lower limit of the adjustable power range. for Time-of-use static var compensator Optimize the adjustment power; The mathematical expression for the operating constraints of the thermal storage tank is: in, This indicates the heat and cold storage status of the thermal storage tank. This indicates the output cooling and heating capacity of the thermal storage tank. This indicates the cooling and heating capacity input to the thermal storage tank. This refers to the upper limit of the heat storage tank's capacity to store heat or cold. This represents the lower limit of the heat storage capacity of the thermal storage tank. for The upper limit of heat or cold capacity that a time-limited thermal storage tank can store. for The upper limit of heat or cold energy released by the thermal storage tank during a given period; The mathematical expression for the operating constraints of the heat pump and electric heating equipment is as follows: in, This indicates the power of the heat pump. Indicates the power of the heating equipment. This indicates the maximum operating power of the heat pump. This indicates the maximum operating power of the electric heating equipment.

[0011] Preferably, both the Actor network and the Critic network adopt a heterogeneous Transformer network architecture, including: The Actor network is an Encoder-only Transformer-based Actor network that uses the historical state sequence as input to the Transformer encoder. The mathematical expression for the self-attention mechanism in Transformer is: in, For query vector, For matching vectors, For the target vector, The dimension of the sequence; the first dimension in an Actor network. The mathematical expression for the output of the layer Transformer encoder is: in, Indicates that the input sequence is mapped to Trainable parameters, This indicates the number of heads in the multi-head attention mechanism. Indicates the number of layers in the Transformer encoder; The Critic network is an Encoder-Decoder Transformer-based network that uses the global state sequence as the input to the encoder. ,in The number of agents is specified; the node correlation coefficient matrix is ​​obtained using the node adjacency matrix and the voltage and power data of each node as location information to encode the input state; a multi-head attention mechanism is used to extract the inter-state correlation features; the output memory sequence is then executed. As a decoder Value and Value input; global action sequence as decoder input. By combining geographic information and utilizing a multi-head attention mechanism to extract action-related features, the output is used as... Input value; comprehensively consider the relationship between the states and actions of each agent to generate a global evaluation value. .

[0012] To achieve the above objectives, this application also provides a voltage regulation device for large-scale residential distributed resource coordination distribution networks, which applies the voltage regulation method for large-scale residential distributed resource coordination distribution networks as described above, including: A voltage regulation framework construction module is used to pre-build a distribution network voltage regulation framework that includes residential distributed resources, multiple agents, and the distribution network environment; wherein, the residential distributed resources take residential buildings as the smallest control unit; The voltage regulation network construction module is used to perform distributed resource-coordinated voltage regulation of the distribution network based on a decentralized Actor network and a centralized Critic network within the voltage regulation framework of the distribution network; wherein, both the Actor network and the Critic network adopt a heterogeneous Transformer network architecture; The distributed resource setting module is used to set standardized action spaces and confidence levels for heterogeneous distributed resources of residents; The action value output module is used to combine the local state information and historical data of each agent to obtain a time series, and use the time series as the input of the corresponding Actor network to output the corresponding action value. The evaluation value output module is used to take the action values ​​and global state sequence of the multi-agent as input to the corresponding Critic network, and output the corresponding evaluation value in combination with the geographic information of the agent. The coordinated voltage regulation module is used to continuously perform coordinated voltage regulation based on the action value and the evaluation value, combined with the reward function, by using a Critic network to centrally train the model, and then using an Actor network to perform the decentralized execution. The reward function includes at least a voltage deviation term, which is used to characterize the voltage deviation of the distribution network.

[0013] To achieve the above objectives, this application also provides a large-scale residential distributed resource coordination distribution network voltage regulation computer device, including at least one processor, at least one memory and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which calls the program instructions to execute the large-scale residential distributed resource coordination distribution network voltage regulation method as described above.

[0014] To achieve the above objectives, this application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the large-scale residential distributed resource coordination distribution network voltage regulation method described above.

[0015] Beneficial effects: The collaborative voltage regulation method, device, computer equipment and storage medium of this application adopt a novel heterogeneous approach, using Transformer to replace Actor network and Critic network, which is suitable for large-scale intelligent agent control environment, and solves the problems of insufficient utilization of distributed resources in residential areas, complex model voltage regulation method and low training efficiency of multi-agent. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating a large-scale residential distributed resource coordination distribution network voltage regulation method provided in this application embodiment; Figure 2 This is a structural block diagram of the power distribution network voltage regulation framework provided in the embodiments of this application; Figure 3 A schematic diagram of the improved power distribution system of the IEEE-33 node provided in the embodiments of this application; Figure 4 The global reward fluctuation curve of the ST-Het-TMATD3 agent provided in the embodiments of this application; Figure 5 The SVC output of node 32 provided in this application embodiment is shown in the continuous real-time scheduling results over 240 hours; Figure 6 The energy absorption from the power grid by node 32 in the continuous real-time scheduling results over 240 hours, as provided in this embodiment of the application. Figure 7 The voltage curve of node 13 provided in this application embodiment is shown in the continuous real-time scheduling results over 240 hours. Figure 8 The voltage curve of node 14 provided in this application embodiment is shown in the continuous real-time scheduling results over 240 hours. Figure 9 The voltage curve of node 32 provided in this application embodiment is shown in the continuous real-time scheduling results over 240 hours. Figure 10A comparison of test results for different solutions provided in the embodiments of this application; Figure 11 This application provides a comparative scheme for ablation experiments. Figure 12 Cumulative reward curves for different algorithms provided in the embodiments of this application; Figure 13 Test results for different algorithms provided in the embodiments of this application; Figure 14 A detailed flowchart of the distribution network voltage regulation method for large-scale residential distributed resource coordination provided in the embodiments of this application; Figure 15 A structural block diagram of a large-scale residential distributed resource coordination distribution network voltage regulation device provided in the embodiments of this application.

[0018] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0020] In this document, the term "comprising" is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0021] In recent years, large-scale distributed photovoltaic (PV) grid connection has brought a series of challenges to the distribution network, including voltage amplitude fluctuations and voltage quality degradation. To address these issues, traditional distribution networks typically employ voltage regulators. However, traditional voltage regulators suffer from long response times, poor continuous voltage regulation capabilities, and the potential for frequent operation to reduce equipment reliability and increase maintenance costs. Compared to traditional voltage regulators, residential distributed resources are characterized by their large scale, fast response speed, and flexible control. Therefore, residential distributed resources are gradually becoming an important resource for enhancing the flexible regulation capabilities of the power system.

[0022] Currently, research on distributed resource participation in distribution network regulation mainly focuses on microgrid-level regulation. For example, multiple microgrids are combined into microgrid clusters to achieve complementary and interconnected distributed resources, thereby optimizing the distribution network. Another example is considering multiple microgrids as units to regulate continuously and discretely operating devices, thus optimizing the active and reactive power distribution of the distribution network. Yet another example is a novel mixed-integer linear programming model that utilizes microgrids to optimize voltage recovery in the distribution network. However, these existing technologies all use microgrids as the smallest unit for distributed resource regulation, failing to regulate distributed resources at the resident level. Furthermore, centralized control at the microgrid level requires communication, which not only increases decision-making time but may also raise user privacy concerns.

[0023] Traditional voltage control methods for distribution networks mainly consist of centralized control and distributed control based on physical models. Centralized optimization methods collect network-wide data from a dispatch center for optimization and dispatch calculations, then issue control commands to each controlled unit. For example, centralized voltage control technologies such as model predictive control and mixed-integer second-order cone programming are used to optimize active and reactive power in the distribution network. However, centralized optimization methods require users to fully share internal information. Distributed resources belong to different stakeholders, and the operating data of equipment within each area and the loads of sensitive users are private data. Centralized control systems cannot collect detailed information about the controlled objects or directly control each area. Distributed optimization methods, on the other hand, are based on the idea of ​​"decomposition and coordination." They decompose complex global optimization problems with multiple variables and constraints into multiple sub-problems of lower complexity, which are solved independently by each entity. Privacy is protected through multiple iterations and the exchange of necessary algorithmic information. For example, multi-block alternating direction multiplier methods and improved alternating direction multiplier methods are used to solve the problem of massive communication in centralized dispatch. However, when dealing with large-scale distribution networks and a large number of distributed resources, the modeling difficulty and scale of physical model-based distribution network optimization methods increase exponentially, and the computational speed cannot meet the needs of real-time decision-making.

[0024] To address the optimization of large-scale distributed resources, scholars both domestically and internationally have begun to focus on data-driven deep reinforcement learning (DRL) methods. These methods train agent neural network models offline in a simulation environment to progressively fit the optimal scheduling strategy, avoiding the difficulties in iterative optimization or ill-conditioned solutions that may exist in model-driven methods, thus significantly saving decision generation time. However, single-agent algorithms suffer from poor stability and low learning efficiency, making them unsuitable for multi-agent control tasks. Multi-Agent Deep Reinforcement Learning (MADRL) applies the ideas and algorithms of DRL to the learning and control of multi-agent systems. For example, an improved multi-agent Actor-Critic algorithm is used to solve the coordination optimization problem of an electro-aero-thermal coupled system, improving decision generation speed and cost. Another example is the use of the multi-agent Actor-Double-Critic algorithm to solve the charging control problem of multiple electric vehicles in a dynamically operating electric vehicle charging station. However, most MADRL algorithms use traditional Multi-Layer Perceptron (MLP) networks for their Actor and Critic networks. When dealing with dependencies between large-scale data, multiple hidden layers are needed to pass information, leading to an increase in the number of network layers and an exponential increase in computational cost with sequence length. Therefore, the currently popular MADRL algorithms suffer from low learning efficiency, slow convergence speed, and unstable training results in complex environments with multiple control entities.

[0025] In summary, current voltage regulation strategies for distribution networks both domestically and internationally have the following problems: The regulation of distributed resources has not been fully implemented at the residential level, resulting in residential distributed resources not being able to fully participate in the voltage optimization of the distribution network. Traditional voltage control methods rely on physical models, but with the increase of distributed resources, the difficulty and scale of modeling increase exponentially, and the computing speed cannot meet the needs of real-time decision-making. The currently popular MADRL algorithm suffers from low learning efficiency and unstable training results when applied to large-scale multi-agent control environments where distributed resident resources are the scheduling objects.

[0026] Reference Figure 1 and Figure 2 , Figure 1 This is a flowchart of the large-scale residential distributed resource coordination distribution network voltage regulation method in this embodiment. Figure 2 This is a structural block diagram of the distribution network voltage regulation framework in this embodiment.

[0027] like Figure 1As shown, this embodiment discloses a voltage regulation method for large-scale residential distributed resource coordination in distribution networks, applicable to, for example... Figure 2 The pre-constructed distribution network voltage regulation framework shown includes residential distributed resources, multi-agent systems, and the distribution network environment, in such cases... Figure 2 In the distribution network voltage regulation framework shown, we use multiple agents to control distributed resources of residents, with residential buildings as the smallest unit, thereby coordinating voltage regulation with the distribution network environment. The methods include: S1: A pre-constructed distribution network voltage regulation framework including residential distributed resources, multiple agents, and the distribution network environment; wherein, residential distributed resources take residential buildings as the smallest control unit; S2: Distributed resource-coordinated distribution network voltage regulation is carried out within the distribution network voltage regulation framework based on decentralized Actor networks and centralized Critic networks; wherein, both Actor networks and Critic networks adopt heterogeneous Transformer network architecture; S3: Set up standardized action space and confidence level for heterogeneous distributed residential resources; S4: Combine the local state information and historical data of each agent to obtain a time series, use the time series as the input of the corresponding Actor network, and output the corresponding action value; S5: The action values ​​and global state sequence of the multi-agent network are used as inputs to the corresponding Critic network. Combined with the geographic information of the agents, the corresponding evaluation value is output. S6: Based on action values ​​and evaluation values, and combined with the reward function, the model is trained centrally using the Critic network, and then the Actor network is used to perform distributed execution to continuously regulate voltage. The reward function includes at least a voltage deviation term, which is used to characterize the voltage deviation of the distribution network.

[0028] In this embodiment, the training efficiency of agents in heterogeneous devices is improved by standardizing the action space and confidence level; the local state information and historical data of each agent are combined to obtain a time series, which is used as the input of the Actor network to output action values; the action values ​​of multiple agents and the global state sequence are used as the input of the Critic network, which, combined with the geographical information of the agents, outputs evaluation values; based on the action values ​​and evaluation values, and combined with the reward function, the Critic network is used to centrally train the model, and then the Actor network is used to perform distributed execution to continuously regulate the voltage of the distribution network.

[0029] Based on the above, this embodiment proposes a residential distributed resource collaborative distribution network voltage regulation strategy based on ST-Het-TMATD3 (Spatio-Temporal Heterogeneous Transformer-based Multi-Agent Twin Delayed3) to address the problems existing in current domestic and international distribution network voltage regulation strategies. Specifically, this paper uses the Transformer, which has achieved great success in natural language processing in recent years, to replace the traditional Actor network and Critic network in the MATD3 (Multi-Agent Twin Delayed3) algorithm. Compared with existing models such as recurrent neural networks and long short-term memory networks, the computational complexity of the Transformer is linear when processing long sequences, rather than increasing exponentially with the sequence length. The model-free ST-Het-TMATD3 algorithm based on the CTDE (Centralized Training and Decentralized Execution) architecture is used to regulate the distribution network, with residential buildings as the smallest unit of distributed resource regulation. In this architecture, the decentralized Actor network only needs local information for decision-making, which is beneficial for protecting resident privacy and improving real-time decision-making speed. The centralized Critic network requires global information for training. To adapt to multi-agent, multi-dimensional, and continuous control and regulation tasks, the Actor and Critic networks are replaced by novel heterogeneous Transformers. The local observation states of the agents have time-series characteristics; an Encoder-only Transformer is used instead of the Actor network, combining historical states to form a time-series input, and multi-head attention is used to extract features. The Critic network uses an Encoder-DecoderTransformer to process global information, combining geographic location information and multi-head attention to extract state-location association features. The agent's action sequence serves as the decoder input; multi-head attention extracts action-location association features, and a cross-multi-head attention mechanism integrates state and action features, avoiding learning difficulties caused by simply concatenating inputs. This method fully considers the interrelationships between resident states and the decisions of various agents, thereby improving the accuracy of Actor network decisions and Critic network evaluations.

[0030] Specifically, the reward function also includes a residential electricity purchase cost item, which represents the cost of electricity purchase for residents.

[0031] Specifically, the reward function is used to solve the preset objective function, which aims to minimize the combined target of each resident's electricity purchase cost and the voltage deviation of the distribution network.

[0032] In the specific application of this embodiment, we use the following formula to describe the objective function of voltage regulation in residential distributed resource cooperative distribution networks: in, Indicates residents The cost of purchasing electricity, This represents the electricity purchase cost coefficient. This indicates the voltage deviation in the distribution network. This represents the voltage deviation cost coefficient.

[0033] Based on the objective function, we describe the residents using the following formula. Electricity purchase cost : in, Indicates in Residents during the period Power exchanged with the power grid Indicates the purchase of electricity. Indicates electricity sales. For scheduling time intervals, This indicates the price residents pay when purchasing electricity from the grid. This indicates the price at which residents sell electricity to the power grid.

[0034] In response to Residents during the period Power exchanged with the power grid We describe it using the following formula: in, This indicates the electricity load demand of residents.

[0035] Based on the objective function, we use the following formula to describe the voltage deviation in the distribution network. : in, Indicates the number of nodes. To control the cycle, Indicates in Residents at the node during the time period voltage amplitude, This indicates the voltage reference value.

[0036] Specifically, the objective function is subject to preset constraints, including power balance constraints of the distribution network, safe operation constraints of the distribution network, operation constraints of controllable equipment, and heat and cold energy balance constraints of residential buildings.

[0037] In the specific application of this embodiment, the power balance constraint of the distribution network is described by the following formula: in, For nodes The active power injected by the photovoltaic system For nodes The reactive power injected by the photovoltaic system For nodes Active power of load For nodes Reactive power of load For nodes With nodes The branch conductance between For nodes With nodes The susceptance between them For nodes With nodes voltage between and voltage The phase angle difference.

[0038] In the specific application of this embodiment, the safety operation constraints of the distribution network are described by the following formula: in, for Time period nodes Voltage amplitude, For nodes The upper limit of voltage amplitude, For nodes The lower limit of voltage amplitude for Time-of-day branch The current value, branch road The upper limit of the current amplitude, This represents the maximum exchange power between the regional distribution network and the main grid interconnection line. This is the upper limit of the maximum exchange power between the regional distribution network and the main grid interconnection line. This is the lower limit of the maximum power exchanged between the regional distribution network and the main grid tie line. The purpose of the power exchange constraint with the main grid tie line is to avoid the negative impact of large fluctuations in the power of the regional distribution network on the upstream transmission network and lines.

[0039] In the specific application of this embodiment, the thermal energy balance constraint of residential buildings is described by the following formula: in, This indicates the cooling capacity provided by the heat pump to the residents. This indicates the amount of heat provided to residents by the electric heating equipment. This indicates the residents' cooling load demand. This indicates the residents' heat load demand.

[0040] Specifically, the operational constraints on controllable equipment include at least the operational constraints on adjustable photovoltaic systems, batteries, static var compensators, thermal storage tanks, and heat pumps and electric heating equipment.

[0041] In the specific application of this embodiment, the operating constraints of adjustable photovoltaics are described by the following formula: in, express The active power of photovoltaics can be adjusted during different time periods. This indicates the maximum power generation capacity of the photovoltaic system. This indicates the maximum apparent power of the adjustable photovoltaic system.

[0042] In the specific application of this embodiment, the operating constraints of the battery are described by the following formula: in, This indicates the battery's discharge power. This indicates the charging power of the battery. Indicates the state of charge of the battery. This indicates the maximum charging power of the battery. This indicates the maximum discharge power of the battery. This represents the minimum state of charge of the battery. This represents the battery's maximum state of charge.

[0043] In the specific application of this embodiment, the operating constraints of the static var compensator are described by the following formula: in, for Time-of-use static var compensator The upper limit of the adjustable power range, for Time-of-use static var compensator The lower limit of the adjustable power range. for Time-of-use static var compensator Optimized power adjustment.

[0044] In the specific application of this embodiment, the operating constraints of the thermal storage tank are described by the following formula: in, This indicates the heat and cold storage status of the thermal storage tank. This indicates the output cooling and heating capacity of the thermal storage tank. This indicates the cooling and heating capacity input to the thermal storage tank. This refers to the upper limit of the heat storage tank's capacity to store heat or cold. This represents the lower limit of the heat storage capacity of the thermal storage tank. for The upper limit of heat or cold capacity that a time-limited thermal storage tank can store. for The upper limit of heat or cold released by the thermal storage tank during a given period.

[0045] In the specific application of this embodiment, the operating constraints of the heat pump and the electric heating equipment are described by the following formula: in, This indicates the power of the heat pump. Indicates the power of the heating equipment. This indicates the maximum operating power of the heat pump. This indicates the maximum operating power of the electric heating equipment.

[0046] Specifically, the residential load in the distributed resources of residents corresponds to heterogeneous loads, which include at least adjustable photovoltaics, batteries and thermal storage tanks.

[0047] Specifically, the Actor network and Critic network are updated using a double-delay deterministic policy gradient update.

[0048] In the residential distributed resource cooperative distribution network voltage regulation strategy based on ST-Het-TMATD3 proposed in this embodiment, the voltage regulation problem of residential distributed resource cooperative distribution network is modeled as a reinforcement learning (RL) model based on MDP (Markov Decision Process), resulting in a MADRL-based voltage regulation framework for residential distributed resource cooperative distribution network. In this framework, each household is regarded as an agent responsible for data collection, equipment status information, and voltage status monitoring. The centralized Critic network in the MATD3 algorithm processes all agent information during training, solving the non-stationarity problem in multi-agent learning. After training, the agents use a decentralized Actor network to make decisions, shortening the decision time and meeting real-time requirements. At the same time, the decentralized nature of the CTDE (Centralized Training and Decentralized Execution) architecture allows agents to make independent decisions based on local information, protecting privacy data.

[0049] The ST-Het-TMATD3 algorithm proposed in this embodiment will now be described in detail.

[0050] The Dual-Delay Deep Deterministic Policy Gradient (TD3) algorithm, based on the Actor-Critic framework, is a popular DRL algorithm. MATD3, which extends TD3 to multi-agent scenarios, solves the problem of overestimation of the value function in the traditional MADDPG (Multi-Agent Deep Deterministic Policy Gradient). This embodiment integrates the Transformer network into reinforcement learning, leveraging its excellent sequence processing capabilities. Local observation states and historical data form the input to the Actor network, while the global state sequence, action sequence, and location information constitute the input to the Critic network. This approach fully considers spatiotemporal information and is suitable for voltage regulation scenarios in residential distributed resource collaborative distribution networks.

[0051] In the ST-Het-TMATD3 algorithm proposed in this embodiment, heterogeneous residential distributed energy systems have incompatible operational dimensions and differentiated dynamic response characteristics, making joint optimization under a unified control strategy extremely complex. To address this challenge, we design a collaborative control scheme for heterogeneous devices based on a confidence-gated strategy. First, a standardized action space is introduced. To characterize the general control signal, parameter sharing is used to improve the training efficiency of multi-agent systems. Furthermore, to enhance system stability and prevent oscillating control behavior that may occur in the continuous action space, a confidence-gated output mechanism is introduced. The final executed action is defined as a weighted interpolation between historical actions and the current network output, with the specific mathematical expression as follows: Among them, the confidence score (denoted as Constrained within the interval Within the system, computation is performed using a dedicated MLP based on input observations. This mechanism enables each agent to dynamically interpolate between new actions proposed by the policy network and their previous actions, thereby smoothing the output action sequence. Confidence scoring is constrained by applying device-specific thresholds to ensure system stability. Safe to operate and compliant Physical constraints. This method effectively achieves scalable and stable control over heterogeneous distributed residential resources.

[0052] In the ST-Het-TMATD3 algorithm proposed in this embodiment, the Actor network is an Encoder-only Transformer-based Actor network. Compared to the sequential processing method of RNN (Recurrent Neural Network), the Transformer can process all positions of the entire sequence simultaneously, without relying on previous time steps like an RNN. Traditional Actor networks only receive the observed state at the current moment as input; however, the observed states of each agent exhibit time-series characteristics during learning, so using the historical state sequence as input to the Transformer encoder is the approach used in this embodiment. .

[0053] We describe the self-attention mechanism in Transformer using the following formula: in, For query vector, For matching vectors, For the target vector, Indicates the dimension of the sequence.

[0054] We will use the first Actor in the Actor network in this embodiment. The output of the layer Transformer encoder is written as: in, Indicates that the input sequence is mapped to Trainable parameters, This indicates the number of heads in the multi-head attention mechanism. This indicates the number of layers in the Transformer encoder.

[0055] In the ST-Het-TMATD3 algorithm proposed in this embodiment, the Critic network is an Encoder-DecoderTransformer-based Critic network. Traditional centralized Critic networks using MLP networks suffer from computational burden in large-scale cooperative environments. Simply concatenating the global action sequence and state sequence as input can easily lead to the agent's inability to effectively learn key features. This embodiment employs an Encoder-Decoder Transformer-based Critic network, using the global state sequence as the encoder input. ,in The number of agents is specified. The node correlation coefficient matrix is ​​obtained using the node adjacency matrix and the voltage and power data of each node as location information to encode the input state. A multi-head attention mechanism is used to extract the inter-state correlation features. The output is a memory sequence. As a decoder Value and Value input. The global action sequence is used as the decoder input in this embodiment. By combining geographic information and utilizing a multi-head attention mechanism to extract action-related features, the output is used as... Input value. A global evaluation value is generated by comprehensively considering the relationships between the states and actions of each agent. .

[0056] In the residential distributed resource collaborative distribution network voltage regulation strategy based on ST-Het-TMATD3 proposed in this embodiment, network parameter updates are performed within the CTDE architecture. The length is... state sequence As input to the Actor network, the output action is obtained. .in, Represents the global state sequence of a centralized Critic network. According to the strategy Output global action sequence To address the value estimation problem, the TD3 algorithm employs a dual-objective Critic network and objective policy smoothing regularization. The dual-objective Critic network uses two identical target Critic networks, selecting the minimum value between them as the target value, thus suppressing the overestimation problem of the value network. and These represent the network parameters of the Critic network and two different objectives, respectively. Objective policy smoothing regularization calculates the objective using the objective actions. This value facilitates a smooth estimation of the target, hence in the next state sequence The action is , This represents random noise. To make the target action more closely resemble the original action, it is generally normally distributed, and the sampled noise is truncated. , For variance, These represent the upper and lower limits of the noise. To reduce the error caused by updating the Actor network before the Critic network has stabilized, the Actor network parameters are updated only after the Critic network has converged. The parameters of the target Actor network and the target Critic network are updated using a soft update method.

[0057] in, For soft update rate, For two different current Critic network parameters, These are the current Actor network parameters. The parameters of the target Actor network.

[0058] In the specific application of this embodiment, in the distribution network voltage regulation optimization model, MADRL uses a multi-agent Markov decision process as its basic framework and can be represented as a tuple. .in, For the joint state of the agents, Represents intelligent agents exist The state of the time period; For joint actions of intelligent agents, Represents intelligent agents exist Actions selected during specific time periods; The state transition probability of the agent; For joint rewards for intelligent agents, Represents intelligent agents exist Rewards earned during the specified time period; The MADRL algorithm aims to find the optimal strategy that maximizes the cumulative discount reward, using a discount factor as the reward.

[0059] Based on the above tuples, we further describe the state space. The state space is the environmental information perceived by the agent. Time period The state of an agent is represented by the following formula: in, Indicates residents exist The energy storage status of the thermal storage tank for hot water during a given time period. Indicates residents exist The energy storage status of the thermal storage tank for cold water during a given period. Indicates the battery status. Indicates photovoltaic power generation capacity. express The relative voltage of the nodes.

[0060] Based on the above tuples, we further describe the action space. The action space consists of relevant decision variables, which we will... Time period The actions of an agent are represented by the following formula: in, This indicates the amount of cooling capacity stored and released in the thermal storage tank. This indicates the amount of heat released and stored in the thermal storage tank. Indicates the battery's charging and discharging power. This represents the reactive power output of the photovoltaic system.

[0061] Based on the above tuples, we clarify that the training objective of MADRL is to find the optimal policy. To maximize the expected cumulative return, and described by the following formula: in, Representation Strategy The trajectory formed This represents the number of steps per round. This embodiment transforms the problem of minimizing the total system cost into maximizing the cumulative reward in deep reinforcement learning, thus setting an immediate reward. This includes the cost of electricity purchased by residents and the voltage deviation item.

[0062] Regarding the cost of residential electricity purchases, taking into account residents' willingness to participate in regulation, the cost of residential electricity purchases after the agent makes scheduling decisions is used to measure the willingness to regulate and is set as follows: , Described using the following formula: For the voltage deviation term, the system power flow distribution after each agent makes a scheduling decision is calculated, and based on the obtained voltage of each node, it is set as follows: , Described using the following formula: In this embodiment, the environment provides immediate rewards to each agent at each time interval. Described using the following formula: Based on the above, this embodiment discloses a complete collaborative voltage regulation method based on ST-Het-TMATD3, and achieves at least the following technical effects: The ST-Het-TMATD3 model-free method solves the problem of exponential growth in modeling volume in large-scale residential distributed resource regulation; Based on the CTDE architecture, Transformer is introduced into MADRL, and a new heterogeneous approach is adopted to replace the Actor and Critic networks, which can adapt to the distributed resource regulation tasks of residents with multiple subjects, multiple dimensions, and continuous control. Distributed Actor networks use residential buildings as the smallest decision-making unit, protecting privacy and ensuring real-time decision-making speed. They combine local state information with historical data as time-series input to improve decision-making accuracy. The centralized Critic network uses multi-agent actions and global state sequences as inputs to the Transformer encoder and decoder, fully considering the influence of resident states and agent decisions, solving the problem of difficulty in learning key features caused by simple concatenation of inputs, and improving evaluation accuracy and agent training efficiency.

[0063] Reference Figure 3 , Figure 3 This is a schematic diagram of the improved power distribution system of the IEEE-33 node in this embodiment.

[0064] To verify the above technical effects, this embodiment was tested on the improved IEEE-33 node power distribution system, using measured data from four types of residential buildings in a southeastern region in 2023, totaling 30 days of data. The first 20 days were used for reinforcement learning training, and the remaining 10 days were used for testing. The training set was used for centralized learning, and the test set was used for distributed evaluation. Each node, excluding grid access node 1, has 5 residential buildings.

[0065] Reference Figure 4 , Figure 4 This is the global reward fluctuation curve of the ST-Het-TMATD3 agent in this embodiment.

[0066] For the ST-Het-TMATD3 training process, the power flow calculation in this embodiment is performed on Pandapower, and all algorithms in this embodiment are implemented on the PyTorch neural network framework. The computing platform hardware information is: CPU 13th Gen Intel(R) Core(TM) i7-13700K, GPU NVIDIA GeForce RTX 4070. Through training, the global reward fluctuation curve of the ST-Het-TMATD3 agent is as follows... Figure 4 As shown, the Epoch on the horizontal axis represents the process by which the model has fully learned the entire training dataset. The agent's cumulative reward gradually increases and then stabilizes at a relatively high reward value, indicating that training is complete.

[0067] To analyze the voltage optimization results of the distribution network and verify the effectiveness of the proposed voltage regulation scheme in voltage optimization, the default electricity consumption strategy (setting the thermal storage tank to timed storage and not using batteries and photovoltaics) without interfering with residents was used as the baseline control method (Rule-Based Control, RBC). Ten days of test data were used, and the following four voltage regulation modes (schemes) were set: Mode 1: Residents use RBC control and do not use SVC (Static Var Compensator) for voltage regulation.

[0068] Mode 2: After all residents use RBC, SVC is used for voltage regulation.

[0069] Mode 3: Residents do not use SVC voltage regulation after using ST-Het-TMATD3 control.

[0070] Mode 4: Residents use ST-Het-TMATD3 control and then use SVC voltage regulation.

[0071] Reference Figure 5 , Figure 5 This shows the SVC output of node 32 in this embodiment during 240 hours of continuous real-time scheduling.

[0072] Taking node 32 in the improved IEEE-33 node power distribution system as an example, the SVC output in Mode 2 and Mode 4 is as follows: Figure 5 As shown, after RBC control, the SVC output fluctuates significantly and is relatively high. However, ST-Het-TMATD3 control reduces the SVC output by regulating the distributed resources of residents, thereby reducing the installed capacity of the SVC and reducing the output fluctuation of the SVC, which can extend the service life of the SVC.

[0073] Reference Figure 6 , Figure 6 This describes the situation of node 32 absorbing electrical energy from the power grid in the continuous real-time scheduling results over 240 hours in this embodiment.

[0074] The situation of nodes absorbing electrical energy from the grid under different control schemes is as follows: Figure 6 As shown. Mode 1 uses RBC control and does not utilize batteries for energy storage, therefore it sells excess electricity to the grid when photovoltaic output is high. Mode 3, on the other hand, uses batteries to store excess electricity and releases it during peak load periods, achieving peak shaving and valley filling. Mode 4 uses SVC control, eliminating concerns about node voltage deviation caused by purchasing electricity during off-peak periods. Therefore, ST-Het-TMATD3 can minimize electricity purchase costs, purchasing more electricity during low load periods and releasing battery power during high load periods, resulting in smaller fluctuations and better performance compared to Mode 3.

[0075] To verify the optimization effect of the ST-Het-TMATD3 algorithm on voltage distribution, this embodiment selects node 13, CB2 compensation node 14, and SVC compensation node 32, which are far from the reactive power compensation equipment, and compares and analyzes their voltage distribution. Figures 7 to 9 As shown.

[0076] Reference Figure 7 , Figure 7 This is the voltage curve of node 13 in this embodiment during 240 hours of continuous real-time scheduling.

[0077] like Figure 7 As shown, the overall voltage of node 13 is low. Mode 2 demonstrates the voltage boosting effect of SVC. Mode 3 uses ST-Het-TMATD3 to increase the overall voltage. Mode 4 uses both SVC and ST-Het-TMATD3 methods to increase the overall voltage and reduce fluctuations, thus maintaining voltage stability.

[0078] Reference Figure 8 , Figure 8 This is the voltage curve of node 14 in this embodiment during 240 hours of continuous real-time scheduling.

[0079] like Figure 8 As shown, the overall voltage of node 14 is too high, the voltage of Mode 2 is further increased due to the influence of SVC, while Mode 3 and Mode 4 adjust the voltage through distributed resources to reduce the deviation.

[0080] Reference Figure 9 , Figure 9 This is the voltage curve of node 32 in this embodiment during 240 hours of continuous real-time scheduling.

[0081] like Figure 9 As shown, Mode3 has a similar voltage deviation to Mode1 compared to Mode4, but Mode4 has smaller voltage fluctuations and is more stable.

[0082] In summary, Mode 1 showed the largest deviation, Mode 2 and Mode 3 showed some improvement, and Mode 4 performed the best with the smallest voltage fluctuation, demonstrating a significant optimization effect.

[0083] Reference Figure 10 , Figure 10 This is a comparison of the test results of different schemes in this embodiment.

[0084] To quantitatively evaluate the effectiveness of distribution network voltage optimization, this embodiment analyzes the results of 960 power flow calculations conducted over a 10-day test period using different schemes. Voltage data from 30,720 nodes (excluding grid connection points) were collected. The frequency of voltage deviations exceeding ±0.5 was defined as the overvoltage frequency, and test results for different control methods were obtained, such as... Figure 10 As shown. From Figure 10 The data shows that compared to Scheme 1, Schemes 2, 3, and 4 reduced the overvoltage frequency by 2.05%, 1.21%, and 4.53%, respectively. Schemes 1 and 2 use RBC control, keeping the residential electricity purchase cost unchanged; while Schemes 3 and 4 reduce costs by 7.43 yuan and 15.5 yuan, respectively, showing that the control in this embodiment can reduce electricity purchase costs. However, Scheme 3 does not use SVC control, resulting in a larger overall voltage deviation, only slightly reducing electricity purchase costs to avoid voltage problems. In contrast, Scheme 4 has a smaller overall voltage deviation, and the algorithm in this embodiment can significantly reduce electricity purchase costs.

[0085] Based on the above analysis, we conclude that: in the absence of SVC, ST-Het-TMATD3 regulation can replace SVC control and reduce overvoltage situations; under SVC voltage regulation, ST-Het-TMATD3 regulation further optimizes voltage distribution and reduces residential electricity purchase costs, while also reducing SVC installation capacity and output fluctuations.

[0086] Reference Figure 11 , Figure 11 This is the ablation experiment comparison scheme for this embodiment.

[0087] Reference Figure 12 , Figure 12 This is the cumulative reward curve for different algorithms in this embodiment.

[0088] Reference Figure 13 , Figure 13 These are the test results for different algorithms in this embodiment.

[0089] To verify the rationality of the improvements to each part of the proposed ST-Het-TMATD3 algorithm, this paper conducts ablation experiments by replacing each improved part. The settings of each algorithm model are as follows: Figure 11 The training results are as follows Figure 12 As shown. After the model training was completed, MADRL control and SVC voltage regulation were used as the control scheme. The test results of different algorithms in the ablation experiment were compared, as shown. Figure 13 As shown.

[0090] Combination Figure 12 and Figure 13 The results show that using Transformer instead of traditional MLP networks improves the training speed and cumulative reward of multi-agent networks, and yields better test results. When the Critic network inputs are concatenated with actions and states, the Transformer's Encoder-Decoder replacement slightly outperforms the Encoder-only replacement, especially in terms of training convergence speed, which shows no significant improvement. This is because simple concatenation does not fully utilize the encoder and decoder capabilities of the Transformer, making it difficult to quickly extract key features. Compared to TE-MATD3, ST-Het-TMATD3-cat shows limited improvement. However, under the same input method and Transformer replacement, ST-Het-TMATD3, while showing slight improvement in test results, converges faster, indicating that extracting action and state features separately improves the feature extraction efficiency of the Transformer.

[0091] Based on the above verification, we draw the following conclusions: This embodiment proposes the ST-Het-TMATD3 algorithm for voltage regulation strategies in residential distributed resource collaborative distribution networks. This algorithm adopts a novel heterogeneous approach, using Transformer to replace the Actor and Critic networks in the CTDE architecture, making it suitable for large-scale agent control environments. It solves problems such as insufficient utilization of residential distributed resources, complex model-based voltage regulation methods, and low training efficiency of multi-agent systems. Four control schemes were compared in the improved IEEE-33 node system. The results show that without SVC control, ST-Het-TMATD3 can achieve similar voltage regulation effects to SVC; while with SVC control, ST-Het-TMATD3 can reduce SVC output fluctuations and installed capacity, better optimizing residential electricity purchase costs, thus verifying the effectiveness of the algorithm in distribution network voltage regulation and reducing residential electricity purchase costs. Ablation experiments verified the rationality and superiority of the improvements made to each part of the algorithm.

[0092] Reference Figure 14 , Figure 14 This is a detailed flowchart of the coordinated voltage regulation method in this embodiment.

[0093] In summary, as Figure 14 As shown, we will provide a detailed explanation of the collaborative voltage regulation method based on residential distributed resources and the ST-Het-TMATD3 algorithm in this embodiment: Establish a building-level voltage regulation model for heterogeneous residential distributed resource collaborative distribution networks; The established residential distributed resource collaborative distribution network voltage regulation model is designed as a Markov decision problem; Standardize the action space and confidence level for heterogeneous distributed residential resources; The ST-Het-TMATD3 algorithm is used to solve the Markov decision problem. Specifically, the local state information and historical data of each agent are combined to obtain a time series, which is used as the input of the Actor network to output action values. The action values ​​of multiple agents and the global state sequence are used as the input of the Critic network, which is combined with the geographical information of the agents to output evaluation values. Obtain a decentralized Actor network with an optimal resident distributed resource control strategy; Deploy a distributed Actor network with optimal resident distributed resource control strategy among various resident users; in particular: collect the state information of resident distributed resources; and output the control quantities of various heterogeneous resident distributed resources.

[0094] Reference Figure 15 , Figure 15 This is a structural block diagram of the voltage regulation of a large-scale residential distributed resource collaborative distribution network in this embodiment.

[0095] like Figure 15 As shown, this embodiment also discloses a voltage regulation device for large-scale residential distributed resource coordination distribution networks, which applies the voltage regulation method for large-scale residential distributed resource coordination distribution networks described above, including: A voltage regulation framework construction module is used to pre-build a distribution network voltage regulation framework that includes residential distributed resources, multiple agents, and the distribution network environment; wherein, the residential distributed resources take residential buildings as the smallest control unit; The voltage regulation network construction module is used to perform distributed resource-coordinated voltage regulation of the distribution network based on a decentralized Actor network and a centralized Critic network within the voltage regulation framework of the distribution network; wherein, both the Actor network and the Critic network adopt a heterogeneous Transformer network architecture; The distributed resource setting module is used to set standardized action spaces and confidence levels for heterogeneous distributed resources of residents; The action value output module is used to combine the local state information and historical data of each agent to obtain a time series, and use the time series as the input of the corresponding Actor network to output the corresponding action value. The evaluation value output module is used to take the action values ​​and global state sequence of the multi-agent as input to the corresponding Critic network, and output the corresponding evaluation value in combination with the geographic information of the agent. The coordinated voltage regulation module is used to continuously perform coordinated voltage regulation based on the action value and the evaluation value, combined with the reward function, by using a Critic network to centrally train the model, and then using an Actor network to perform the decentralized execution. The reward function includes at least a voltage deviation term, which is used to characterize the voltage deviation of the distribution network.

[0096] This embodiment also discloses a large-scale residential distributed resource coordination distribution network voltage regulation computer device, including at least one processor, at least one memory and a data bus; The processor and memory communicate with each other via a data bus; The memory stores program instructions that can be executed by the processor, which calls the program instructions to execute the voltage regulation method for large-scale residential distributed resource coordination in the distribution network as described above.

[0097] This embodiment also discloses a storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the voltage regulation method for large-scale residential distributed resource coordination distribution networks as described above.

[0098] It should be noted that the voltage regulation device, computer equipment, and storage medium for large-scale residential distributed resource collaborative distribution networks in this embodiment correspond to the aforementioned voltage regulation method for large-scale residential distributed resource collaborative distribution networks. Therefore, any content not specifically described in the voltage regulation device, computer equipment, and storage medium for large-scale residential distributed resource collaborative distribution networks in this embodiment, including but not limited to functional definitions, working principles, and technical effects, can be referred to the description in the aforementioned voltage regulation method for large-scale residential distributed resource collaborative distribution networks, and will not be repeated here.

[0099] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor may be implemented in one or more of the following: application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to implement the functions described herein, or combinations thereof. For software implementation, some or all of the processes of the embodiments may be performed by a computer program instructing the associated hardware. During implementation, the program may be stored in a computer-readable storage medium or transmitted as one or more instructions or code on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible to a computer. Computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having the form of instructions or data structures and accessible to a computer.

[0100] Finally, it should be noted that the above description is only a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for voltage regulation in a large-scale residential distributed resource collaborative distribution network, characterized in that, The method includes: A pre-constructed distribution network voltage regulation framework includes residential distributed resources, multiple agents, and the distribution network environment; wherein, the residential distributed resources are based on residential buildings as the smallest control unit; Within the aforementioned distribution network voltage regulation framework, distributed resource-coordinated distribution network voltage regulation is performed based on a decentralized Actor network and a centralized Critic network; wherein, both the Actor network and the Critic network adopt a heterogeneous Transformer network architecture; Standardize the action space and confidence level for heterogeneous distributed residential resources; The local state information and historical data of each agent are combined to obtain a time series. The time series is used as the input of the corresponding Actor network to output the corresponding action value. The action values ​​and global state sequences of the multi-agent network are used as inputs to the corresponding Critic network, and combined with the geographic information of the agents, the corresponding evaluation values ​​are output. Based on the action value and the evaluation value, after training the model centrally using the Critic network in conjunction with the reward function, the coordinated voltage regulation is continuously performed using the distributed execution method of the Actor network; wherein, the reward function includes at least a voltage deviation term, which is used to characterize the voltage deviation of the distribution network.

2. The method for voltage regulation of a large-scale residential distributed resource collaborative distribution network according to claim 1, characterized in that, The residential loads in the distributed residential resources correspond to heterogeneous loads, which include at least adjustable photovoltaics, batteries, and thermal storage tanks.

3. The method for voltage regulation of a large-scale residential distributed resource collaborative distribution network according to claim 1, characterized in that, The reward function also includes a residential electricity purchase cost term, which represents the cost of electricity purchased by residents; the mathematical expression of the reward function is: in, This refers to the immediate rewards that the environment provides to each agent at each time interval. The willingness to regulate is measured by the cost of electricity purchased by residents after the intelligent agent makes scheduling decisions. For the voltage of each node, and These are the preset weight parameters.

4. The method for voltage regulation of a large-scale residential distributed resource collaborative distribution network according to claim 3, characterized in that, The reward function is used to solve a preset objective function, which aims to minimize the combined cost of electricity purchase for each resident and the voltage deviation of the distribution network. The mathematical expression of the objective function is: in, Indicates residents The cost of purchasing electricity, This represents the electricity purchase cost coefficient. This indicates the voltage deviation in the distribution network. This represents the voltage deviation cost coefficient; and , .

5. The method for voltage regulation of a large-scale residential distributed resource collaborative distribution network according to claim 4, characterized in that, The objective function is subject to preset constraints, including power balance constraints of the distribution network, safe operation constraints of the distribution network, operation constraints of controllable equipment, and heat and cold energy balance constraints of the residential buildings. The mathematical expression for the power balance constraint of the distribution network is: in, For nodes The active power injected by the photovoltaic system For nodes The reactive power injected by the photovoltaic system For nodes Active power of load For nodes Reactive power of load, For nodes With nodes The branch conductance between For nodes With nodes The susceptance between them For nodes With nodes voltage between and voltage The phase angle difference; The mathematical expression for the safety operation constraints of the distribution network is: in, for Time period nodes Voltage amplitude, For nodes The upper limit of voltage amplitude, For nodes The lower limit of voltage amplitude for Time-of-day branch The current value, branch road The upper limit of the current amplitude, This represents the maximum exchange power between the regional distribution network and the main grid interconnection line. This is the upper limit of the maximum exchange power between the regional distribution network and the main grid interconnection line. This is the lower limit of the maximum power exchanged between the regional distribution network and the main grid tie line. The purpose of the power exchange constraint with the main grid tie line is to avoid the negative impact of large fluctuations in the power of the regional distribution network on the upstream transmission network and lines. The mathematical expression for the thermal energy balance constraint of the residential building is: in, This indicates the cooling capacity provided by the heat pump to the residents. This indicates the amount of heat provided to residents by the electric heating equipment. This indicates the residents' cooling load demand. This indicates the residents' heat load demand.

6. The method for voltage regulation of a large-scale residential distributed resource collaborative distribution network according to claim 5, characterized in that, The operational constraints of the controllable equipment include at least the operational constraints on adjustable photovoltaic systems, batteries, static var compensators, thermal storage tanks, and heat pumps and electric heating equipment. The mathematical expression for the operational constraints of the adjustable photovoltaic system is: in, express The active power of photovoltaics can be adjusted during different time periods. This indicates the maximum power generation capacity of the photovoltaic system. This indicates the maximum apparent power of the adjustable photovoltaic system; The mathematical expression for the operating constraints of the battery is: in, This indicates the battery's discharge power. This indicates the charging power of the battery. Indicates the state of charge of the battery. This indicates the maximum charging power of the battery. This indicates the maximum discharge power of the battery. This represents the minimum state of charge of the battery. This represents the battery's maximum state of charge. The mathematical expression for the operating constraints of the static var compensator is: in, for Time-of-use static var compensator The upper limit of the adjustable power range, for Time-of-use static var compensator The lower limit of the adjustable power range. for Time-of-use static var compensator Optimize power adjustment; The mathematical expression for the operating constraints of the thermal storage tank is: in, This indicates the heat and cold storage status of the thermal storage tank. This indicates the output cooling and heating capacity of the thermal storage tank. This indicates the cooling and heating capacity input to the thermal storage tank. This refers to the upper limit of the heat storage tank's capacity to store heat or cold. This represents the lower limit of the heat storage capacity of the thermal storage tank. for The upper limit of heat or cold capacity that a time-limited thermal storage tank can store. for The upper limit of heat or cold energy released by the thermal storage tank during a given period; The mathematical expression for the operating constraints of the heat pump and electric heating equipment is as follows: in, This indicates the power of the heat pump. Indicates the power of the heating equipment. This indicates the maximum operating power of the heat pump. This indicates the maximum operating power of the electric heating equipment.

7. The method for voltage regulation of a large-scale residential distributed resource collaborative distribution network according to claim 1, characterized in that, Both the Actor network and the Critic network adopt a heterogeneous Transformer network architecture, including: The Actor network is an Encoder-only Transformer-based Actor network that uses the historical state sequence as input to the Transformer encoder. The mathematical expression for the self-attention mechanism in Transformer is: in, For query vector, For matching vectors, For the target vector, The dimension of the sequence; the first dimension in an Actor network. The mathematical expression for the output of the layer Transformer encoder is: in, Indicates that the input sequence is mapped to Trainable parameters, This indicates the number of heads in the multi-head attention mechanism. Indicates the number of layers in the Transformer encoder; The Critic network is an Encoder-Decoder Transformer-based network that uses the global state sequence as the input to the encoder. ,in The number of agents is specified; the node correlation coefficient matrix is ​​obtained through the node adjacency matrix and the voltage and power data of each node as position information to encode the input state; a multi-head attention mechanism is used to extract the correlation features between states; the output memory sequence is then executed. as a decoder Value and Value input; global action sequence as decoder input. By combining geographic information and utilizing a multi-head attention mechanism to extract action-related features, the output is used as... Input value; comprehensively consider the relationship between the states and actions of each agent to generate a global evaluation value. .

8. A voltage regulation device for a large-scale residential distributed resource collaborative distribution network, employing the voltage regulation method for a large-scale residential distributed resource collaborative distribution network as described in any one of claims 1 to 7, characterized in that, include: A voltage regulation framework construction module is used to pre-build a distribution network voltage regulation framework that includes residential distributed resources, multiple agents, and the distribution network environment; wherein, the residential distributed resources are based on residential buildings as the smallest control unit; The voltage regulation network construction module is used to perform distributed resource-coordinated voltage regulation of the distribution network based on a decentralized Actor network and a centralized Critic network within the voltage regulation framework of the distribution network; wherein, both the Actor network and the Critic network adopt a heterogeneous Transformer network architecture; The distributed resource setting module is used to set standardized action spaces and confidence levels for heterogeneous distributed resources of residents; The action value output module is used to combine the local state information and historical data of each agent to obtain a time series, and use the time series as the input of the corresponding Actor network to output the corresponding action value. The evaluation value output module is used to take the action values ​​and global state sequence of the multi-agent as input to the corresponding Critic network, and output the corresponding evaluation value in combination with the geographic information of the agent. The coordinated voltage regulation module is used to continuously perform coordinated voltage regulation based on the action value and the evaluation value, combined with the reward function, by using a Critic network to centrally train the model, and then using an Actor network to perform the decentralized execution. The reward function includes at least a voltage deviation term, which is used to characterize the voltage deviation of the distribution network.

9. A computer device for voltage regulation in a large-scale residential distributed resource coordination distribution network, characterized in that, Includes at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the large-scale residential distributed resource coordination distribution network voltage regulation method according to any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the large-scale residential distributed resource coordination distribution network voltage regulation method according to any one of claims 1 to 7.