Low-voltage power distribution network multi-objective collaborative optimization method and system based on reinforcement learning

By introducing a multi-agent reinforcement learning network structure into the low-voltage distribution network and using an Actor-Critic network for centralized training, multi-objective phasor embedding is achieved, solving the problem of multi-objective collaborative optimization under distributed photovoltaic access and improving the overall optimization effect of the power grid.

CN121546641BActive Publication Date: 2026-04-21STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED
Filing Date
2026-01-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve multi-objective collaborative optimization control in large-scale distributed photovoltaic (PV) grid connection scenarios, exhibiting problems such as complex parameter optimization, long computation time, excessive communication latency, and insufficient dynamic response.

Method used

We adopt a multi-objective collaborative optimization method based on reinforcement learning. By constructing a multi-agent environment and using an Actor-Critic network structure for centralized training, we introduce a multi-objective phasor embedding mechanism to phasor the reward function and weights, thereby achieving collaborative optimization of multi-dimensional objectives and dynamic adaptive adjustment of weights.

Benefits of technology

In the scenario of large-scale distributed photovoltaic (PV) grid connection, multi-objective collaborative optimization control is achieved, which improves the overall optimization effect of voltage deviation, line loss and PV absorption rate, avoids the local optimum problem caused by fixed weights in traditional methods, and improves decision-making efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121546641B_ABST
    Figure CN121546641B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-objective collaborative optimization method and system for low-voltage distribution networks based on reinforcement learning. The method includes the following steps: Step S01. Treating adjustable devices in the low-voltage distribution network as agents, and acquiring historical operating data of multiple agents to form a training dataset; Step S02. Constructing a multi-objective optimization model and defining a phasor-form reward function and multi-objective weights; Step S03. Performing centralized training on the deep reinforcement learning model, during which the multi-objective weights are input into the neural network, learning to maximize the cumulative reward according to the reward phasor, and fusing the multi-objective weight vector and Q-value vector through a phasor fusion layer; Step S04. Acquiring real-time operating data of multiple agents in the low-voltage distribution network, and using the learned strategies to control the actions of each agent. This invention can achieve multi-objective collaborative optimization control in large-scale distributed photovoltaic grid connection scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of low-voltage distribution network technology, and in particular to a multi-objective collaborative optimization method and system for low-voltage distribution networks based on reinforcement learning. Background Technology

[0002] With the increasing abundance of flexible and controllable resources in distribution networks, flexibility has become one of the typical characteristics of new distribution networks. On the power source side, a high proportion of distributed energy sources are connected to the grid through power electronic devices, which not only complete the conversion of electrical energy forms but also provide functions such as power control, reactive power compensation, voltage regulation, and grid support. On the user side, the aggregation of new loads such as electric vehicles and data centers forms flexible demand response capabilities, which will become an important flexible resource for distribution network operation. On the grid side, new equipment such as power electronic transformers and intelligent soft switches can achieve real-time, precise, and continuous power and voltage regulation, promoting the evolution of distribution networks from open-loop radial networks to flexible multi-ring networks. Although the above-mentioned new elements enhance the flexibility of the distribution system at the physical level, the sheer number, spatial dispersion, significant differences in characteristics, and diverse stakeholders of controllable resources, along with the mixing of multiple time scales, strong uncertainties, and strong time-varying factors at the system level, have created complex large-scale coordinated operation and control problems that are difficult to solve with traditional technologies. With the introduction of new high-performance equipment and the continuous improvement of communication networks, the multi-cluster collaborative operation mode on the edge side has become an important means to improve the flexible operation of the power distribution system. Based on the autonomous control capability of the edge-side cluster, multiple intelligent terminals in the distribution area can interact with each other to exchange boundary information, realize the distributed solution of large-scale operation optimization problems, avoid the communication pressure caused by the aggregation of massive data, and improve the decision-making efficiency under source-load fluctuations.

[0003] Currently, the main methods for optimizing voltage control in power distribution networks are as follows:

[0004] 1. Control method based on multi-level cooperation

[0005] This type of multi-level collaborative control method employs a distributed energy regional autonomous regulation system. Real-time control targets are transmitted from the provincial automatic generation control module to the regional dispatch automatic power control module for decomposition and tracking execution. The regional dispatch automation system progressively uploads grid models and aggregated real-time data, while simultaneously issuing control commands to the county distribution automation system, which executes the decomposed commands. However, this multi-level collaborative control method still suffers from limitations in real-time perception of distributed energy output and excessive communication latency; that is, directly using existing multi-level models for collaborative autonomy has limitations.

[0006] 2. An Active Distribution Network Regional Multi-Agent Autonomous Collaborative Optimization Method Based on the Objective Cascading Method

[0007] This type of active distribution network regional multi-agent autonomous collaborative optimization method, based on the objective cascading method, treats the upper-level excitation signal and several controlled units such as loads, tie lines, wind turbines, photovoltaics, and batteries as a single agent. It considers the power supply and demand balance within the agent, introduces a multi-level distribution network-agent system, and decomposes the global scheduling problem by decoupling tie line power. This leads to the construction of a multi-agent optimization model, enabling collaborative control of multiple agents within the microgrid. However, this type of method suffers from complex parameter optimization and long computation time, making it difficult to guarantee the timeliness of collaborative optimization. Summary of the Invention

[0008] The technical problem to be solved by this invention is: in view of the above-mentioned problems existing in the prior art, this invention provides a multi-objective collaborative optimization method and system for low-voltage distribution networks based on reinforcement learning, which can realize the collaborative optimization of multi-dimensional objectives and dynamic weight adaptive adjustment, and can realize multi-objective collaborative optimization control in the scenario of large-scale distributed photovoltaic access.

[0009] To solve the above-mentioned technical problems, the technical solution proposed by this invention is as follows:

[0010] A multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning, comprising the following steps:

[0011] Step S01. Treat the adjustable equipment in the low-voltage distribution network as an intelligent agent in the low-voltage distribution network environment, and obtain the historical operation data of multiple intelligent agents in the low-voltage distribution network to form a training dataset;

[0012] Step S02. Construct a multi-objective optimization model. The optimization objectives in the model include the lowest voltage deviation, the lowest line loss, and the highest photovoltaic absorption rate. Define a reward function in phasor form and multi-objective weights to form a reward phasor and a weighted phasor. The reward phasor includes multiple complex phasor components. Each complex phasor component corresponds to an optimization objective. The complex phasor component encodes the strength of the optimization objective and the preferred phase.

[0013] Step S03. Use the training dataset to perform centralized training on the deep reinforcement learning model based on the Actor-Critic network structure. During the training process, multi-objective weights are input into the neural network. The Actor network and Critic network of the neural network learn to maximize the cumulative reward according to the reward phasor. The Critic network outputs a Q-value vector containing the Q-values ​​corresponding to multiple optimization objectives. The multi-objective weight vector and the Q-value vector are fused through the phasor fusion layer to obtain the fused Q-value vector.

[0014] Step S04. Obtain real-time operating data of multiple intelligent agents in the low-voltage distribution network, and control the actions of each intelligent agent using the strategies learned through centralized training based on the obtained real-time operating data to achieve collaborative optimization of the low-voltage distribution network area.

[0015] Further, in step S02, the expression for the reward phasor is:

[0016]

[0017] in, Represents the reward service vector. This represents the preference phase vector representing the optimization objective. Indicates the first The strength of each optimization objective Indicates the first The preference phase of each optimization objective is used to characterize the optimization direction and preference of the objective. The preference phase is obtained by combining the pitch angle and azimuth angle of the optimization objective. This indicates the number of optimization objectives.

[0018] Further, in step S02, the calculation expression for fusing the multi-objective weight vector and the Q-value vector through the phasor fusion layer to obtain the fused Q-value vector is as follows:

[0019]

[0020]

[0021]

[0022] in, Denotes a complex weight vector, the first... The complex weight vector of each optimization objective is , It is the first The weights of each optimization objective, It is the first The preference phase of an optimization objective Indicates taking the real part, This indicates the complex conjugate transpose. Indicates the first An optimization objective is in state ,action and weight vector Q value under, Represents the fused Q-value vector. Indicates the number of optimization objectives. It is a discount factor used to ensure that later rewards have a smaller impact on the reward function. T This is the state transition function. Indicates action strategy, Indicates the first An optimization objective is in state ,action and weight The reward below, Representing state Execute action After transitioning to state The probability of.

[0023] Furthermore, step S03, the step of centralized training of the deep reinforcement learning model based on the Actor-Critic network structure, includes:

[0024] Initialize the agent's policy network With the value network q and the corresponding target network, each agent collects local observation information O from the state space S and inputs it into the deep neural network to generate a weight vector. ;

[0025] The weight vector The agent receives the weight vector input into the Actor network and then applies it to the network according to the current policy. and adaptive discrimination network output action ;

[0026] Current environment Based on the output action Generate the next state And calculate the reward at the current moment. Collect tuples of state, action, reward, and next state. And store it in the experience replay pool;

[0027] A portion of the data is randomly sampled from the experience pool to update the network;

[0028] Determine if the conditions for stopping training are met. If not, return to retrieve the agent's running data and iterate again until the stopping conditions are met.

[0029] Furthermore, the deep reinforcement learning model also includes a multi-head attention mechanism for feature-weighted aggregation of the state embedding vectors of each agent, wherein the calculation expression of the multi-head attention mechanism is:

[0030]

[0031] in, ~ Each represents a different attention point. Indicates the number of attention heads. Indicates the output projection matrix;

[0032] The calculation expression for each attention head is:

[0033]

[0034] in, Let the query vector represent the first... The needs of an intelligent agent , , For the first The input state vector of each agent. Let the key vector represent the first Characteristics of an intelligent agent , Let the values ​​be vectors to represent the effects of the actions of other agents. , , which is the scaling factor for the feature dimension;

[0035] In multi-head attention mechanisms, the agent State embedding Intelligent agents with neighbors Features Through intelligent agents For intelligent agents attention weights Weighted aggregation generates context vectors. :

[0036]

[0037]

[0038] in, Represents intelligent agents The neighborhood group, , For hyperparameters, For intelligent agents With intelligent agents The degree of power coupling between them For intelligent agents With intelligent agents The similarity between them.

[0039] Furthermore, in step S03, the loss function used during training is:

[0040]

[0041] in, Indicates Smooth L1 loss, Indicates the state ,action and weight Below Q value, B This indicates the number of training samples sampled from the experience replay pool. Indicates the first A vector of target Q-values ​​for each sample, where λ represents the regularization coefficient. This represents the phase change of phasor weights between adjacent training iterations;

[0042] The critic network is updated by minimizing the loss function, the actor network executes actions based on the Q-values ​​fed back by the critic network, and the policy network is updated by maximizing the Q-values.

[0043] Furthermore, in step S03, during the training process, if the agent cannot find a feasible solution in the current multi-objective space, a safe photovoltaic reduction strategy is activated to gradually reduce photovoltaic output until the agent finds the optimal solution within a safe range. The safe photovoltaic reduction strategy is to first reduce the output according to the initial reduction ratio, then recalculate the distribution network state after reduction. If there are still nodes exceeding the limit, the reduction ratio is increased according to the specified ratio to continue the reduction. This process is repeated until the voltage is restored to the preset safe range and the minimum reduction ratio is recorded. During the training process, the minimum reduction ratio is used as the safe baseline state for training.

[0044] Furthermore, a soft update mechanism is used to update the target parameter networks of the Actor and Critic. During the update, a stable Q-value estimate of the target network is used, where the main network parameters are determined by... and The components are respectively the Actor network and the Critic network, and the target network parameters are composed of... and The components, corresponding to the main network, are proportionally updated after each training iteration or after each preset K steps.

[0045] A multi-objective collaborative optimization system for low-voltage distribution networks based on reinforcement learning, comprising:

[0046] The data acquisition module is used to treat the adjustable equipment in the low-voltage distribution network as intelligent agents in the low-voltage distribution network environment, and to acquire historical operating data of multiple intelligent agents in the low-voltage distribution network to form a training dataset.

[0047] The optimization model construction module is used to construct a multi-objective optimization model. The optimization objectives in the model include voltage deviation, line loss and photovoltaic absorption rate. The model defines a reward function in phasor form and multi-objective weights to form a reward phasor and a weighted phasor. The reward phasor includes multiple complex phasor components, each of which corresponds to an optimization objective. The complex phasor components encode the strength of the optimization objective and the preferred phase.

[0048] The model training module is used to perform centralized training on a deep reinforcement learning model based on the Actor-Critic network structure using a training dataset. During the training process, multi-objective weights are input into the neural network, and the Actor network and Critic network of the neural network learn to maximize the cumulative reward according to the reward phasor. The Critic network outputs a Q-value vector containing the Q-values ​​corresponding to multiple optimization objectives, and the multi-objective weight vector and the Q-value vector are fused through a phasor fusion layer to obtain a fused Q-value vector.

[0049] The real-time control module is used to acquire real-time operating data of multiple intelligent agents in the low-voltage distribution network. Based on the acquired real-time operating data, it uses strategies learned through centralized training to control the actions of each intelligent agent, thereby achieving collaborative optimization of the low-voltage distribution network areas.

[0050] An electronic device includes a processor and a memory, the memory being used to store a computer program, and the processor being used to execute the computer program to perform the method described above.

[0051] Compared with existing technologies, the advantages of this invention are as follows: By introducing a multi-objective phasor embedding mechanism into the multi-agent reinforcement learning network structure, this invention expands the reward function and weights from scalars to phasors, making it applicable to multi-objective optimization scenarios. This avoids the problem of repeatedly determining the weights of different objectives in the multi-objective optimization process of traditional reinforcement learning. At the same time, the multi-objective weights are phased and added to the Actor and Critic neural networks for synchronous iterative updates, allowing learning to be performed directly in the vector-valued function space. This enables the optimization process to simultaneously consider the trade-offs and cooperative relationships between objectives, allowing the agent to make decisions directly in the multi-dimensional phasor reward space. This achieves cooperative optimization and dynamic weight adaptive adjustment of multi-dimensional objectives, enabling cooperative optimization control of multiple objectives in distributed photovoltaic large-scale grid connection scenarios. Attached Figure Description

[0052] Figure 1 This is a schematic diagram illustrating the implementation process of the multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning in this embodiment.

[0053] Figure 2 This is a schematic diagram of the V-MATD3 framework constructed in this embodiment.

[0054] Figure 3 This is a schematic diagram illustrating the principle of the Actor and critic neural network structure in this embodiment.

[0055] Figure 4 This is a schematic diagram illustrating the principle of centralized training and distributed execution in this embodiment.

[0056] Figure 5This is a detailed flowchart illustrating the centralized training and distributed execution process in this embodiment.

[0057] Figure 6 This is a schematic diagram of the reward curve test results obtained using the traditional MATD3 algorithm in a specific application embodiment.

[0058] Figure 7 This is a schematic diagram of the reward curve test results obtained by using the present invention in a specific application embodiment.

[0059] Figure 8 This is a schematic diagram of the results of the three-dimensional Pareto front in a specific application embodiment. Detailed Implementation

[0060] The present invention will be further described below with reference to the accompanying drawings and specific preferred embodiments, but this does not limit the scope of protection of the present invention.

[0061] To facilitate understanding, the relevant technical background of this invention will be explained first.

[0062] With the increasing penetration of distributed photovoltaic (PV) systems, voltage fluctuations at nodes in the distribution network are intensifying. Traditional centralized voltage control strategies struggle to respond to dynamic changes in real time. Furthermore, the large-scale integration of distributed PV into low-voltage distribution networks leads to voltage exceedances and increased line losses, while reducing PV output is not economically viable. Introducing reinforcement learning (RL) technology into distribution network optimization control can improve real-time response efficiency. However, single-objective RL methods (such as DDPG and TD3) can only optimize a single performance index and cannot simultaneously address multiple objectives such as voltage deviation, line losses, and PV grid integration rate. Existing multi-objective RL methods often rely on fixed preference weights, making it difficult to adaptively adjust the trade-offs between different objectives and easily leading to suboptimal solutions. Furthermore, traditional multi-agent reinforcement learning algorithms (such as MADDPG and MATD3) typically employ a scalar reward mechanism, which compresses multiple objectives into a single scalar form through weighted summation for gradient updates. However, this approach has the following problems: 1) Strong objective coupling: There is competition between different optimization objectives (such as minimizing voltage deviation vs. maximizing photovoltaic output), and scalarization loses the balance information between objectives; 2) Difficulty in capturing the Pareto optimal boundary: A single scalar objective compresses the solution space into a single direction, making it difficult to search for the complete Pareto front; 3) Gradient propagation distortion: When the gradient directions of different objectives conflict with each other, the gradient update after scalar weighting may deviate from the optimal solution.

[0063] Low-voltage distribution networks typically need to balance multiple conflicting objectives. When maximizing one objective comes at the expense of another, decision-makers must find a compromise. This invention addresses this type of multi-objective optimization problem by vectorizing rewards based on reinforcement learning methods. This approach is suitable for multi-objective optimization scenarios, avoiding the drawback of traditional reinforcement learning which requires repeatedly determining the weights of different objectives during multi-objective optimization. If single-objective reinforcement learning methods are used in multi-objective optimization scenarios, simply combining multiple optimization objectives linearly (with pre-set weight coefficients for each objective) into a single objective task, the fixed weights, being manually set, are somewhat arbitrary and may not align with the maximum reward explored by the current strategy. In other words, the maximum reward may not be the optimal solution with the current multi-objective weights. This invention, based on reward vectorization, phasors the multi-objective weights to form a weight phasor. This weight phasor is then added to the Actor and Critic neural networks for synchronous iterative updates, enabling the neural network to find the optimal solution with the current weights in a multi-dimensional phasor space.

[0064] This invention achieves voltage optimization control of distribution networks by constructing a phase-quantized multi-objective reinforcement learning architecture. It introduces a multi-objective phasor embedding mechanism into the multi-agent reinforcement learning network structure to realize the collaborative optimization of multi-dimensional objectives and dynamic adaptive adjustment of weights. It transforms the reward function and weight phasors from scalars to phasors, enabling multi-objective collaborative optimization control in the scenario of large-scale distributed photovoltaic grid connection. Specifically, for multi-objective optimization problems, this invention extends the reward function and weights from scalars to vectors, realizing the phasing of rewards and weights to be applicable to multi-objective optimization scenarios. This avoids the problem of repeatedly determining the weights of different objectives in the multi-objective optimization process of traditional reinforcement learning. At the same time, the multi-objective weights are phased and the weight phasors are added to the Actor and Critic neural networks for synchronous iterative updates, so that learning can be carried out directly in the vector-valued function space. This allows the optimization process to consider the trade-offs and synergistic relationships between various objectives simultaneously. Thus, the agent can make decisions directly in the multi-dimensional phasor reward space. This can effectively overcome the uncertainty of distributed photovoltaic output and electricity load in low-voltage distribution networks, the risk of voltage exceeding limits and excessive line losses, but the contradiction of poor economic efficiency due to reducing photovoltaic output. It can also solve the problem of local optima caused by the mutual influence of multiple objectives in low-voltage distribution networks.

[0065] like Figure 1 As shown, the steps of the multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning in this embodiment include:

[0066] Step S01. Treat the adjustable equipment in the low-voltage distribution network as an intelligent agent in the low-voltage distribution network environment, and obtain the historical operation data of multiple intelligent agents in the low-voltage distribution network to form a training dataset;

[0067] Step S02. Construct a multi-objective optimization model. The optimization objectives in the model include voltage deviation, line loss and photovoltaic absorption rate. Define a reward function in phasor form and multi-objective weights to form a reward phasor and a weighted phasor. The reward phasor includes multiple complex phasor components. Each complex phasor component corresponds to an optimization objective. The complex phasor component encodes the strength of the optimization objective and the preferred phase.

[0068] Step S03. Use the training dataset to perform centralized training on the deep reinforcement learning model based on the Actor-Critic network structure. During the training process, multi-objective weights are input into the neural network. The Actor network and Critic network of the neural network learn to maximize the cumulative reward according to the reward phasor. The Critic network outputs a Q-value vector containing the Q-values ​​corresponding to multiple optimization objectives. The multi-objective weight vector and the Q-value vector are fused through the phasor fusion layer to obtain the fused Q-value vector.

[0069] Step S04. Obtain real-time operating data of multiple intelligent agents in the low-voltage distribution network, and control the actions of each intelligent agent using the strategies learned through centralized training based on the obtained real-time operating data to achieve collaborative optimization of the low-voltage distribution network area.

[0070] In this embodiment, a multi-objective reinforcement learning framework (Vectorized Multi-Objective TD3, V-MATD3) based on phase-quantized rewards is constructed based on the MATD3 (Multi-Agent Twin Delayed Deep Deterministic Policy Gradient) network. Figure 2 As shown, a multi-objective phasor embedding mechanism is introduced into the MATD3 network structure. A phasor-quantized reward module is used to realize the collaborative optimization of multi-dimensional objectives and the dynamic adaptive adjustment of weights. The reward function and weight phasors are transformed from scalars to phasors, and a phasor fusion layer is introduced to extend the output of the Critic network to a phasor Q function. The Pareto front is found in the multi-objective phasor space, so that the optimal solution under the current weight can be found in the multi-dimensional phasor space, and the collaborative optimization control of multiple objectives in the scenario of large-scale distributed photovoltaic access can be efficiently realized.

[0071] In step S01 of this embodiment, multi-objective parameters of the low-voltage distribution network are analyzed. Adjustable devices in the low-voltage distribution network (such as distributed photovoltaic inverters, charge and discharge controllers, and adjustable flexible loads such as charging piles) are considered as intelligent agents in the low-voltage distribution network environment. The set of intelligent agents is denoted as N. Historical operating data of different intelligent agents, including voltage, current, and power, are collected and preprocessed. Specifically, historical data such as voltage, current, and power of different intelligent agents such as distributed power sources, energy storage devices, and flexible loads can be collected through devices such as smart meters and distributed energy controllers. This data is then uploaded to the main station via the HPLC communication network and normalized for model training.

[0072] In step S02 of this embodiment, the voltage deviation is used as the basis for calculation. Minimum, line loss Minimum and photovoltaic grid connection rate The highest level constructs a multi-objective optimization model for low-voltage distribution networks with the optimization objective as the goal. The multi-objective optimization task of the system is defined as follows:

[0073] (1)

[0074] in, ~ Let each represent an optimization objective, and m represent the number of optimization objectives.

[0075] Voltage deviation Line loss and photovoltaic absorption rate The calculation expressions are as follows:

[0076] (2)

[0077] (3)

[0078] (4)

[0079] in, Indicates the actual voltage value. Indicates the rated voltage. Represents current. This represents the equivalent resistance of the circuit. This indicates the actual output active power of the photovoltaic system. This indicates the rated maximum output power of the photovoltaic system.

[0080] It is understandable that, in addition to the optimization objectives mentioned above, other types of optimization objectives can also be adopted. The type and number of optimization objectives can be configured according to actual needs.

[0081] Multi-agent reinforcement learning is a process in which multiple self-controlled, interactive agents perceive the state of an environment through sensors and perform actions. Historical data and constraints are input into a low-voltage distribution network environment to simulate a real low-voltage distribution network. Reinforcement learning training is achieved through information interaction between the agents and the external environment. First, the agent selects actions to respond to the environment based on its perceived state. Then, it adjusts its policy by observing the results of these actions. Finally, the trained agent selects its policy based on its actions to achieve the optimal response to the environment and obtain the maximum reward.

[0082] In step S03 of this embodiment, the multi-objective collaborative optimization problem of the low-voltage distribution network is transformed into a distributed partially observable Markov decision-making process. The substation collaborative optimization problem is modeled as a multi-agent extension of the distributed partially observable Markov decision-making process, which can be described as a tuple containing 9 elements, namely (N, R, S, O, W, A, T, π, γ). Here, π represents the joint strategy, and γ is a discount factor that ensures that later rewards have a smaller impact on the reward function, including uncertainty about future rewards. Taking the adjustable devices of the low-voltage distribution network as agents, at time t, the state set S is defined as the set of global voltage, current, power, and other information at the current time. The observation set is defined as... ,in Represents intelligent agents The local observation information is used. The weight space composed of the weight vectors w output after processing the observation information is defined as W. Actions are defined: at time t, A is the joint action set, where the action set of agent n includes photovoltaic active and reactive power output, SVC output, and energy storage charging and discharging power. The state transition probability function is defined as T to represent the transition from the current state... and actions Transition to the state s of the next moment t+1 The probability of the reward R is the sum of the values ​​of the reward R and the reward R, which represents the probability of the agent n in state n. The reward R obtained from performing an action is typically designed based on the specific environment and learning objectives.

[0083] In step S02 of this embodiment, multi-dimensional rewards are calculated with the optimization objectives of minimizing voltage deviation, minimizing line loss, and maximizing photovoltaic absorption rate. The reward function is formed into a reward phasor using phasing, which can be expressed as:

[0084] (5)

[0085] in, Represents the reward service vector. This represents the preference phase vector representing the optimization objective. Indicates the first The strength of each optimization objective (i.e., amplitude, such as voltage deviation, line loss power, photovoltaic output ratio). Indicates the first The preference phase of two optimization objectives is used to characterize the optimization direction and preference of the objectives. When the phases of two objectives are close, it indicates that their optimization directions are consistent (i.e., cooperative objectives); when the phase difference is close to 180°, it indicates that there is a conflict in their optimization. This indicates the number of optimization objectives.

[0086] As shown in equation (5), in this embodiment, the multi-objective reward function R is no longer a scalar, but a phasor form, exploring the global optimal solution in the multi-objective space through complex weighted phasor. It can achieve a dynamic balance between competition and cooperation among objectives. Through dynamic phase tradeoffs, a controllable balance between competition and cooperation among multiple objectives can be achieved during training.

[0087] Specifically, preferred phase This is used to represent the optimization direction and preference for each optimization objective, with the preference phase... The position of each objective in the multidimensional optimization space is characterized by its optimization direction. Taking voltage deviation, line loss, and photovoltaic (PV) grid connection rate as examples, each objective has its own optimization direction. For instance, a smaller voltage deviation is better, so its optimization direction is to decrease; while a larger PV grid connection rate is better, so its optimization direction is to increase. By mapping each objective to a three-dimensional phase space, its preferred phase can be calculated. In three-dimensional space, the optimization direction of a target can be represented by the angle of a vector. Assume that the contribution of each target can be represented by a vector in three-dimensional space. To represent, then the vector The preferred phase can be determined by two angles in a spherical coordinate system (pitch angle). Azimuth Let ) represent, where:

[0088] Pitch angle Is the target vector and The angle between planes is calculated using the following formula:

[0089] (6)

[0090] in, , and These correspond to the projections of each target onto the three coordinate axes.

[0091] Azimuth It is the target vector and The angle between the axes is calculated using the following formula:

[0092] (7)

[0093] The preferred phase of each target can then be represented by a combination of pitch and azimuth angles:

[0094] (8)

[0095] This embodiment employs a reward function in the form of a complex phasor, enabling the trade-offs between multiple objectives to be encoded simultaneously in the form of "amplitude + phase". During training, the agent does not need to preset fixed weighting coefficients. Through the preference parameter (preference phase), it can adaptively explore the optimal balance point of different objectives, thereby improving the interpretability and generalization ability of the model.

[0096] In this embodiment, the multi-objective weight w is also in phasor form. , To optimize the number of objectives, for example, the weights for three optimization objectives can be expressed as: , ,in, These represent the weights of the optimization targets for voltage deviation, line loss, and photovoltaic absorption rate, respectively.

[0097] Weights are used to represent the importance of different objectives, and the preference phase is also considered. This indicates the direction of the objective optimization. The two are combined using a complex phasor to jointly represent the relative importance of the objective and the optimization direction. (Preference weight phasor) It is determined by the weight of each objective. and preferred phase Specifically, the weights and phases of each objective are combined into a complex phasor:

[0098] (9)

[0099] in, It is the first The weights of each optimization objective, It is the first The preference phase of each optimization objective.

[0100] In practical reinforcement learning, weights and preferred phase It is dynamic. As training progresses, the adjustment of weights and phase can adaptively balance the relationship between multiple objectives, thereby avoiding the limitations of fixed weights in traditional methods and helping the model adaptively find the optimal balance point between different objectives.

[0101] In this embodiment, the weight vector The corresponding values ​​indicate the relative importance of voltage deviation, line loss, and photovoltaic absorption rate, such as... Figure 3 As shown, this embodiment uses phasor-form weights w, which are then fed into the neural network. This allows the Actor and Critic networks to learn directly in the form of "vector reward / vector value," rather than simply weighting at the loss function level. This differs from traditional methods that map the reward phasor R to scalar utility. Unlike direct linear weighted training strategies, this embodiment inputs the weight phasor w into the neural network, and the Critic network trains according to the weight vector. State space and actions To calculate the multi-objective Q-value, the vector Q-value output by Critic is:

[0102] (10)

[0103] in, Indicates the first i An optimization objective is in state ,action and weight The Q value.

[0104] Specifically, in multi-objective optimization, the first i An optimization objective is in state ,action and weight vector The Q value can be calculated using the following formula:

[0105] (11)

[0106] in, Indicates the first There is an optimization objective, where γ is a discount factor to ensure that later rewards have a smaller impact on the reward function. This is the state transition function. Indicates action strategy, Indicates the first An optimization objective is in state ,action and weight The reward below, Representing state Execute action After transitioning to state The probability of.

[0107] The Critic network receives a state-action pair and the weight phasor of each objective Then, the network will evaluate each objective based on the current strategy, calculate the Q-value for each objective, and for each objective... The output of the Critic network Based on the expected return of target i under its current state and action, the Q-value is calculated using weight phasors. The relative importance of each objective is dynamically adjusted through weighted phasors. The calculation of the Q value for each objective affects the contribution of different objectives in the optimization process.

[0108] In this embodiment, during multi-objective optimization, the Critic network outputs vector Q-values ​​instead of traditional scalars. Each objective has an independent Q-value representing its expected reward, ultimately forming a multi-dimensional Q-value vector. This transforms Q-value calculation from an evaluation of a single objective to a comprehensive evaluation of multiple objectives. Scalar Q-value calculations consider only a single objective and are typically updated using a weighted average of immediate and future rewards. In contrast, this embodiment uses vector Q-values ​​to calculate the independent Q-value for each objective and employs dynamic weights... The influence of this can achieve a balance between multiple objectives without manually setting fixed weights, thus avoiding the suboptimal solution problem caused by fixed weights in traditional methods.

[0109] Simultaneously, a phasor fusion layer is introduced for fusion:

[0110] (12)

[0111] in, This represents a complex weight vector that can be dynamically adjusted. This indicates taking the real part, used to map the backscalar TD error. This represents the complex conjugate transpose, used to capture the interaction relationships between different targets. This represents the fused Q-value vector.

[0112] That is, by using the phasor fusion layer to fuse the multi-objective weight vector and the Q-value vector according to the above formula (8), the fused Q-value vector is obtained, so that different objectives remain directionally independent during gradient propagation. When optimizing the Actor, it no longer depends on a single scalar objective, but is based on multi-objective Pareto optimization and preference direction update of vector Q. This enables the model to capture the complete Pareto optimal surface during a training process, rather than a single point of a preference. This can avoid the problem of local optimal solutions caused by blindly using artificially fixed weights in multi-objective optimization, which leads to the reward obtained by the agent exploration not corresponding to the current weight direction.

[0113] When the reward function simply linearly combines multiple objectives with fixed weights, the preferences and policies become misaligned, and multiple objectives are coupled in the same direction, leading to local optima. The reward curve will always remain at a midpoint during exploration. This embodiment extends the traditional scalar TD3 framework to a vectorized form, forming the V-MATD3 architecture. After expanding the reward function from a scalar to a phasor, the preference vector w (multi-objective weight phasor) is added to the Actor and Critic neural networks for iterative processing. This allows the agent to simultaneously optimize multiple interrelated objectives in a high-dimensional objective space. The Pareto front is dynamically guided by the preference vector, exploring in the multi-objective phasor space. This eliminates coupling between multiple objectives and allows optimization from multiple objective directions, thus quickly obtaining the global optimum.

[0114] In traditional scalar reward mechanisms, the gradient propagation between multiple objectives is affected by the weighting coefficients when multiple objectives are weighted and summed. This can lead to some objectives being ignored or overemphasized during optimization. For example, if two objectives have opposite optimization directions (e.g., minimizing voltage deviation and maximizing photovoltaic absorption), traditional scalar methods will merge the gradients of these two objectives, resulting in inaccurate gradient updates or even deviations from the optimal solution. This embodiment introduces vector Q-values, where each objective is optimized in its own independent space. The Q-value of each objective is represented as an independent component in a vector, allowing collaboration and conflict between objectives to be explicitly handled during training.

[0115] This embodiment introduces phasor-based rewards and weights, ensuring that the preference phase of each objective remains independent during gradient propagation. In multi-objective reinforcement learning, cooperation or conflict between objectives can affect the direction of gradient propagation, thus impacting optimization performance. Furthermore, by introducing weighted phasors, the weights and phases of each objective can be dynamically adjusted based on the current training progress, ensuring the independence between objectives. In phasor form, the optimization direction between objectives is represented by the phase, while the magnitude of each objective represents its importance in the optimization process. This avoids gradient interference between objectives, allowing each objective to maintain its independent optimization direction during training. In this way, the model can independently optimize in a multi-dimensional space, thereby achieving the global optimum for multi-objective optimization. This embodiment further enhances the preservation of objective independence through a phasor fusion layer. The Q-value of each objective is optimized not only through traditional gradient backpropagation during training but also by considering the interactions between objectives. This fusion mechanism ensures the independence between objectives while allowing the model to flexibly weigh multiple objectives.

[0116] In step S03 of this embodiment, the intelligent agent is trained with a large amount of data. Multiple agents interact with the low-voltage distribution network environment. A multi-dimensional phasor reward function is used to maximize the cumulative reward, and a deep reinforcement learning model is trained centrally. Based on real-time state information, the expected cumulative reward for actions in a given distribution network area is evaluated using an Actor and critic neural network with a Q-value function. For example... Figure 4 As shown, the detailed steps for centralized training of a deep reinforcement learning model based on the Actor-Critic network structure include:

[0117] Initialize the agent's policy network π, value network q, and corresponding target network. Each agent collects local observation information O from the state space S and inputs it into the deep neural network to generate a weight vector. ;

[0118] Weight vector The agent, fed into the Actor network and receiving the weight vector, outputs an action based on the current policy network and the adaptive discriminant network. ;

[0119] Current environment Based on the output action Generate the next state And calculate the reward at the current moment. Collect tuples of state, action, reward, and next state. And store it in the experience replay pool;

[0120] A portion of the data is randomly sampled from the experience pool to update the network;

[0121] Determine if the conditions for stopping training are met. If not, return to retrieve the agent's running data and iterate again until the stopping conditions are met.

[0122] Specifically, the detailed process of solving the model using a centralized-trained multi-agent deep learning algorithm in this embodiment is as follows: Figure 5 As shown, during centralized training, the agent's policy network π and value network q, as well as the corresponding target network, are first initialized. Agent n collects local observation information O from the state space S and inputs it into the deep neural network to generate a weight vector w. The weight vector w is then input into the Actor network. Upon receiving the weight vector, the agent, based on the current policy network... and adaptive discrimination network output action The current environment according to Generate the next state s t+1 And calculate the reward at this moment. Collect tuples of state, action, reward, and next state. The data is stored in the experience replay pool; after sufficient experience is collected, a portion of the data is randomly sampled from the experience pool to update the network.

[0123] In this embodiment, the deep reinforcement learning model also includes a multi-head attention mechanism to perform feature-weighted aggregation of the state embedding vectors of each agent. The calculation expression for the multi-head attention mechanism is as follows:

[0124] (13)

[0125] in, ~ Each represents a different attention point. Indicates the number of attention heads. Indicates the output projection matrix;

[0126] The calculation expression for each attention head is:

[0127] (14)

[0128] in, Let the query vector represent the first... i The needs of an intelligent agent , , For the first i The input state vector of each agent. Let the key vector represent the first i Characteristics of an intelligent agent , Let the values ​​be vectors to represent the effects of the actions of other agents. , is the scaling factor for the feature dimension.

[0129] This embodiment introduces an attention layer into the V-MATD3 framework to introduce a multi-head attention mechanism, which performs feature-weighted aggregation on the state embedding vectors of each agent. Therefore, the Q-value of each agent will not only depend on its own state. It also dynamically relies on global context features obtained through attention aggregation, enabling each agent to adaptively adjust its attention scope under dynamic topology. At the same time, voltage coupling relationships are implicitly modeled (without explicit power flow equations).

[0130] In this embodiment, the attention weight is determined by both state similarity and power coupling, and is used to measure the influence between different stations. This mechanism can avoid blind decision-making by the agent and can dynamically focus on the neighboring node that has the greatest impact on it according to the attention weight, dynamically allocate communication and control resources, and achieve an adaptive transition from local optimum to global optimum.

[0131] Specifically, intelligent agents For intelligent agents attention weights By intelligent agents State characteristics and intelligent agents The similarity of the state features and their power coupling degree are calculated together. First, the state similarity is defined as... For intelligent agents and intelligent agents At the same time step Their states are respectively and State similarity is used to measure the degree of similarity between the current state characteristics of two agents in the operation of the power grid, and is calculated using cosine similarity:

[0132] (15)

[0133] Define the degree of power coupling as , representing intelligent agents With intelligent agents The physical coupling strength between the distribution substations or nodes in the power system is calculated using distribution network power flow data:

[0134] (16)

[0135] in , The first and the The amount of active and reactive power exchanged between nodes. This is the normalized base.

[0136] Combining the above two points, we define an intelligent agent. For intelligent agents Attention weights:

[0137] (17)

[0138] in , These are trainable hyperparameters used to adjust the weighting of the two terms in the attention weights; Represents intelligent agents The set of neighbors.

[0139] In multi-head attention mechanisms, the agent State embedding Will with neighboring intelligent agents Features pass Weighted aggregation generates context vectors. :

[0140] (18)

[0141] Aggregated context Will be with intelligent agents Self-state embedding The data is then spliced ​​or fused and fed into the Actor and Critic networks. In this way, the agent's value estimation or policy decision takes into account both its own state and the influence of its neighbors through the attention mechanism.

[0142] This embodiment embeds a multi-head attention mechanism into the state of each agent. Features of other intelligent agents This allows agents in multi-agent scenarios with centralized training and distributed execution to flexibly focus on their most influential neighbors, thereby achieving dynamic resource allocation. This contrasts with traditional multi-head attention mechanisms (MHA) which only apply to QK... T This embodiment differs in its attention weight calculation. It introduces the power coupling degree unique to the power system as an influencing factor, injecting domain-specific coupling into the attention mechanism. This makes the model more closely resemble the physical characteristics of the distribution network, rather than simply relying on the similarity of abstract feature spaces. Through this mechanism, each agent can selectively focus on which neighboring nodes to pay attention to and how much resource to allocate during decision-making. This transitions from considering only itself to considering the local neighborhood, and finally to focusing on global characteristics, achieving an adaptive resource allocation mechanism. This is particularly suitable for distribution network environments because the coupling relationships between nodes change with load, photovoltaic access, and topology variations. By focusing attention, redundant attention to neighboring nodes is reduced, resulting in higher training and execution efficiency. Traditional MHA generally lacks this dynamic switching of attention scope. Furthermore, this embodiment implicitly models physical coupling relationships such as voltage and power through the attention mechanism, allowing multi-head attention to autonomously discover the most influential nodes by learning the distribution network state embedding and coupling characteristics, without requiring explicit power flow equation calculations.

[0143] Because the Pareto boundary contains a large number of discrete solutions, the loss function exhibits significant fluctuations. Furthermore, since the loss function is based on mean squared error (MSE), which squares the error, it amplifies the influence of a few outliers, leading to large gradient oscillations and an uneven optimization path. This can result in deep valleys or high peaks on the surface of the loss function, complicating the optimization path. In other words, the mean squared error loss function leads to numerous local optima, making the optimization process prone to getting trapped in local optima and unable to find the global optimum. To address these issues, this embodiment further employs a loss function more robust to outliers: the Smooth L1 loss. Its behavior is similar to MSE for small errors but similar to MAE for large errors, without amplifying the magnitude, thus reducing the influence of discrete points and making the gradient more robust. Simultaneously, a phase regularization term is added, which is equivalent to adding a "directional consistency constraint" to the loss space. Essentially, this smooths the gradient of the vector Q output in the complex plane, ensuring the continuity of the multi-objective optimization process in the direction space and avoiding training oscillations caused by gradient conflicts between different objectives.

[0144] Specifically, in step S03 of this embodiment, the expression for the overall loss function used during training is as follows:

[0145] (19)

[0146] in, Indicates Smooth L1 loss, Indicates the state ,action and weight Below Q value, B This indicates the number of training samples sampled from the experience replay pool. Indicates the first A vector of target Q-values ​​for each sample, where λ represents the regularization coefficient. This represents the phase change of phasor weights between adjacent training iterations.

[0147] This embodiment updates the critic network by minimizing the aforementioned loss function. The actor network executes actions based on the Q-value fed back by the critic network, and updates the policy network by maximizing the Q-value. This avoids the optimization process from getting stuck in local optima, quickly finds the global optimum, and ensures the continuity of the multi-objective optimization process in the direction space, avoiding training oscillations caused by gradient conflicts between different objectives.

[0148] In step S03 of this embodiment, during the training process, if the agent cannot find a feasible solution in the current multi-objective space, a safe photovoltaic reduction strategy is activated to gradually reduce photovoltaic output until the agent finds the optimal solution within a safe range.

[0149] Specifically, the triggering conditions for the reduction strategy can be configured as follows: when the distributed photovoltaic system is already at full capacity, that is, its output power reaches the preset maximum output or actual power generation level; energy storage devices and adjustable load resources have been adjusted to their limits (remaining capacity is close to zero); the system detects that the voltage exceeds the limit; the intelligent embodiment strategy has no action combination in the multi-objective space under this state, that is, it is impossible to bring the voltage back to the safe range while maintaining other objective constraints.

[0150] Once the above conditions are met, the reduction process will begin: First, an initial reduction ratio of a0 = 5% is set as the first reduction step. After reduction, the distribution network state is recalculated to check if the voltage of all nodes is within the allowable range. If there are still nodes exceeding the limit, an additional 5% reduction ratio will be added, and the reduction will continue. The above steps are repeated until the system voltage recovers to the preset safe range. After reducing to a safe state, the minimum reduction ratio that makes the system safe is recorded. During training, the agent can then continue training as a safe baseline state, that is, continue to explore action strategies under this reduction ratio condition. Furthermore, a gradual recovery mechanism can be set up. For example, in a future training round or after a state change, the photovoltaic output ratio can be gradually restored to assess whether the system can return to full power while remaining safe. Finally, after the photovoltaic safety reduction strategy is triggered, the state-action pair of the environmental feedback will be added to the training memory pool. The agent will be penalized according to its reward function, the severity of the voltage exceeding the limit, and the reduction in photovoltaic utilization caused by the output reduction, and at the same time, the exploration of the current strategy will be stopped.

[0151] By introducing this safety reduction strategy, the system can effectively deal with infeasible solutions in multi-objective reinforcement learning training, avoiding training from getting stuck in a dead loop or continuously exceeding the limit.

[0152] In this embodiment, under the condition of prioritizing full photovoltaic power generation, if there is no remaining capacity in energy storage and adjustable load, and the voltage still exceeds the limit, that is, when the agent cannot find a feasible solution in the current multi-objective space, a safe photovoltaic reduction strategy will be activated to gradually reduce photovoltaic output until the agent finds the optimal solution within a safe range. Finally, the agent updates the target network through soft updates to alleviate the high-frequency parameter changes during training and improve training stability. After the above steps are completed, it is checked whether the conditions for stopping training are met. If not, the collection and observation information is returned and iterated again until the stopping conditions are met. The trained agent can be deployed to the actual distribution network system and then enter the distributed execution stage.

[0153] Specifically, after the reduction is complete and the system is restored to safety, a soft update mechanism is executed to update the target parameter networks of the Actor and Critic to alleviate the instability caused by frequent parameter changes during training. A single target network is used to stabilize the Q-value estimation. In this embodiment, the target network parameters are updated using a soft update mechanism, while the main network parameters are... and The components are respectively the Actor network and the Critic network, and the target network parameters are composed of... and The components, each corresponding to the main network, are defined in each training iteration or after each preset execution. After the step, instead of directly copying the main network parameters to the target network, a proportional hybrid update is performed, as shown in the following formula:

[0154] (20)

[0155] (twenty one)

[0156] The soft update coefficient τ typically ranges from 0.001 to 0.1. A smaller value results in slower updates to the target network and higher stability. Preferably, τ = 0.001 is used to control the target network to slowly follow the main network.

[0157] Since the Critic outputs multi-objective vector Q-values ​​and employs weight phasors w and an attention mechanism context, this embodiment combines a soft update approach. This not only ensures smooth changes in the target network parameters but also works in conjunction with the following elements: ensuring the predictive stability of the target network's multi-objective vector Q-value outputs, preventing changes due to weight adjustments. and attention The system oscillates due to significant adjustments; it improves the robustness of Q-value estimation in multi-objective and phasor spaces during training, avoiding distortion of the training signal in the target network due to rapid changes in the main network; and it improves the scalar performance when switching or adjusting weights. When using an attention mechanism structure, the target network follows slowly, making training more gradual and stable. This is due to the complex structure of "multi-target + phasor representation + attention aggregation". If hard updates or frequent updates are used, the target network may update too quickly, introducing high-frequency perturbations and affecting Q-vector estimation and weights. Adjustment, attention Dynamic allocation often fails to converge, but this embodiment improves the robustness of the target network by using a soft update mechanism, which updates the target network slowly during the mixing process.

[0158] like Figure 5As shown, in the distributed execution phase, each agent independently makes decisions based on feature vectors obtained from local observations (local observation data) and a centrally trained policy network. An adaptive discriminant network then provides corresponding actions to achieve coordinated optimization of distribution network areas. Specifically, agents with autonomous decision-making capabilities are trained using a deep reinforcement learning model. These agents continuously perceive the current state of the environment, select and execute corresponding actions based on preset or learned policies, and the environment provides feedback on the agent's actions to the next state and rewards the agent for the effectiveness of those actions. The agents iteratively maximize cumulative rewards, optimizing their policies through trial and error and learning, achieving efficient interaction with the environment. Through exploration and learning, they continuously experiment in the environment to find the optimal policy. This application trains devices with continuous adjustment capabilities based on the MATD3 framework. During training, parameters are optimized through coordination among multiple agents. Compared to traditional model-based optimization algorithms, this approach does not require a precise and complete physical model of the distribution network, offering better scalability. Furthermore, based on an offline centralized training and online distributed execution framework, agents only need to make decisions by observing local states, eliminating the need for complex communication equipment and prediction data, resulting in better real-time performance.

[0159] In specific application embodiments, compared with traditional TD3 and weighted MADDPG, the method of the present invention can reduce the mean square voltage deviation by about 12.3%, the system line loss by about 9.7%, and the photovoltaic absorption rate by about 8.5% under the same training rounds. Simultaneously, the training curve fluctuation amplitude is reduced by about 30%, and the reward curve of the agent during the training process is as follows: Figure 6 , 7 As shown, the horizontal axis represents the number of training rounds, and the vertical axis represents the reward value. The reward curve reflects the agent's performance and learning progress at different training stages. This curve allows observation of the agent's exploration process during training and its eventual convergence. Figure 6 In Figures (a) and (b), the reward curves of the traditional MATD3 algorithm are shown. In the early stages of training, the agent gradually gains rewards by exploring the environment, and the reward value shows an upward trend. However, after a certain number of rounds, the reward value tends to decrease and eventually stabilizes at a relatively low, intermediate position, indicating that the agent is trapped in a local optimum during training. Figure 7 Images (a) and (b) show the reward curves of the improved V-MATD3 multi-objective reinforcement learning algorithm based on phase quantization reward, as described in this invention. Figure 6As can be seen, unlike the traditional MATD3 algorithm, this invention introduces a phase quantization reward mechanism. The agent also undergoes an exploration period in the early stages of training, with the reward value gradually increasing. Furthermore, unlike traditional methods, the reward value of the agent in this invention continues to rise after a certain number of training rounds, eventually stabilizing at a relatively high value. This indicates that the agent has found a globally optimal solution through exploration and learning. This verifies that this invention can effectively avoid the dilemma of local optima and achieve better global optimization performance.

[0160] The three-dimensional Pareto front obtained using real training data, such as... Figure 8 The diagram illustrates the trade-offs between three optimization objectives (voltage deviation, line loss, and photovoltaic grid integration rate). Each point in the diagram represents a state-action pair, and the objective is to find the optimal solution by optimizing the combination of these objectives. The x-axis represents 1-voltage deviation, the y-axis represents 1-line loss, and the z-axis represents photovoltaic grid integration rate. Higher values ​​for these three objectives indicate better optimization performance. The Pareto front calculation method involves first obtaining the three objective values ​​for each solution. For each solution, if there are other solutions that are superior to the current solution in all objectives, then it does not belong to the Pareto front. By comparing the dominance relationships of the solutions, a set of solutions is found. This set of solutions constitutes the Pareto front in the objective space. By calculating the relative superiority relationships of all solutions, a set of non-dominated solutions (i.e., points on the Pareto front) is obtained. These solutions are then plotted in the objective space. Figure 8 The diagram shows multiple solutions under different conditions, with the optimal solution located on the Pareto front, meaning that no single objective can be further optimized without sacrificing the performance of other objectives. The optimal point in the diagram (marked with a pentagram) represents the best balance achieved through a trade-off between voltage deviation, line loss, and photovoltaic utilization.

[0161] This embodiment of the low-voltage distribution network multi-objective collaborative optimization system based on reinforcement learning includes:

[0162] The data acquisition module is used to treat the adjustable equipment in the low-voltage distribution network as intelligent agents in the low-voltage distribution network environment, and to acquire historical operating data of multiple intelligent agents in the low-voltage distribution network to form a training dataset.

[0163] The optimization model construction module is used to construct a multi-objective optimization model. The optimization objectives in the model include voltage deviation, line loss and photovoltaic absorption rate. The model defines a reward function in phasor form and multi-objective weights to form a reward phasor and a weighted phasor. The reward phasor includes multiple complex phasor components, each of which corresponds to an optimization objective. The complex phasor components encode the strength of the optimization objective and the preferred phase.

[0164] The model training module is used to perform centralized training on a deep reinforcement learning model based on the Actor-Critic network structure using a training dataset. During the training process, multi-objective weights are input into the neural network, and the Actor network and Critic network of the neural network learn to maximize the cumulative reward according to the reward phasor. The Critic network outputs a Q-value vector containing the Q-values ​​corresponding to multiple optimization objectives, and the multi-objective weight vector and the Q-value vector are fused through a phasor fusion layer to obtain a fused Q-value vector.

[0165] The real-time control module is used to acquire real-time operating data of multiple intelligent agents in the low-voltage distribution network. Based on the acquired real-time operating data, it uses strategies learned through centralized training to control the actions of each intelligent agent, thereby achieving collaborative optimization of the low-voltage distribution network areas.

[0166] The reinforcement learning-based multi-objective collaborative optimization system for low-voltage distribution networks in this embodiment corresponds one-to-one with the aforementioned reinforcement learning-based multi-objective collaborative optimization method for low-voltage distribution networks, and will not be described in detail here.

[0167] This embodiment further provides a computer device, including a processor and a memory, the memory being used to store a computer program, and the processor being used to execute the computer program to perform the method as described above.

[0168] It is understood that the method described in this embodiment can be executed by a single device, such as a computer or server, or it can be applied in a distributed scenario where multiple devices cooperate to complete the task. In a distributed scenario, one of the multiple devices may execute only one or more steps of the method described in this embodiment, and the multiple devices interact to complete the method. The processor can be implemented using a general-purpose CPU, microprocessor, application-specific integrated circuit, or one or more integrated circuits, and is used to execute relevant programs to implement the method described in this embodiment. The memory can be implemented using read-only memory (ROM), random access memory (RAM), static storage devices, and dynamic storage devices. The memory can store the operating system and other applications. When the method described in this embodiment is implemented through software or firmware, the relevant program code is stored in the memory and called and executed by the processor.

[0169] Those skilled in the art will understand that the above embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0170] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention should fall within the protection scope of the present invention.

Claims

1. A multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning, characterized by the following steps: include: Step S01. Treat the adjustable equipment in the low-voltage distribution network as an intelligent agent in the low-voltage distribution network environment, and obtain the historical operation data of multiple intelligent agents in the low-voltage distribution network to form a training dataset; Step S02. Construct a multi-objective optimization model. The optimization objectives in the model include the lowest voltage deviation, the lowest line loss, and the highest photovoltaic absorption rate. Define a reward function in phasor form and multi-objective weights to form a reward phasor and a weight phasor. The reward phasor includes multiple complex phasor components, each of which corresponds to an optimization objective. The complex phasor components encode the strength and preference phase of the optimization objective. Step S03. Use the training dataset to perform centralized training on the deep reinforcement learning model based on the Actor-Critic network structure. During the training process, multi-objective weights are input into the neural network. The Actor network and Critic network of the neural network learn to maximize the cumulative reward according to the reward phasor. The Critic network outputs a Q-value vector containing the Q-values ​​corresponding to multiple optimization objectives. The multi-objective weight vector and the Q-value vector are fused through the phasor fusion layer to obtain a fused Q-value vector. Step S04. Obtain real-time operating data of multiple intelligent agents in the low-voltage distribution network, and control the actions of each intelligent agent using the strategies learned through centralized training based on the obtained real-time operating data to achieve collaborative optimization of the low-voltage distribution network area. In step S02, the expression for the reward phasor is: in, Represents the reward service vector. This represents the preference phase vector representing the optimization objective. Indicates the first The strength of each optimization objective Indicates the first The preference phase of each optimization objective is used to characterize the optimization direction and preference of the objective. The preference phase is obtained by combining the pitch angle and azimuth angle of the optimization objective. This indicates the number of optimization objectives.

2. The multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning according to claim 1, characterized in that, In step S03, the calculation expression for the fused Q-value vector obtained by fusing the multi-objective weight vector and the Q-value vector through the phasor fusion layer is as follows: in, Denotes a complex weight vector, the first... The complex weight vector of each optimization objective is , It is the first The weights of each optimization objective, It is the first The preference phase of an optimization objective Indicates taking the real part, This indicates the complex conjugate transpose. Indicates the first An optimization objective is in state ,action and weight vector Q value under, Represents the fused Q-value vector. Indicates the number of optimization objectives. It is a discount factor used to ensure that later rewards have a smaller impact on the reward function. T This is the state transition function. Indicates action strategy, Indicates the first An optimization objective is in state ,action and weight The reward below, Representing state Execute action After transitioning to state The probability of.

3. The multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning according to claim 1, characterized in that, Step S03, the steps for centralized training of the deep reinforcement learning model based on the Actor-Critic network structure, include: Initialize the agent's policy network With the value network q and the corresponding target network, each agent collects local observation information O from the state space S and inputs it into the deep neural network to generate a weight vector. ; The weight vector The agent receives the weight vector from the input to the Actor network and then applies it to the network according to the current policy. and adaptive discrimination network output action ; Current environment Based on the output action Generate the next state And calculate the reward at the current moment. Collect tuples of state, action, reward, and next state. And store it in the experience replay pool; A portion of the data is randomly sampled from the experience pool to update the network; Determine if the conditions for stopping training are met. If not, return to retrieve the agent's running data and iterate again until the stopping conditions are met.

4. The multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning according to claim 1, characterized in that, The deep reinforcement learning model also includes a multi-head attention mechanism to perform feature-weighted aggregation of the state embedding vectors of each agent, wherein the calculation expression of the multi-head attention mechanism is: in, ~ Each represents a different attention point. Indicates the number of attention heads. Indicates the output projection matrix; The calculation expression for each attention head is: in, Let the query vector represent the first... The needs of an intelligent agent , , For the first The input state vector of each agent. Let the key vector represent the first Characteristics of an intelligent agent , Let the values ​​be vectors to represent the effects of the actions of other agents. , , which is the scaling factor for the feature dimension; In multi-head attention mechanisms, the agent State embedding Intelligent agents with neighbors Features Through intelligent agents For intelligent agents attention weights Weighted aggregation generates context vectors : in, Represents intelligent agents The neighborhood group, , For hyperparameters, For intelligent agents With intelligent agents The degree of power coupling between them For intelligent agents With intelligent agents The similarity between them.

5. The multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning according to any one of claims 1 to 4, characterized in that, In step S03, the loss function used during training is: in, Indicates Smooth L1 loss, Indicates the state ,action and weight Below Q value, B This indicates the number of training samples sampled from the experience replay pool. Indicates the first A vector of target Q-values ​​for each sample, where λ represents the regularization coefficient. This represents the phase change of phasor weights between adjacent training iterations; The critic network is updated by minimizing the loss function, the actor network executes actions based on the Q-values ​​fed back by the critic network, and the policy network is updated by maximizing the Q-values.

6. The multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning according to any one of claims 1 to 4, characterized in that, In step S03, during the training process, if the agent cannot find a feasible solution in the current multi-objective space, a safe photovoltaic reduction strategy is activated to gradually reduce photovoltaic output until the agent finds the optimal solution within the safe range. The safe photovoltaic reduction strategy is to first reduce the output according to the initial reduction ratio, and then recalculate the distribution network state. If there are still nodes that exceed the limit, the reduction ratio is increased according to the specified ratio to continue the reduction. This process is repeated until the voltage is restored to the preset safe range and the minimum reduction ratio is recorded. During the training process, the minimum reduction ratio is used as the safe baseline state for training.

7. The multi-objective cooperative optimization method for low-voltage distribution networks based on reinforcement learning according to claim 6, characterized in that, The target parameter networks of the Actor and Critic are updated using a soft update mechanism. During the update, a stable Q-value estimate of the target network is used, where the main network parameters are determined by... and The components are respectively the Actor network and the Critic network, and the target network parameters are composed of... and The components, corresponding to the main network, are configured in each training iteration or each preset execution. After the first step, perform a mixed update proportionally.

8. A multi-objective collaborative optimization system for low-voltage distribution networks based on reinforcement learning, characterized in that, include: The data acquisition module is used to treat the adjustable equipment in the low-voltage distribution network as intelligent agents in the low-voltage distribution network environment, and to acquire historical operating data of multiple intelligent agents in the low-voltage distribution network to form a training dataset. The optimization model construction module is used to construct a multi-objective optimization model. The optimization objectives in the model include voltage deviation, line loss, and photovoltaic absorption rate. A reward function in phasor form and multi-objective weights are defined to form a reward phasor and a weight phasor. The reward phasor includes multiple complex phasor components, each corresponding to an optimization objective. The complex phasor components encode the strength and preference phase of the optimization objective. The expression for the reward phasor is: in, Represents the reward service vector. This represents the preference phase vector representing the optimization objective. Indicates the first The strength of each optimization objective Indicates the first The preference phase of each optimization objective is used to characterize the optimization direction and preference of the objective. The preference phase is obtained by combining the pitch angle and azimuth angle of the optimization objective. Indicates the number of optimization objectives; The model training module is used to perform centralized training on a deep reinforcement learning model based on the Actor-Critic network structure using a training dataset. During the training process, multi-objective weights are input into the neural network, and the Actor network and Critic network of the neural network learn to maximize the cumulative reward according to the reward phasor. The Critic network outputs a Q-value vector containing the Q-values ​​corresponding to multiple optimization objectives, and the multi-objective weight vector and the Q-value vector are fused through a phasor fusion layer to obtain a fused Q-value vector. The real-time control module is used to acquire real-time operating data of multiple intelligent agents in the low-voltage distribution network. Based on the acquired real-time operating data, it uses strategies learned through centralized training to control the actions of each intelligent agent, thereby achieving collaborative optimization of the low-voltage distribution network areas.

9. An electronic device comprising a processor and a memory, the memory being used to store a computer program, characterized in that, The processor is used to execute the computer program to perform the multi-objective collaborative optimization method for low-voltage distribution networks based on reinforcement learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Power distribution network region voltage control method based on multi-agent deep reinforcement learning

    CN119070315A

  • Micro-grid cooperative scheduling method and device

    CN121052610A