A Multi-Objective Optimization Method for Flexible Power Distribution Systems Based on Reward Decoupling Normalization Reinforcement Learning

By employing a reward decoupling normalization strategy optimization method, which independently processes the rewards of each objective and combines them with a conditional reward mechanism, the problem of reward signal confusion in multi-objective collaborative optimization is solved. This achieves efficient and stable optimization of the power system and is applicable to the collaborative optimization of economy, safety and cleanliness in environments with a high proportion of renewable energy.

CN122371173APending Publication Date: 2026-07-10HOHAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HOHAI UNIV
Filing Date
2026-06-09
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies suffer from reward signal confusion and optimization objective ambiguity in multi-objective collaborative optimization, resulting in low efficiency in power system operation optimization and making it difficult to achieve coordinated optimization of economy, safety and cleanliness in environments with a high proportion of renewable energy.

Method used

We employ a reward decoupling normalization strategy optimization method. By independently normalizing the rewards of each objective and combining it with a conditional reward mechanism, we can accurately express the priority relationship between objectives and construct a deep reinforcement learning framework to achieve multi-objective optimization.

Benefits of technology

It achieves refined synergistic optimization among economy, safety and cleanliness in environments with a high proportion of renewable energy. The training process is stable, the computational efficiency is high, and it is suitable for online applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122371173A_ABST
    Figure CN122371173A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-objective optimization method for flexible distribution systems based on reward decoupling normalization reinforcement learning. This method constructs a multi-objective deep reinforcement learning model based on group reward decoupling normalization strategy optimization to perform multi-objective optimization scheduling for complex constraints in flexible distribution systems, such as renewable energy fluctuations, load changes, and scheduling resource coordination. The method designs the three objectives—minimizing operating costs, minimizing voltage offset, and minimizing photovoltaic curtailment—as independent reward functions, guiding the flexible distribution system's operation strategy to focus on multiple dimensions, including economy, environmental friendliness, and safety. The group reward decoupling normalization strategy optimization algorithm independently normalizes each reward, avoiding the collapse problem of multiple reward signals in advantage estimation and effectively preserving the relative differences between the objectives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of source-grid-load-storage collaborative optimization technology, specifically involving a multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning. Background Technology

[0002] As the energy transition deepens, the power system is shifting from a traditional centralized, fossil fuel-dominated model to a distributed model integrating a high proportion of renewable energy. This transition presents unprecedented challenges to the optimization of power system operation across multiple objectives. Economically, power market reforms require the system to minimize costs while meeting safety constraints. In terms of security, traditional safety issues such as voltage stability, frequency stability, and line load are exacerbated by the intermittency and volatility of renewable energy. Environmentally, reducing fossil fuel consumption and maximizing renewable energy utilization have become common goals. These three objectives are complexly coupled and even conflicting: for example, maximizing photovoltaic (PV) utilization may require increasing spinning reserves or adjusting network topology, thus increasing operating costs; maintaining voltage security may require limiting renewable energy output, leading to curtailment of solar and wind power.

[0003] While traditional mathematical programming methods can yield exact solutions, they heavily rely on model linearization and prior weight settings, making it difficult to characterize the complex nonlinear trade-offs between objectives. Intelligent optimization algorithms (such as multi-objective evolutionary algorithms), while capable of generating Pareto fronts, suffer from low computational efficiency and limitations in online applications. In recent years, deep reinforcement learning, with its powerful environmental interaction and sequential decision-making capabilities, has offered new insights into these problems. However, when transforming multiple objectives into multi-reward signals, mainstream policy optimization algorithms often suffer from signal collapse due to reward mixing and normalization, blurring the differences between objectives and causing optimization bias and training instability. Therefore, designing a deep reinforcement learning framework that accurately preserves the relative information of multiple rewards and achieves efficient and stable optimization has become a key challenge in promoting the practical application of intelligent scheduling. Summary of the Invention

[0004] The purpose of this invention is to provide a multi-objective reinforcement learning method for power systems based on a reward decoupling normalization strategy optimization, overcoming the problems of reward signal confusion and optimization objective ambiguity in existing technologies when dealing with multi-objective collaborative optimization. Specifically, it aims to effectively maintain the differences between cost, voltage security, and photovoltaic consumption by independently normalizing the rewards of each objective; accurately express the priority relationship between objectives through a conditional reward mechanism; and ultimately achieve refined collaborative optimization among economy, safety, and cleanliness in a high-proportion renewable energy environment, providing a new intelligent decision-making tool for new power systems that is convergent, stable, clearly decoupled, and easy to apply online.

[0005] Technical Solution: To address the aforementioned technical problems, this invention provides a multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning. This method includes the following steps:

[0006] Step 1: Obtain initial grid topology information, grid user load information, and scheduling resource configuration information for photovoltaic, energy storage, and gas turbines;

[0007] Step 2: Based on the network, load information and scheduling resource configuration information obtained in Step 1, and considering the power grid balance, line connectivity constraints and scheduling resource operation constraints, construct a multi-objective scheduling model for the flexible distribution system.

[0008] Step 3: Based on the multi-objective scheduling model of the flexible power distribution system obtained in Step 2, and combined with the reinforcement learning algorithm, set the state space, action space, state transition and reward function of the decision-making agent. The reward function consists of three objectives: minimizing operating cost, minimizing voltage deviation and minimizing photovoltaic curtailment.

[0009] Step 4: Based on the decision-making agent settings in Step 3, and combined with the group reward decoupling normalization strategy optimization algorithm, perform independent normalization processing on each reward;

[0010] Step 5: Based on the reward normalization processing results in Step 4, and combined with historical photovoltaic power output data, train the scheduling model, and schedule the distribution network according to the training results to obtain the scheduling strategy.

[0011] Furthermore, in step 2, the multi-objective scheduling model for the flexible power distribution system is as follows:

[0012] 1) Objective function

[0013] (1)

[0014] in, For scheduling time; This is the total scheduling period; Indicates the system's first One node; For scheduling time scale; For the set of nodes in the upper-level power grid; For electricity purchase costs; Active power purchased by the distribution network from the upper-level power grid; A set of nodes; This is the load shedding penalty factor; This refers to the load shedding power; This refers to the set of nodes connected to the gas turbine. This represents the cost coefficient for gas turbines. To output active power for the gas turbine; A set of nodes connected to energy storage; The cost coefficient for energy storage charging and discharging; and These are the energy storage charging and discharging power, respectively; This refers to the node voltage amplitude. The set of nodes connected to photovoltaics; The cost of penalties for abandoning light; This refers to the amount of solar power that has been curtailed.

[0015] 2) Power balance constraints

[0016] (2)

[0017] (3)

[0018] (4)

[0019] (5)

[0020] (6)

[0021] in, Indicates the first A set of end nodes whose first node is a _________ node; Indicates the first A set of first and last nodes where each node is the last node; For distribution network branch collection; , The first The injected active power and injected reactive power of each node; , They are respectively the branches through which the flow passes Active power and reactive power; and Branch roads The resistance and reactance values; For flow through branch road The square of the current; Reactive power purchased from the upper-level power grid for distribution network; To output reactive power for the gas turbine; , Injecting active and reactive power into photovoltaic systems; , For flexible soft switching, there are nodes To the node Transmitted active and reactive power; , For flexible soft switching, there are nodes To the node Transmitted active and reactive power; , Active and reactive loads; For reactive load shedding; The square of the node voltage amplitude; Equations (2)-(5) are the power balance constraints of AC nodes in the distribution network; Equation (6) is the branch capacity constraint after second-order cone relaxation;

[0022] 3) Line connectivity constraints

[0023] (7)

[0024] (8)

[0025] (9)

[0026] (10)

[0027] (11)

[0028] (12)

[0029] in, This represents the maximum value of the line's transmitted current. , These represent the maximum active and reactive power values ​​for line transmission, respectively. , These are the squares of the minimum and maximum node voltages, respectively; Equations (7) and (8) constrain the magnitude of the node voltages at both ends of the line; Equations (9)-(12) constrain the current and power on the line;

[0030] 4) Energy storage operation constraints

[0031] (13)

[0032] (14)

[0033] (15)

[0034] (16)

[0035] (17)

[0036] (18)

[0037] (19)

[0038] (20)

[0039] in, , To represent the 0-1 variables of energy storage charge and discharge states, This indicates that the energy storage is in a charging state. This indicates that the energy storage is in a discharging state. This indicates that the energy storage is in a non-charging state. This indicates that the energy storage is in a non-discharge state; This represents the maximum charging and discharging power. The amount of electricity stored for energy storage; , These are the upper and lower limits of energy storage capacity; , These are the energy storage charging and discharging efficiencies, respectively. The energy storage charge / discharge coefficient during the total scheduling period; This represents the maximum number of charge / discharge cycles per day.

[0040] 5) Operating constraints of photovoltaic units

[0041] (twenty one)

[0042] (twenty two)

[0043] (twenty three)

[0044] in, This refers to the upper limit of the active power output of the photovoltaic unit; This refers to the upper limit of reactive power output of photovoltaic units; This refers to the rated capacity of the photovoltaic unit.

[0045] 6) Power purchase constraints from the upper-level power grid

[0046] (twenty four)

[0047] in, Maximum power consumption for grid connection;

[0048] 7) Gas turbine operating constraints

[0049] (25)

[0050] (26)

[0051] in, This represents the upper limit of the active power output of the gas turbine. This represents the upper limit of reactive power output of the gas turbine.

[0052] 8) Gas turbine ramping power constraint

[0053] (27)

[0054] in, This represents the maximum ramp power of the gas turbine.

[0055] 9) Load shedding constraint

[0056] (28)

[0057] (29)

[0058] in, The power factor of the load node;

[0059] 10) Flexible soft-switching power constraints

[0060] (30)

[0061] (31)

[0062] (32)

[0063] (33)

[0064] (34)

[0065] in, Transmitting apparent power for flexible soft switching; , These are the upper limits for active and reactive power transmission of flexible soft switches, respectively.

[0066] Furthermore, in step 3, the state space, action space, state transition, and reward function of the decision-making agent are set as follows, in conjunction with the reinforcement learning algorithm:

[0067] 1) State Space

[0068] The state space includes the predicted photovoltaic output, the active power of the load, the energy storage capacity at the previous moment, and the electricity price.

[0069] (35)

[0070] In the formula, This represents the state space of the system at time t. It covers the predicted output values ​​of all photovoltaic units. This represents the set of load conditions for all nodes. This is the collection of the stored energy of all energy storage devices at the previous moment;

[0071] 2) Action Space

[0072] Actions of the agent at time t Defined as:

[0073] (36)

[0074] in, For all photovoltaic units in The collection of active power output at all times; For all photovoltaic units in The collection of reactive power that is constantly being generated; For all energy storage The set of output power at any given moment; For all active loads in The set of time consumption; For all reactive loads in The set of time consumption; For all SOPs in The collection of active power transmitted at any given moment; For all SOPs in The collection of reactive power transmitted at any given moment;

[0075] 3) State transition

[0076] The state transition process from time step t to t+1 can be represented as the optimal power flow solution process for a flexible power distribution system, influenced by the current state of the environment. Actions of the intelligent agent The influence of the environment at time t on the agent's behavior. Execute its operation Having obtained all its own scheduling decisions, the environment transforms according to the optimal power flow state. ;

[0077] Energy storage devices cannot be charged and discharged simultaneously. SoC status of individual energy storage devices By mutually exclusive charge and discharge quantity Energy limit and power limit and charge / discharge coefficient The common constraints are shown in the following formula:

[0078] (37)

[0079] (38)

[0080] in, For the set of energy storage actions in equation (36) The Middle The operating values ​​of an energy storage device; energy storage capacity. The transition from time t to time t+1 is expressed as equation (16);

[0081] 4) Reward Function

[0082] The reward function of the intelligent agent comprehensively considers three objectives: system economic cost, safe operation, and photovoltaic grid integration rate, and is expressed as follows:

[0083] (39)

[0084] (40)

[0085] (41)

[0086] (42)

[0087] in, This represents the reward value obtained by the agent at time t; As a reward discount factor; This represents the overall operating cost of the system at time t; This indicates the penalty for exceeding the system voltage limit at time t; This represents the penalty for solar curtailment at time t.

[0088] Furthermore, in step 4, the group reward decoupling normalization strategy optimization algorithm is used to perform independent normalization processing on each reward:

[0089] 1) Within-group advantage normalization for each reward calculation group

[0090] exist At any given moment, the agent receives a reward for each action, and the sample is processed. One action As a group, execute this separately. For each action, three sets of reward values ​​are obtained based on equations (40), (41), and (42). , , , among which, the The quality of an action is represented by its corresponding strategy advantage value:

[0091] (43)

[0092] in, Indicates for the first The time period, the first The first strategic action Within-group normalized strategy advantage value for each objective. ; The size of the group; Indicates for the first The time period, the first The first strategic action The target value of each objective. These correspond to equations (40), (41), and (42), respectively. For the first The time period, the first The first strategic action The average of the target values; For the first The time period, the first The first strategic action The variance of each target value;

[0093] 2) Perform normalized summation of advantages for multiple objectives and batch normalization.

[0094] The normalized advantage values ​​of each objective are summed to construct the final stable advantage value. Used to guide policy network updates;

[0095] (44)

[0096] (45)

[0097] in, Indicates the first The time period, the first Multi-objective fusion advantage value under each strategy action; This is the stable advantage value that is ultimately used for updating the reinforcement learning strategy; This represents the set of time periods included in the current training batch; For stable parameters; This is the average of the multi-objective fusion advantage values ​​across all time periods and policy actions in the current training batch. This represents the variance of the multi-objective fusion advantage value across all time periods and policy actions in the current training batch.

[0098] Furthermore, in step 5, the scheduling model is trained and solved, and the training and solving process is as follows:

[0099] 5.1) Initialize model parameters and training environment

[0100] Initialize the policy network and value network parameters of the agent, set the experience replay buffer, configure the dynamic learning rate scheduler, set a relatively high initial learning rate to accelerate model convergence, and set a learning rate decay mechanism with training rounds.

[0101] 5.2) State perception and control action generation

[0102] At the beginning of each training round, the agent acquires the current operating status data of the flexible power distribution system. The agent processes the status data based on the current policy network and outputs the corresponding grid dispatch actions, including the photovoltaic curtailment ratio, energy storage charging and discharging amount, load shedding amount, and SOP transmission power.

[0103] 5.3) Environmental Interaction and Reward Signal Calculation

[0104] The scheduling action generated in step 5.2) is sent to the flexible power distribution system environment for execution. The environment performs optimal power flow calculation based on the scheduling action to solve the new operating state of the power grid. Guided by minimizing the power grid operating cost, voltage over-limit penalty and photovoltaic curtailment penalty, the environment calculates and returns the corresponding reward signal to the agent.

[0105] 5.4) Experience tuple storage and replay mechanism

[0106] Combine the current state generated in steps 5.2) and 5.3), the action performed, the reward signal obtained, and the new state after performing the action into an experience tuple. The experience tuples are stored in the experience replay buffer pool, and the complete trajectory data from the initial state to the final state is continuously accumulated through continuous interaction with the environment.

[0107] 5.5) Network model update based on dominance estimation

[0108] When the amount of data in the buffer pool reaches a preset threshold or the set number of rounds is completed, the model is updated. Batch experience data is randomly sampled from the experience replay buffer pool, and the policy experience in the batch data is extracted. The advantage function value of each scheduling action is calculated using the reward decoupling normalization algorithm to evaluate the merits of the action relative to the average level. Based on the calculated advantage function value, the loss function of the policy network is constructed by combining the policy gradient algorithm. The loss function of the value network is calculated using mean squared error, and the gradient is calculated using the Adam optimizer to update the network weight parameters of the agent, so that the agent can choose the scheduling action that can obtain higher total reward in future interactions.

[0109] 5.6) Iterative Loop and Policy Convergence

[0110] Repeat steps 5.2) to 5.5). The training process includes multiple training rounds, each round containing a fixed number of training rounds. During this iterative learning process, the learning rate scheduler dynamically adjusts the learning rate of the Adam optimizer until the network model converges or reaches the maximum number of training rounds. Finally, the trained agent model is output, which contains the optimal scheduling strategy that minimizes the grid operating cost.

[0111] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:

[0112] This invention introduces a reward decoupling normalization mechanism, bringing significant benefits to multi-objective optimization of power systems. The group reward decoupling normalization strategy optimization algorithm effectively overcomes the defect of reward signal confusion in traditional multi-reward reinforcement learning, and performs better on the comprehensive Pareto front of the three objectives of cost, voltage security, and photovoltaic consumption. Decoupling normalization avoids the collapse of advantage estimates, making the training curve converge more smoothly and eliminating the late-stage performance collapse phenomenon common in traditional methods. By processing each reward independently, decision-makers can more clearly observe the competition and cooperation relationships between different objectives, and accurately realize strategic intentions such as safety priority and economic consideration through conditional reward mechanism. The proposed framework has low dependence on model linearity, can directly handle continuous-discrete mixed action space, and has fast online inference speed, providing a feasible path from simulation testing to actual dispatching systems. Attached Figure Description

[0113] Figure 1 This is a flowchart of the method of the present invention;

[0114] Figure 2 This is a comparison chart of training results between different algorithms;

[0115] Figure 3 This is a diagram showing the scheduling results of energy storage devices;

[0116] Figure 4 This is a diagram showing the gas turbine scheduling results;

[0117] Figure 5 This is a screenshot of the results of purchasing electricity online. Detailed Implementation

[0118] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0119] like Figure 1 As shown, this invention provides a multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning. The method includes the following steps:

[0120] Step 1: Obtain initial grid topology information, grid user load information, and scheduling resource configuration information for photovoltaic, energy storage, and gas turbines;

[0121] Step 2: Based on the network, load information and scheduling resource configuration information obtained in Step 1, and considering the power grid balance, line connectivity constraints and scheduling resource operation constraints, construct a multi-objective scheduling model for the flexible distribution system.

[0122] Step 3: Based on the multi-objective scheduling model of the flexible power distribution system obtained in Step 2, and combined with the reinforcement learning algorithm, set the state space, action space, state transition and reward function of the decision-making agent. The reward function consists of three objectives: minimizing operating cost, minimizing voltage deviation and minimizing photovoltaic curtailment.

[0123] Step 4: Based on the decision-making agent settings in Step 3, and combined with the group reward decoupling normalization strategy optimization algorithm, perform independent normalization processing on each reward;

[0124] Step 5: Based on the reward normalization processing results in Step 4, and combined with historical photovoltaic power output data, train the scheduling model, and schedule the distribution network according to the training results to obtain the scheduling strategy.

[0125] Furthermore, in step 2, the multi-objective scheduling model for the flexible power distribution system is as follows:

[0126] 1) Objective function

[0127] (1)

[0128] in, For scheduling time; This is the total scheduling period; Indicates the system's first One node; For scheduling time scale; For the set of nodes in the upper-level power grid; For electricity purchase costs; Active power purchased by the distribution network from the upper-level power grid; A set of nodes; This is the load shedding penalty factor; This refers to the load shedding power; This refers to the set of nodes connected to the gas turbine. This represents the cost coefficient for gas turbines. To output active power for the gas turbine; A set of nodes connected to energy storage; The cost coefficient for energy storage charging and discharging; and These are the energy storage charging and discharging power, respectively; This refers to the node voltage amplitude. The set of nodes connected to photovoltaics; The cost of penalties for abandoning light; This refers to the amount of solar power that has been curtailed.

[0129] 2) Power balance constraints

[0130] (2)

[0131] (3)

[0132] (4)

[0133] (5)

[0134] (6)

[0135] in, Indicates the first A set of end nodes whose first node is a _________ node; Indicates the first A set of first and last nodes where each node is the last node; For distribution network branch collection; , The first The injected active power and injected reactive power of each node; , They are respectively the branches through which the flow passes Active power and reactive power; and Branch roads The resistance and reactance values; For flow through branch road The square of the current; Reactive power purchased from the upper-level power grid for distribution network; To output reactive power for the gas turbine; , Injecting active and reactive power into photovoltaic systems; , For flexible soft switching, there are nodes To the node Transmitted active and reactive power; , For flexible soft switching, there are nodes To the node Transmitted active and reactive power; , Active and reactive loads; For reactive load shedding; The square of the node voltage amplitude; Equations (2)-(5) are the power balance constraints of AC nodes in the distribution network; Equation (6) is the branch capacity constraint after second-order cone relaxation;

[0136] 3) Line connectivity constraints

[0137] (7)

[0138] (8)

[0139] (9)

[0140] (10)

[0141] (11)

[0142] (12)

[0143] in, This represents the maximum value of the line's transmitted current. , These represent the maximum active and reactive power values ​​for line transmission, respectively. , These are the squares of the minimum and maximum node voltages, respectively; Equations (7) and (8) constrain the magnitude of the node voltages at both ends of the line; Equations (9)-(12) constrain the current and power on the line;

[0144] 4) Energy storage operation constraints

[0145] (13)

[0146] (14)

[0147] (15)

[0148] (16)

[0149] (17)

[0150] (18)

[0151] (19)

[0152] (20)

[0153] in, , To represent the 0-1 variables of energy storage charge and discharge states, This indicates that the energy storage is in a charging state. This indicates that the energy storage is in a discharging state. This indicates that the energy storage is in a non-charging state. This indicates that the energy storage is in a non-discharge state; This represents the maximum charging and discharging power. The amount of electricity stored for energy storage; , These are the upper and lower limits of energy storage capacity; , These are the energy storage charging and discharging efficiencies, respectively. The energy storage charge / discharge coefficient during the total scheduling period; This represents the maximum number of charge / discharge cycles per day.

[0154] 5) Operating constraints of photovoltaic units

[0155] (twenty one)

[0156] (twenty two)

[0157] (twenty three)

[0158] in, This refers to the upper limit of the active power output of the photovoltaic unit; This refers to the upper limit of reactive power output of photovoltaic units; This refers to the rated capacity of the photovoltaic unit.

[0159] 6) Power purchase constraints from the upper-level power grid

[0160] (twenty four)

[0161] in, Maximum power consumption for grid connection;

[0162] 7) Gas turbine operating constraints

[0163] (25)

[0164] (26)

[0165] in, This represents the upper limit of the active power output of the gas turbine. This represents the upper limit of reactive power output of the gas turbine.

[0166] 8) Gas turbine ramping power constraint

[0167] (27)

[0168] in, This represents the maximum ramp power of the gas turbine.

[0169] 9) Load shedding constraint

[0170] (28)

[0171] (29)

[0172] in, The power factor of the load node;

[0173] 10) Flexible soft-switching power constraints

[0174] (30)

[0175] (31)

[0176] (32)

[0177] (33)

[0178] (34)

[0179] in, Transmitting apparent power for flexible soft switching; , These are the upper limits for active and reactive power transmission of flexible soft switches, respectively.

[0180] Furthermore, in step 3, the state space, action space, state transition, and reward function of the decision-making agent are set as follows, in conjunction with the reinforcement learning algorithm:

[0181] 1) State Space

[0182] The state space includes the predicted photovoltaic output, the active power of the load, the energy storage capacity at the previous moment, and the electricity price.

[0183] (35)

[0184] In the formula, This represents the state space of the system at time t. It covers the predicted output values ​​of all photovoltaic units. This represents the set of load conditions for all nodes. This is the collection of the stored energy of all energy storage devices at the previous moment;

[0185] 2) Action Space

[0186] Actions of the agent at time t Defined as:

[0187] (36)

[0188] in, For all photovoltaic units in The collection of active power output at all times; For all photovoltaic units in The collection of reactive power that is constantly being generated; For all energy storage The set of output power at any given moment; For all active loads in The set of time consumption; For all reactive loads in The set of time consumption; For all SOPs in The collection of active power transmitted at any given moment; For all SOPs in The collection of reactive power transmitted at any given moment;

[0189] 3) State transition

[0190] The state transition process from time step t to t+1 can be represented as the optimal power flow solution process for a flexible power distribution system, influenced by the current state of the environment. Actions of the intelligent agent The influence of the environment at time t on the agent's behavior. Execute its operation Having obtained all its own scheduling decisions, the environment transforms according to the optimal power flow state. ;

[0191] Energy storage devices cannot be charged and discharged simultaneously. SoC status of individual energy storage devices By mutually exclusive charge and discharge quantity Energy limit and power limit and charge / discharge coefficient The common constraints are shown in the following formula:

[0192] (37)

[0193] (38)

[0194] in, For the set of energy storage actions in equation (36) The Middle The operating values ​​of an energy storage device; energy storage capacity. The transition from time t to time t+1 is expressed as equation (16);

[0195] 4) Reward Function

[0196] The reward function of the intelligent agent comprehensively considers three objectives: system economic cost, safe operation, and photovoltaic grid integration rate, and is expressed as follows:

[0197] (39)

[0198] (40)

[0199] (41)

[0200] (42)

[0201] in, This represents the reward value obtained by the agent at time t; As a reward discount factor; This represents the overall operating cost of the system at time t; This indicates the penalty for exceeding the system voltage limit at time t; This represents the penalty for solar curtailment at time t.

[0202] Furthermore, in step 4, the group reward decoupling normalization strategy optimization algorithm is used to perform independent normalization processing on each reward:

[0203] 1) Within-group advantage normalization for each reward calculation group

[0204] exist At any given moment, the agent receives a reward for each action, and the sample is processed. One action As a group, execute this separately. For each action, three sets of reward values ​​are obtained based on equations (40), (41), and (42). , , , among which, the The quality of an action is represented by its corresponding strategy advantage value:

[0205] (43)

[0206] in, Indicates for the first The time period, the first The first strategic action Within-group normalized strategy advantage value for each objective. ; The size of the group; Indicates for the first The time period, the first The first strategic action The target value of each objective. These correspond to equations (40), (41), and (42), respectively. For the first The time period, the first The first strategic action The average of the target values; For the first The time period, the first The first strategic action The variance of each target value;

[0207] 2) Perform normalized summation of advantages for multiple objectives and batch normalization.

[0208] The normalized advantage values ​​of each objective are summed to construct the final stable advantage value. Used to guide policy network updates;

[0209] (44)

[0210] (45)

[0211] in, Indicates the first The time period, the first Multi-objective fusion advantage value under each strategy action; This is the stable advantage value that is ultimately used for updating the reinforcement learning strategy; This represents the set of time periods included in the current training batch; For stable parameters; This is the average of the multi-objective fusion advantage values ​​across all time periods and policy actions in the current training batch. This represents the variance of the multi-objective fusion advantage value across all time periods and policy actions in the current training batch.

[0212] Furthermore, in step 5, the scheduling model is trained and solved, and the training and solving process is as follows:

[0213] 5.1) Initialize model parameters and training environment

[0214] Initialize the policy network and value network parameters of the agent, set the experience replay buffer, configure the dynamic learning rate scheduler, set a relatively high initial learning rate to accelerate model convergence, and set a learning rate decay mechanism with training rounds.

[0215] 5.2) State perception and control action generation

[0216] At the beginning of each training round, the agent acquires the current operating status data of the flexible power distribution system. The agent processes the status data based on the current policy network and outputs the corresponding grid dispatch actions, including the photovoltaic curtailment ratio, energy storage charging and discharging amount, load shedding amount, and SOP transmission power.

[0217] 5.3) Environmental Interaction and Reward Signal Calculation

[0218] The scheduling action generated in step 5.2) is sent to the flexible power distribution system environment for execution. The environment performs optimal power flow calculation based on the scheduling action to solve the new operating state of the power grid. Guided by minimizing the power grid operating cost, voltage over-limit penalty and photovoltaic curtailment penalty, the environment calculates and returns the corresponding reward signal to the agent.

[0219] 5.4) Experience tuple storage and replay mechanism

[0220] Combine the current state generated in steps 5.2) and 5.3), the action performed, the reward signal obtained, and the new state after performing the action into an experience tuple. The experience tuples are stored in the experience replay buffer pool, and the complete trajectory data from the initial state to the final state is continuously accumulated through continuous interaction with the environment.

[0221] 5.5) Network model update based on dominance estimation

[0222] When the amount of data in the buffer pool reaches a preset threshold or the set number of rounds is completed, the model is updated. Batch experience data is randomly sampled from the experience replay buffer pool, and the policy experience in the batch data is extracted. The advantage function value of each scheduling action is calculated using the reward decoupling normalization algorithm to evaluate the merits of the action relative to the average level. Based on the calculated advantage function value, the loss function of the policy network is constructed by combining the policy gradient algorithm. The loss function of the value network is calculated using mean squared error, and the gradient is calculated using the Adam optimizer to update the network weight parameters of the agent, so that the agent can choose the scheduling action that can obtain higher total reward in future interactions.

[0223] 5.6) Iterative Loop and Policy Convergence

[0224] Repeat steps 5.2) to 5.5). The training process includes multiple training rounds, each round containing a fixed number of training rounds. During this iterative learning process, the learning rate scheduler dynamically adjusts the learning rate of the Adam optimizer until the network model converges or reaches the maximum number of training rounds. Finally, the trained agent model is output, which contains the optimal scheduling strategy that minimizes the grid operating cost.

[0225] Case Analysis

[0226] The following example illustrates the superiority of the multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning described in this invention. This invention employs an improved IEEE 33-bus power system. To compare the superiority of the proposed method, the group reward decoupling normalization strategy optimization algorithm and the swarm relative strategy optimization algorithm used in this paper are compared, and simultaneously compared with traditional mathematical programming methods.

[0227] Based on this example, the prediction performance and scheduling results under different algorithms are shown in [link to results]. Figure 2 The scheduling results for each are compared in Table 1; and the results for flexible resource response are also shown in Table 1. Figures 3-5 ,in, Figure 3 This is a diagram showing the scheduling results of energy storage devices; Figure 4 This is a diagram showing the gas turbine scheduling results; Figure 5 This is a graph showing the results of online electricity purchase. It's clear that the group reward decoupling normalization strategy optimization algorithm retains fine-grained guidance signals for multiple objectives, resulting in more accurate decision-making. This algorithm avoids the information collapse of group relative strategy optimization, allowing the relative merits of each objective to be independently transmitted to reinforcement learning, making training more stable and preventing later convergence failures. Batch normalization in group reward decoupling normalization strategy optimization is crucial for stable training, while group relative strategy optimization lacks batch normalization, and total reward normalization leads to uncontrolled numerical scaling, making it prone to objective rebound in the later stages of training. Furthermore, the multi-objective trade-offs are more balanced, aligning with actual dispatching needs. The core of power system dispatching is multi-objective balance rather than single-objective optimization. Group reward decoupling normalization strategy optimization, through decoupling normalization, ensures equal guidance strength for each objective, avoiding the dominance of easily optimized objectives. It also supports "fixed weights + reward conditionalization," allowing for flexible adjustment of objective priorities with more precise results. Compared to traditional algorithms, both reinforcement learning algorithms have higher computational efficiency and are more suitable for engineering applications.

[0228] Table 1 Comparison of the impact of different optimization methods on the running results

[0229]

[0230] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning, characterized in that, Includes the following steps: Step 1: Obtain initial grid topology information, grid user load information, and scheduling resource configuration information for photovoltaic, energy storage, and gas turbines; Step 2: Based on the network, load information and scheduling resource configuration information obtained in Step 1, and considering the power grid balance, line connectivity constraints and scheduling resource operation constraints, construct a multi-objective scheduling model for the flexible distribution system. Step 3: Based on the multi-objective scheduling model of the flexible power distribution system obtained in Step 2, and combined with the reinforcement learning algorithm, set the state space, action space, state transition and reward function of the decision-making agent. The reward function consists of three objectives: minimizing operating cost, minimizing voltage deviation and minimizing photovoltaic curtailment. Step 4: Based on the decision-making agent settings in Step 3, and combined with the group reward decoupling normalization strategy optimization algorithm, perform independent normalization processing on each reward; Step 5: Based on the reward normalization processing results in Step 4, and combined with historical photovoltaic power output data, train the scheduling model, and schedule the distribution network according to the training results to obtain the scheduling strategy.

2. The multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning according to claim 1, characterized in that, In step 2, the multi-objective scheduling model for the flexible power distribution system is as follows: 1) Objective function (1) in, For scheduling time; This is the total scheduling period; Indicates the system's first One node; For scheduling time scale; For the set of nodes in the upper-level power grid; For electricity purchase costs; Active power purchased by the distribution network from the upper-level power grid; A set of nodes; This is the load shedding penalty factor; This refers to the load shedding power; This refers to the set of nodes connected to the gas turbine. This represents the cost coefficient for gas turbines. To output active power for the gas turbine; A set of nodes connected to energy storage; The cost coefficient for energy storage charging and discharging; and These are the energy storage charging and discharging power, respectively; This refers to the node voltage amplitude. The set of nodes connected to photovoltaics; The cost of penalties for abandoning light; This refers to the amount of solar power that has been curtailed. 2) Power balance constraints (2) (3) (4) (5) (6) in, Indicates the first A set of end nodes whose first node is a _________ node; Indicates the first A set of first and last nodes where each node is the last node; For distribution network branch collection; , The first The injected active power and injected reactive power of each node; , They are respectively the branches through which the flow passes Active power and reactive power; and Branch roads The resistance and reactance values; For flow through branch road The square of the current; Reactive power purchased from the upper-level power grid for distribution network; To output reactive power for the gas turbine; , Injecting active and reactive power into photovoltaic systems; , For flexible soft switching, there are nodes To the node Transmitted active and reactive power; , For flexible soft switching, there are nodes To the node Transmitted active and reactive power; , Active and reactive loads; For reactive load shedding; The square of the node voltage amplitude; Equations (2)-(5) are the power balance constraints of AC nodes in the distribution network; Equation (6) is the branch capacity constraint after second-order cone relaxation; 3) Line connectivity constraints (7) (8) (9) (10) (11) (12) in, This represents the maximum value of the line's transmitted current. , These represent the maximum active and reactive power values ​​for line transmission, respectively. , These are the squares of the minimum and maximum node voltages, respectively; Equations (7) and (8) constrain the magnitude of the node voltages at both ends of the line; Equations (9)-(12) constrain the current and power on the line; 4) Energy storage operation constraints (13) (14) (15) (16) (17) (18) (19) (20) in, , To represent the 0-1 variables of energy storage charge and discharge states, This indicates that the energy storage is in a charging state. This indicates that the energy storage is in a discharging state. This indicates that the energy storage is in a non-charging state. This indicates that the energy storage is in a non-discharge state; This represents the maximum charging and discharging power. The amount of electricity stored for energy storage; , These are the upper and lower limits of energy storage capacity; , These are the energy storage charging and discharging efficiencies, respectively. The energy storage charge / discharge coefficient during the total scheduling period; This represents the maximum number of charge / discharge cycles per day. 5) Operating constraints of photovoltaic units (21) (22) (23) in, This refers to the upper limit of the active power output of the photovoltaic unit; This refers to the upper limit of reactive power output of photovoltaic units; This refers to the rated capacity of the photovoltaic unit. 6) Power purchase constraints from the upper-level power grid (24) in, Maximum power consumption for grid connection; 7) Gas turbine operating constraints (25) (26) in, This represents the upper limit of the active power output of the gas turbine. This represents the upper limit of reactive power output of the gas turbine. 8) Gas turbine ramping power constraint (27) in, This represents the maximum ramp power of the gas turbine. 9) Load shedding constraint (28) (29) in, The power factor of the load node; 10) Flexible soft-switching power constraints (30) (31) (32) (33) (34) in, Transmitting apparent power for flexible soft switching; , These are the upper limits for active and reactive power transmission of flexible soft switches, respectively.

3. The multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning according to claim 2, characterized in that, In step 3, using reinforcement learning algorithms, the state space, action space, state transition, and reward function of the decision-making agent are set as follows: 1) State Space The state space includes the predicted photovoltaic output, the active power of the load, the energy storage capacity at the previous moment, and the electricity price. (35) In the formula, This represents the state space of the system at time t. It covers the predicted output values ​​of all photovoltaic units. This represents the set of load conditions for all nodes. This is the collection of the stored energy of all energy storage devices at the previous moment; 2) Action Space Actions of the agent at time t Defined as: (36) in, For all photovoltaic units in The collection of active power output at all times; For all photovoltaic units in The collection of reactive power that is constantly being generated; For all energy storage The set of output power at any given moment; For all active loads in The set of time consumption; For all reactive loads in The set of time consumption; For all SOPs in The collection of active power transmitted at any given moment; For all SOPs in The collection of reactive power transmitted at any given moment; 3) State transition The state transition process from time step t to t+1 can be represented as the optimal power flow solution process for a flexible power distribution system, influenced by the current state of the environment. Actions of the intelligent agent The influence of the environment at time t on the agent's behavior. Execute its operation Having obtained all its own scheduling decisions, the environment transforms according to the optimal power flow state. ; Energy storage devices cannot be charged and discharged simultaneously. SoC status of individual energy storage devices By mutually exclusive charge and discharge quantity Energy limit and power limit and charge / discharge coefficient The common constraints are shown in the following formula: (37) (38) in, For the set of energy storage actions in equation (36) The Middle The operating values ​​of an energy storage device; energy storage capacity. The transition from time t to time t+1 is expressed as equation (16); 4) Reward Function The reward function of the intelligent agent comprehensively considers three objectives: system economic cost, safe operation, and photovoltaic grid integration rate, and is expressed as follows: (39) (40) (41) (42) in, This represents the reward value obtained by the agent at time t; As a reward discount factor; This represents the overall operating cost of the system at time t; This indicates the penalty for exceeding the system voltage limit at time t; This represents the penalty for solar curtailment at time t.

4. The multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning according to claim 3, characterized in that, In step 4, the group reward decoupling normalization strategy optimization algorithm is used to perform independent normalization on each reward: 1) Within-group advantage normalization for each reward calculation group exist At any given moment, the agent receives a reward for each action, and the sample is processed. One action As a group, execute this separately. For each action, three sets of reward values ​​are obtained based on equations (40), (41), and (42). , , , among which, the The quality of an action is represented by its corresponding strategy advantage value: (43) in, Indicates for the first The time period, the first The first strategic action Within-group normalized strategy advantage value for each objective. ; The size of the group; Indicates for the first The time period, the first The first strategic action The target value of each objective. These correspond to equations (40), (41), and (42), respectively. For the first The time period, the first The first strategic action The average of the target values; For the first The time period, the first The first strategic action The variance of each target value; 2) Perform normalized summation of advantages for multiple objectives and batch normalization. The normalized advantage values ​​of each objective are summed to construct the final stable advantage value. Used to guide policy network updates; (44) (45) in, Indicates the first The time period, the first Multi-objective fusion advantage value under each strategy action; This is the stable advantage value that is ultimately used for updating the reinforcement learning strategy; This represents the set of time periods included in the current training batch; For stable parameters; This is the average of the multi-objective fusion advantage values ​​across all time periods and policy actions in the current training batch. This represents the variance of the multi-objective fusion advantage value across all time periods and policy actions in the current training batch.

5. The multi-objective optimization method for flexible power distribution systems based on reward decoupling normalization reinforcement learning according to claim 4, characterized in that, In step 5, the scheduling model is trained and solved. The training and solving process is as follows: 5.1) Initialize model parameters and training environment Initialize the policy network and value network parameters of the agent, set the experience replay buffer, configure the dynamic learning rate scheduler, set a relatively high initial learning rate to accelerate model convergence, and set a learning rate decay mechanism with training rounds. 5.2) State perception and control action generation At the beginning of each training round, the agent acquires the current operating status data of the flexible power distribution system. The agent processes the status data based on the current policy network and outputs the corresponding grid dispatch actions, including the photovoltaic curtailment ratio, energy storage charging and discharging amount, load shedding amount, and SOP transmission power. 5.3) Environmental Interaction and Reward Signal Calculation The scheduling action generated in step 5.2) is sent to the flexible power distribution system environment for execution. The environment performs optimal power flow calculation based on the scheduling action to solve the new operating state of the power grid. Guided by minimizing the power grid operating cost, voltage over-limit penalty and photovoltaic curtailment penalty, the environment calculates and returns the corresponding reward signal to the agent. 5.4) Experience tuple storage and replay mechanism Combine the current state generated in steps 5.2) and 5.3), the action performed, the reward signal obtained, and the new state after performing the action into an experience tuple. The experience tuples are stored in the experience replay buffer pool, and the complete trajectory data from the initial state to the final state is continuously accumulated through continuous interaction with the environment. 5.5) Network model update based on dominance estimation When the amount of data in the buffer pool reaches a preset threshold or the set number of rounds is completed, the model is updated; batches of experience data are randomly extracted from the experience replay buffer pool, the strategy experience in the batch data is extracted, and the advantage function value of each scheduling action is calculated using the reward decoupling normalization algorithm to evaluate the superiority or inferiority of the action relative to the average level. Based on the calculated advantage function value, the loss function of the policy network is constructed by combining the policy gradient algorithm; the loss function of the value network is calculated using mean squared error; the Adam optimizer is used to calculate the gradient and update the network weight parameters of the agent, so that the agent can choose the scheduling action that can obtain higher total reward in future interactions. 5.6) Iterative Loop and Policy Convergence Repeat steps 5.2) to 5.5). The training process includes multiple training rounds, each round containing a fixed number of training rounds. During this iterative learning process, the learning rate scheduler dynamically adjusts the learning rate of the Adam optimizer until the network model converges or reaches the maximum number of training rounds. Finally, the trained agent model is output, which contains the optimal scheduling strategy that minimizes the grid operating cost.