User-side multi-target decision execution method and system

By constructing decision state vectors and iteratively training Markov policy networks, combined with multi-objective optimization, decision action vectors are generated and pruned, solving the problem of insufficient multi-objective decision-making accuracy on the user side and achieving higher decision accuracy and responsiveness.

CN122022400BActive Publication Date: 2026-07-10LISHUI POWER SUPPLY COMPANY OF STATE GRID ZHEJIANG ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LISHUI POWER SUPPLY COMPANY OF STATE GRID ZHEJIANG ELECTRIC POWER
Filing Date
2026-04-13
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

After large-scale integration of distributed renewable energy into the user side, existing technologies have failed to meet the target requirements for multi-objective decision-making at the user side. Traditional control strategies lack system-level collaborative information, and intelligent agent methods are highly complex and fail to consider the user side's response to power system commands.

Method used

The system collects operational status data from user-side devices, constructs decision state vectors, iteratively trains Markov policy networks, and generates and trims decision action vectors by combining operational costs, carbon emissions, and command response optimization objectives to ensure the comprehensiveness and accuracy of decision-making.

Benefits of technology

This improves the decision-making accuracy of user-side equipment, fully reflects the user-side response to power grid decision commands, and ensures the executability and accuracy of decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122022400B_ABST
    Figure CN122022400B_ABST
Patent Text Reader

Abstract

This application discloses a user-side multi-objective decision-making execution method and system, relating to the field of intelligent scheduling and decision-making. The method includes: collecting operating status data from multiple user-side devices to construct a decision state vector; inputting the decision state vector into a preset Markov policy network to output a decision action vector; wherein the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives; the action space constraints are constructed based on the decision state vector; the state transition incentives are constructed based on multiple optimization objectives; the multiple optimization objectives include operating cost optimization, carbon emission optimization, and command response optimization; according to the action space constraints, the decision action vector is trimmed to obtain a feasible action vector, and the feasible action vector is sent to multiple user-side devices to execute the decision. The implementation of this application can improve the accuracy of user-side multi-objective decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent scheduling and decision-making, and in particular to a multi-objective decision-making execution method and system oriented towards the user side. Background Technology

[0002] With the large-scale integration of distributed renewable energy sources at the user side, several new dimensions of problems have been added to the power system's decision-making and dispatching at the user side, such as economic operation and carbon emission reduction, posing new challenges. Currently, there are two main decision-making methods for the power system at the user side: those based on traditional control strategies and those based on intelligent agents. Traditional control strategy-based methods rely on local user-side information for monitoring, lacking system-level collaborative information. Their decisions are primarily based on single-objective optimization, making it difficult to consider the added problem dimensions, resulting in the decision-making and dispatching accuracy failing to meet target requirements after the integration of distributed renewable energy at the user side. Intelligent agent-based methods can handle multi-objective optimization and thus address the added problem dimensions, but the complexity of the agents they rely on is high, and they fail to consider the user-side response to power system commands, leaving significant room for improvement in actual decision-making and dispatching accuracy. Therefore, improving the multi-objective decision-making accuracy at the user side under the large-scale integration of distributed renewable energy remains a pressing technical problem that needs to be solved. Summary of the Invention

[0003] This application provides a user-side multi-objective decision execution method and system to solve the technical problem that the accuracy of existing user-side multi-objective decision-making does not meet the target requirements.

[0004] According to a first aspect of the embodiments of this application, a user-side multi-objective decision execution method is provided, comprising:

[0005] Collect operational status data from multiple user-side devices to be decided, and construct a decision status vector;

[0006] The decision state vector is input into a preset Markov policy network, and the output is a decision action vector; wherein, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives; the action space constraints are constructed based on each data column in the decision state vector; the state transition incentives are constructed based on multiple preset optimization objectives; the multiple optimization objectives include operating cost optimization, carbon emission optimization, and command response optimization;

[0007] Based on the action space constraints, the decision action vector is pruned to obtain a feasible action vector, and the feasible action vector is sent to the plurality of user-side devices so that the plurality of user-side devices can execute decisions based on the feasible action vector.

[0008] This application first collects the operating status data of multiple user-side devices to be decided, and then constructs a decision state vector. This vector is then input into a pre-defined Markov policy network to obtain a decision action vector, which is then pruned to obtain an actionable vector. The actionable vector is then distributed to multiple user-side devices to execute the decision. Compared to existing technologies, this application constructs action space constraints for the Markov policy network based on each data column in the decision state vector. It constructs state transition incentives for the Markov policy network based on multiple optimization objectives, including operating cost optimization, carbon emission optimization, and command response optimization. Then, it iteratively trains a feedforward neural network using action space constraints and state transition incentives to obtain the Markov policy network. By constructing multiple optimization objectives, strategies are generated from various decision dimensions, improving the comprehensiveness of the decision and thus increasing the decision accuracy for user-side devices. By constructing optimization objectives for command response optimization, the application can fully reflect the user-side's responsiveness to power grid decision commands. Furthermore, by deeply optimizing decision actions based on these optimization objectives, the accuracy of decision actions is improved, thereby enhancing the decision accuracy for user-side devices.

[0009] In some embodiments of this application, the step of collecting operational status data of multiple user-side devices to be decided and constructing a decision state vector specifically includes:

[0010] The system collects operational status data from the multiple user-side devices; wherein the operational status data includes local measurement data from the multiple user-side devices and group perception information obtained by monitoring the multiple user-side devices through the power grid-side system.

[0011] By integrating the local measurement data from the multiple user-side devices and the group perception information, a decision state vector is constructed.

[0012] This application first collects operational status data from multiple user-side devices, including local measurement data and collective perception information obtained through monitoring user-side devices via the power grid system. Then, it integrates these data to construct a decision-making state vector. By combining user-side local measurement and power grid monitoring, the constructed decision-making state vector is better matched to the current status of the user-side devices, thereby improving monitoring accuracy.

[0013] In some embodiments of this application, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition stimuli, specifically including:

[0014] Based on the device types of the multiple user-side devices and combined with the data columns of the decision state vector, action space constraints are constructed and determined;

[0015] Based on the multiple optimization objectives and the decision state vector, state transition incentives are constructed and determined.

[0016] Based on the decision state vector, the model parameters of the feedforward neural network are set, and combined with the action space constraints and the state transition incentives, the feedforward neural network is iteratively trained to obtain a Markov policy network.

[0017] This application first constructs a definite action space constraint based on the device type of the user-side device and the data columns of the decision state vector. Then, based on multiple optimization objectives and the decision state vector, it constructs a definite state transition incentive. Subsequently, it sets the model parameters of the feedforward neural network and performs iterative training to obtain a Markov policy network. By constructing action space constraints through the device type of the user-side device, it accurately matches the potential action states of the user-side device. At the same time, it determines the state transition incentive through multiple different optimization objective dimensions, comprehensively making decisions for the user-side device, thereby improving the decision accuracy for user-side devices.

[0018] In some embodiments of this application, the step of constructing and determining action space constraints based on the device types of the plurality of user-side devices and in combination with the data columns of the decision state vector specifically includes:

[0019] Based on the device types of the multiple user-side devices, determine the executable action vectors of the Markov policy network;

[0020] Based on each data column of the decision state vector, the action space of the Markov policy network is determined;

[0021] Based on the executable action vectors and the action space, action space constraints are constructed.

[0022] This application first determines the executable action vector based on the device type of the user-side device, and then determines the action space based on each data column of the decision state vector, accurately matching the potential action states of the user-side device, thereby improving the accuracy of the constructed action space constraints and thus improving the decision accuracy when making decisions for the user-side device.

[0023] In some embodiments of this application, the step of constructing and determining state transition stimuli based on the plurality of optimization objectives and the decision state vector specifically includes:

[0024] Based on the operating cost optimization objective, and combining the system electricity price component and the instruction response component in the decision state vector, the operating cost optimization incentive is determined.

[0025] Based on the carbon emission optimization objective, and combined with the system marginal carbon emission intensity component in the decision state vector, the carbon emission optimization incentive is determined.

[0026] Based on the instruction response optimization objective, and in conjunction with the instruction response components in the decision state vector, the instruction response optimization incentive is determined.

[0027] The normalized weighted sum of the operating cost optimization incentive, the carbon emission optimization incentive, and the instruction response optimization incentive is used as the state transition incentive of the Markov policy network.

[0028] This application first determines the corresponding operating cost optimization incentives, carbon emission optimization incentives, and command response optimization incentives based on the operating cost optimization objective, carbon emission optimization objective, and command response optimization objective, respectively, and combines different components in the decision state vector. Then, it normalizes and performs a weighted sum to obtain the state transition incentives. Through multiple different optimization objectives, it comprehensively makes decisions for user-side equipment, thereby improving the decision accuracy when making decisions for user-side equipment.

[0029] In some embodiments of this application, the step of inputting the decision state vector into a preset Markov policy network and outputting a decision action vector specifically includes:

[0030] Based on the control modes of the multiple user-side devices, the output mode of the decision action vector of the Markov policy network is determined; wherein, the output mode includes the output vector length of the decision action vector and the output type of each component;

[0031] The decision state vector is input into the Markov policy network to obtain the decision action vector under the output mode control.

[0032] This application first determines the output mode of the decision action vector based on the control mode of the user-side device, and then obtains the decision action vector output by the decision state vector through the Markov policy network under the control of the output mode. By determining the output mode of the decision action vector through the control mode of the user-side device, the decision action vector can be better matched to the control requirements of the current user-side device, thereby improving the decision accuracy when making decisions for the user-side device.

[0033] In some embodiments of this application, the step of pruning the decision action vector according to the action space constraints to obtain an actionable action vector specifically includes:

[0034] The decision action vector is denormalized to obtain the original action vector;

[0035] The original action vector is clipped, and the clipping is performed based on the action boundary constructed by the action space constraints to obtain a movable action vector.

[0036] This application first performs inverse normalization on the decision action vector to obtain the original action vector, and then performs pruning based on the action space constraints to obtain the actionable action vector, which can ensure the executability of the decision and avoid decision execution failure.

[0037] According to a second aspect of the embodiments of this application, a multi-objective decision execution system for the user side is provided, including a device data acquisition module, a decision action generation module, and a decision action execution module;

[0038] The device data acquisition module is used to collect the operating status data of multiple user-side devices to be decided, and to construct a decision status vector.

[0039] The decision action generation module is used to input the decision state vector into a preset Markov policy network and output a decision action vector; wherein, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives; the action space constraints are constructed based on each data column in the decision state vector; the state transition incentives are constructed based on multiple preset optimization objectives; the multiple optimization objectives include operating cost optimization, carbon emission optimization, and command response optimization;

[0040] The decision action execution module is used to prune the decision action vector according to the action space constraints to obtain a feasible action vector, and send the feasible action vector to the plurality of user-side devices so that the plurality of user-side devices can execute decisions according to the feasible action vector.

[0041] In some embodiments of this application, the device data acquisition module includes a data acquisition unit and a data fusion unit;

[0042] The data acquisition unit is used to collect the operating status data of the multiple user-side devices; wherein, the operating status data includes local measurement data of the multiple user-side devices and group perception information obtained by monitoring the multiple user-side devices through the power grid side system;

[0043] The data fusion unit is used to fuse local measurement data from multiple user-side devices and the group perception information to construct a decision state vector.

[0044] In some embodiments of this application, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition stimuli, specifically including:

[0045] Based on the device types of the multiple user-side devices and combined with the data columns of the decision state vector, action space constraints are constructed and determined;

[0046] Based on the multiple optimization objectives and the decision state vector, state transition incentives are constructed and determined.

[0047] Based on the decision state vector, the model parameters of the feedforward neural network are set, and combined with the action space constraints and the state transition incentives, the feedforward neural network is iteratively trained to obtain a Markov policy network.

[0048] In some embodiments of this application, the step of constructing and determining action space constraints based on the device types of the plurality of user-side devices and in combination with the data columns of the decision state vector specifically includes:

[0049] Based on the device types of the multiple user-side devices, determine the executable action vectors of the Markov policy network;

[0050] Based on each data column of the decision state vector, the action space of the Markov policy network is determined;

[0051] Based on the executable action vectors and the action space, action space constraints are constructed.

[0052] In some embodiments of this application, the step of constructing and determining state transition stimuli based on the plurality of optimization objectives and the decision state vector specifically includes:

[0053] Based on the operating cost optimization objective, and combining the system electricity price component and the instruction response component in the decision state vector, the operating cost optimization incentive is determined.

[0054] Based on the carbon emission optimization objective, and combined with the system marginal carbon emission intensity component in the decision state vector, the carbon emission optimization incentive is determined.

[0055] Based on the instruction response optimization objective, and in conjunction with the instruction response components in the decision state vector, the instruction response optimization incentive is determined.

[0056] The normalized weighted sum of the operating cost optimization incentive, the carbon emission optimization incentive, and the instruction response optimization incentive is used as the state transition incentive of the Markov policy network.

[0057] In some embodiments of this application, the decision action generation module includes a pattern determination unit and an action generation unit;

[0058] The mode determination unit is used to determine the output mode of the decision action vector of the Markov policy network according to the control modes of the plurality of user-side devices; wherein, the output mode includes the output vector length of the decision action vector and the output type of each component;

[0059] The action generation unit is used to input the decision state vector into the Markov policy network to obtain the decision action vector under the output mode control.

[0060] In some embodiments of this application, the decision-making action execution module includes an inverse normalization unit and a vector clipping unit;

[0061] The inverse normalization unit is used to inverse normalize the decision action vector to obtain the original action vector.

[0062] The vector clipping unit is used to clip the original action vector. During clipping, the action boundary constructed based on the action space constraints is clipped to obtain an actionable action vector.

[0063] This application first collects the operating status data of multiple user-side devices to be decided, and then constructs a decision state vector. This vector is then input into a pre-defined Markov policy network to obtain a decision action vector, which is then pruned to obtain an actionable vector. The actionable vector is then distributed to multiple user-side devices to execute the decision. Compared to existing technologies, this application constructs action space constraints for the Markov policy network based on each data column in the decision state vector. It constructs state transition incentives for the Markov policy network based on multiple optimization objectives, including operating cost optimization, carbon emission optimization, and command response optimization. Then, it iteratively trains a feedforward neural network using action space constraints and state transition incentives to obtain the Markov policy network. By constructing multiple optimization objectives, strategies are generated from various decision dimensions, improving the comprehensiveness of the decision and thus increasing the decision accuracy for user-side devices. By constructing optimization objectives for command response optimization, the application can fully reflect the user-side's responsiveness to power grid decision commands. Furthermore, by deeply optimizing decision actions based on these optimization objectives, the accuracy of decision actions is improved, thereby enhancing the decision accuracy for user-side devices.

[0064] According to a third aspect of the embodiments of this application, a computer device is provided, comprising: a processor; a memory; and a computer program stored in the memory and configured to be executed by the processor; wherein the processor executes the computer program to implement a user-side multi-objective decision-making method as described in this application.

[0065] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute a user-side multi-objective decision-making method as described in this application. Attached Figure Description

[0066] Figure 1 This is a flowchart illustrating a user-side multi-objective decision-making execution method according to certain embodiments of this application.

[0067] Figure 2 This is a block diagram of a user-oriented multi-objective decision execution system according to certain embodiments of this application. Detailed Implementation

[0068] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below in conjunction with the accompanying drawings are exemplary and are only used to explain some embodiments of this application, and should not be construed as limiting the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments shown in this application without inventive effort are within the protection scope of this application.

[0069] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, unless otherwise explicitly specified, "a plurality of" or "several" means two or more.

[0070] Currently, power system decision-making and dispatching at the user side mainly employs two methods: traditional control strategies and agent-based intelligent decision-making. Traditional control strategies rely on local user-side monitoring, lacking system-level collaborative information. Their decisions are primarily based on single-objective optimization, failing to account for new problem dimensions, resulting in dispatch accuracy falling short of target requirements after distributed renewable energy is integrated into the user side. Agent-based intelligent decision-making can handle multi-objective optimization, thus addressing new problem dimensions. However, the agents they rely on are highly complex and fail to consider user-side responses to power system commands, leaving significant room for improvement in actual dispatch accuracy. Therefore, improving the accuracy of multi-objective decision-making at the user side remains a critical technical challenge for current technologies in the context of large-scale distributed renewable energy integration.

[0071] Based on the above technical background, please refer to Figure 1 This application provides a user-side multi-objective decision-making execution method, including steps S101 to S103, each step as follows:

[0072] Step S101: Collect the operating status data of multiple user-side devices to be decided, and construct the decision status vector.

[0073] In some embodiments of this application, the step of collecting operational status data of multiple user-side devices to be decided and constructing a decision state vector specifically includes:

[0074] The system collects operational status data from the multiple user-side devices; wherein the operational status data includes local measurement data from the multiple user-side devices and group perception information obtained by monitoring the multiple user-side devices through the power grid-side system.

[0075] By integrating the local measurement data from the multiple user-side devices and the group perception information, a decision state vector is constructed.

[0076] Specifically, the multiple user-side devices can be grouped into a device set. User-side equipment types include distributed photovoltaics, battery storage, electric vehicles, and flexible loads, corresponding to the equipment sets. The subsets can be represented as distributable photovoltaic sets. Battery energy storage collection Electric vehicle collection and flexible load collection Among them, each user-side device Belongs to the distribution network bus set The corresponding unique node in .

[0077] Specifically, during the scheduling period ( (This represents the discrete scheduling step size, preferably 15 minutes or 1 hour), and the local measurement data of the multiple user-side devices. This includes: equipment obtained through irradiance prediction or actual measurement. Photovoltaic power available output ;node Energy storage state of charge Electric vehicle cluster aggregation available adjustable capacity ,in For electric vehicles During the scheduling period Adjustable power margin; flexible load operation status Real-time local load Local node voltage amplitude .

[0078] Specifically, during the scheduling period The group perception information This is a group status quo sent from the upper aggregation and coordination layer (i.e., the grid side, including virtual power plants or distribution network dispatch centers) to the user side via a low-frequency broadcast channel every 15-60 minutes. It characterizes the overall system operation status obtained from monitoring multiple user-side devices, including: ancillary service request commands (i.e. command responses). , respectively representing the need for down-regulation, no demand, and the need for up-regulation; system marginal carbon emission intensity System electricity price Key feeder voltage over-limit warning sign .

[0079] Specifically, during the scheduling period The decision state vector obtained by fusing local measurement data and group perception information ,in For the fusion state dimension.

[0080] This application first collects operational status data from multiple user-side devices, including local measurement data and collective perception information obtained through monitoring user-side devices via the power grid system. Then, it integrates these data to construct a decision-making state vector. By combining user-side local measurement and power grid monitoring, the constructed decision-making state vector is better matched to the current status of the user-side devices, thereby improving monitoring accuracy.

[0081] Step S102: Input the decision state vector into a preset Markov policy network and output a decision action vector; wherein, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives; the action space constraints are constructed based on each data column in the decision state vector; the state transition incentives are constructed based on multiple preset optimization objectives; the multiple optimization objectives include operating cost optimization, carbon emission optimization, and command response optimization.

[0082] In some embodiments of this application, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition stimuli, specifically including:

[0083] Based on the device types of the multiple user-side devices and combined with the data columns of the decision state vector, action space constraints are constructed and determined;

[0084] Based on the multiple optimization objectives and the decision state vector, state transition incentives are constructed and determined.

[0085] Based on the decision state vector, the model parameters of the feedforward neural network are set, and combined with the action space constraints and the state transition incentives, the feedforward neural network is iteratively trained to obtain a Markov policy network.

[0086] In some embodiments of this application, the feedforward neural network includes an input layer, a hidden layer, and an output layer; the input layer receives a decision state vector as input; the hidden layer has 128 neurons, uses ReLU as the activation function, has a hidden weight matrix of size 128×16, and a bias vector of length 128; the output layer uses tanh as the activation function, mapping the original output to... .

[0087] This application first constructs a definite action space constraint based on the device type of the user-side device and the data columns of the decision state vector. Then, based on multiple optimization objectives and the decision state vector, it constructs a definite state transition incentive. Subsequently, it sets the model parameters of the feedforward neural network and performs iterative training to obtain a Markov policy network. By constructing action space constraints through the device type of the user-side device, it accurately matches the potential action states of the user-side device. At the same time, it determines the state transition incentive through multiple different optimization objective dimensions, comprehensively making decisions for the user-side device, thereby improving the decision accuracy for user-side devices.

[0088] In some embodiments of this application, the step of constructing and determining action space constraints based on the device types of the plurality of user-side devices and in combination with the data columns of the decision state vector specifically includes:

[0089] Based on the device types of the multiple user-side devices, determine the executable action vectors of the Markov policy network;

[0090] Based on each data column of the decision state vector, the action space of the Markov policy network is determined;

[0091] Based on the executable action vectors and the action space, action space constraints are constructed.

[0092] Specifically, during the scheduling period The executable action vector , of which components Indicates the energy storage charging and discharging power. Indicates the net adjustable power of electric vehicles. This indicates that flexible loads can reduce power consumption. More specifically, the operating space... The corresponding action space constraints are as follows:

[0093] (1) Photovoltaic operation constraints: Then the feasible range of total photovoltaic power output ;in For scheduling period Photovoltaics Actual output; For scheduling period Photovoltaics Available output; For the actual total output of photovoltaic power, Total feasible output for photovoltaic power;

[0094] To ensure Local net power transmission The feasible constraint interval is ;

[0095] Combined with the fact that the voltage amplitude at the local node is not less than the upper limit of the distribution network voltage Time requirements The voltage constraint can determine the two operational boundaries of photovoltaic operation:

[0096] Preventing negative total power consumption ;

[0097] To prevent photovoltaic power from being unable to support static power transmission ;in For real-time local load;

[0098] (2) Constraints on energy storage operation: Energy storage charging and discharging power components ;in For scheduling period Energy storage unit The charging and discharging power; For scheduling period Energy storage unit The actual maximum allowable charging power; For charging efficiency; This refers to the maximum allowable charging power of the energy storage unit. For scheduling period Energy storage unit The energy storage state of charge; For energy storage units Rated capacity; The sampling interval; For scheduling period Energy storage unit The actual maximum permissible discharge power; For discharge efficiency, This represents the maximum allowable discharge efficiency of the energy storage unit.

[0099] Based on the actual maximum allowable charge and discharge power of the energy storage unit, the evolution of the energy storage state of charge must satisfy... and always maintain ,in These are the minimum and maximum power limits for the energy storage state of charge, respectively, and the power limit for the energy storage state of charge varies with the current... Dynamic changes constitute the operational boundaries of energy storage;

[0100] (3) Electric vehicles and flexible load constraints: Net regulating power component of electric vehicles Flexible loads can reduce power components. ;in For scheduling period electric vehicles Net regulation power, For scheduling period Flexible load The power that can be reduced;

[0101] Electric vehicles need to meet the total energy demand of a single vehicle. ;in Basic charging power; The sampling interval; Minimum daily charging energy;

[0102] Flexible loads must meet continuous operation / interruption time constraints, specifically:

[0103] ;in For scheduling time slots Flexible load The power that can be reduced For indicator functions, when the indicator function contains The value is 1 if the condition is met, and 0 otherwise. Flexible loads Runtime and interrupt duration;

[0104] Flexible loads also need to meet the daily interruption constraint. ,in For flexible loads The maximum number of daily interruptions.

[0105] This application first determines the executable action vector based on the device type of the user-side device, and then determines the action space based on each data column of the decision state vector, accurately matching the potential action states of the user-side device, thereby improving the accuracy of the constructed action space constraints and thus improving the decision accuracy when making decisions for the user-side device.

[0106] In some embodiments of this application, the step of constructing and determining state transition stimuli based on the plurality of optimization objectives and the decision state vector specifically includes:

[0107] Based on the operating cost optimization objective, and combining the system electricity price component and the instruction response component in the decision state vector, the operating cost optimization incentive is determined.

[0108] Based on the carbon emission optimization objective, and combined with the system marginal carbon emission intensity component in the decision state vector, the carbon emission optimization incentive is determined.

[0109] Based on the instruction response optimization objective, and in conjunction with the instruction response components in the decision state vector, the instruction response optimization incentive is determined.

[0110] The normalized weighted sum of the operating cost optimization incentive, the carbon emission optimization incentive, and the instruction response optimization incentive is used as the state transition incentive of the Markov policy network.

[0111] Specifically, the operating cost optimization incentive ;in For system electricity price; This represents the local net power output. Fixed stimulus for instruction response; As an indicator function, when the action Commands to meet ancillary service requirements If the value is 1, then the value is 0; otherwise, the value is 0.

[0112] Specifically, the carbon emission optimization incentives ;in This represents the marginal carbon emission intensity of the system.

[0113] Specifically, the instruction response optimization incentive ;in The response excitation coefficient.

[0114] Specifically, the normalization includes standard normalization. ( For various incentives (mean and standard deviation) and maximum and minimum normalization ( For various incentives (Minimum and maximum values), normalized range is .

[0115] Specifically, the state transition incentive ;in For weighted weights, and .

[0116] This application first determines the corresponding operating cost optimization incentives, carbon emission optimization incentives, and command response optimization incentives based on the operating cost optimization objective, carbon emission optimization objective, and command response optimization objective, respectively, and combines different components in the decision state vector. Then, it normalizes and performs a weighted sum to obtain the state transition incentives. Through multiple different optimization objectives, it comprehensively makes decisions for user-side equipment, thereby improving the decision accuracy when making decisions for user-side equipment.

[0117] In some embodiments of this application, the step of inputting the decision state vector into a preset Markov policy network and outputting a decision action vector specifically includes:

[0118] Based on the control modes of the multiple user-side devices, the output mode of the decision action vector of the Markov policy network is determined; wherein, the output mode includes the output vector length of the decision action vector and the output type of each component;

[0119] The decision state vector is input into the Markov policy network to obtain the decision action vector under the output mode control.

[0120] Specifically, when the control mode is discrete control, the output vector length of the decision action vector is determined to be 5, and the output type of each component is an action type, corresponding to 5 predefined action combinations; when the control mode is continuous control, the output vector length of the decision action vector is determined to be 3, and the output type of each component corresponds to the energy storage charging and discharging power, the electric vehicle net regulation power, and the flexible load reduction power, respectively.

[0121] This application first determines the output mode of the decision action vector based on the control mode of the user-side device, and then obtains the decision action vector output by the decision state vector through the Markov policy network under the control of the output mode. By determining the output mode of the decision action vector through the control mode of the user-side device, the decision action vector can be better matched to the control requirements of the current user-side device, thereby improving the decision accuracy when making decisions for the user-side device.

[0122] Step S103: Based on the action space constraints, the decision action vector is pruned to obtain a feasible action vector, and the feasible action vector is sent to the plurality of user-side devices so that the plurality of user-side devices can execute decisions based on the feasible action vector.

[0123] In some embodiments of this application, the step of pruning the decision action vector according to the action space constraints to obtain an actionable action vector specifically includes:

[0124] The decision action vector is denormalized to obtain the original action vector;

[0125] The original action vector is clipped, and the clipping is performed based on the action boundary constructed by the action space constraints to obtain a movable action vector.

[0126] This application first performs inverse normalization on the decision action vector to obtain the original action vector, and then performs pruning based on the action space constraints to obtain the actionable action vector, which can ensure the executability of the decision and avoid decision execution failure.

[0127] Compared to existing technologies, this application first collects the operating status data of multiple user-side devices to be decided, and then constructs a decision state vector. This vector is then input into a pre-defined Markov policy network to obtain a decision action vector, which is then pruned to obtain an actionable vector. The actionable vector is then distributed to multiple user-side devices to execute the decision. Compared to existing technologies, this application constructs action space constraints for the Markov policy network based on each data column in the decision state vector. It constructs state transition incentives for the Markov policy network based on multiple optimization objectives, including operating cost optimization, carbon emission optimization, and command response optimization. Then, it iteratively trains a feedforward neural network using action space constraints and state transition incentives to obtain the Markov policy network. By constructing multiple optimization objectives, strategies are generated from various decision dimensions, improving the comprehensiveness of the decision and thus increasing the decision accuracy for user-side devices. By constructing optimization objectives for command response optimization, the responsiveness of the user side to power grid decision commands can be fully reflected. Furthermore, the decision actions are deeply optimized based on these optimization objectives, improving the accuracy of the decision actions and thus increasing the decision accuracy for user-side devices.

[0128] For a method corresponding to the one described above, please refer to [link to relevant documentation]. Figure 2 This application provides a user-side multi-objective decision execution system, including a device data acquisition module 210, a decision action generation module 220, and a decision action execution module 230;

[0129] The device data acquisition module 210 is used to collect the operating status data of multiple user-side devices to be decided, and to construct a decision status vector.

[0130] The decision action generation module 220 is used to input the decision state vector into a preset Markov policy network and output a decision action vector; wherein, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives; the action space constraints are constructed based on each data column in the decision state vector; the state transition incentives are constructed based on multiple preset optimization objectives; the multiple optimization objectives include operating cost optimization, carbon emission optimization, and command response optimization;

[0131] The decision action execution module 230 is used to trim the decision action vector according to the action space constraints to obtain an actionable action vector, and send the actionable action vector to the plurality of user-side devices so that the plurality of user-side devices can execute decisions according to the actionable action vector.

[0132] In some embodiments of this application, the device data acquisition module 210 includes a data acquisition unit and a data fusion unit;

[0133] The data acquisition unit is used to collect the operating status data of the multiple user-side devices; wherein, the operating status data includes local measurement data of the multiple user-side devices and group perception information obtained by monitoring the multiple user-side devices through the power grid side system;

[0134] The data fusion unit is used to fuse local measurement data from multiple user-side devices and the group perception information to construct a decision state vector.

[0135] In some embodiments of this application, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition stimuli, specifically including:

[0136] Based on the device types of the multiple user-side devices and combined with the data columns of the decision state vector, action space constraints are constructed and determined;

[0137] Based on the multiple optimization objectives and the decision state vector, state transition incentives are constructed and determined.

[0138] Based on the decision state vector, the model parameters of the feedforward neural network are set, and combined with the action space constraints and the state transition incentives, the feedforward neural network is iteratively trained to obtain a Markov policy network.

[0139] In some embodiments of this application, the step of constructing and determining action space constraints based on the device types of the plurality of user-side devices and in combination with the data columns of the decision state vector specifically includes:

[0140] Based on the device types of the multiple user-side devices, determine the executable action vectors of the Markov policy network;

[0141] Based on each data column of the decision state vector, the action space of the Markov policy network is determined;

[0142] Based on the executable action vectors and the action space, action space constraints are constructed.

[0143] In some embodiments of this application, the step of constructing and determining state transition stimuli based on the plurality of optimization objectives and the decision state vector specifically includes:

[0144] Based on the operating cost optimization objective, and combining the system electricity price component and the instruction response component in the decision state vector, the operating cost optimization incentive is determined.

[0145] Based on the carbon emission optimization objective, and combined with the system marginal carbon emission intensity component in the decision state vector, the carbon emission optimization incentive is determined.

[0146] Based on the instruction response optimization objective, and in conjunction with the instruction response components in the decision state vector, the instruction response optimization incentive is determined.

[0147] The normalized weighted sum of the operating cost optimization incentive, the carbon emission optimization incentive, and the instruction response optimization incentive is used as the state transition incentive of the Markov policy network.

[0148] In some embodiments of this application, the decision action generation module 220 includes a pattern determination unit and an action generation unit;

[0149] The mode determination unit is used to determine the output mode of the decision action vector of the Markov policy network according to the control modes of the plurality of user-side devices; wherein, the output mode includes the output vector length of the decision action vector and the output type of each component;

[0150] The action generation unit is used to input the decision state vector into the Markov policy network to obtain the decision action vector under the output mode control.

[0151] In some embodiments of this application, the decision action execution module 230 includes an inverse normalization unit and a vector clipping unit;

[0152] The inverse normalization unit is used to inverse normalize the decision action vector to obtain the original action vector.

[0153] The vector clipping unit is used to clip the original action vector. During clipping, the action boundary constructed based on the action space constraints is clipped to obtain an actionable action vector.

[0154] This application first collects the operating status data of multiple user-side devices to be decided, and then constructs a decision state vector. This vector is then input into a pre-defined Markov policy network to obtain a decision action vector, which is then pruned to obtain an actionable vector. The actionable vector is then distributed to multiple user-side devices to execute the decision. Compared to existing technologies, this application constructs action space constraints for the Markov policy network based on each data column in the decision state vector. It constructs state transition incentives for the Markov policy network based on multiple optimization objectives, including operating cost optimization, carbon emission optimization, and command response optimization. Then, it iteratively trains a feedforward neural network using action space constraints and state transition incentives to obtain the Markov policy network. By constructing multiple optimization objectives, strategies are generated from various decision dimensions, improving the comprehensiveness of the decision and thus increasing the decision accuracy for user-side devices. By constructing optimization objectives for command response optimization, the application can fully reflect the user-side's responsiveness to power grid decision commands. Furthermore, by deeply optimizing decision actions based on these optimization objectives, the accuracy of decision actions is improved, thereby enhancing the decision accuracy for user-side devices.

[0155] It should be understood that the system provided in this application is corresponding to the aforementioned method. The user-side multi-objective decision execution system provided in this application can implement the user-side multi-objective decision execution method provided in any embodiment of this application.

[0156] Adaptively, embodiments of this application also provide a computer device and a computer-readable storage medium.

[0157] The computer device includes: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor;

[0158] The processor executes the computer program to implement a user-side multi-objective decision execution method according to this application.

[0159] The computer-readable storage medium stores multiple instructions that are adapted for a processor to load and execute a user-oriented multi-objective decision execution method according to this application.

[0160] The above description represents some embodiments of this application, providing a further detailed explanation of the purpose, technical solution, and beneficial effects of this application. It should be understood that the above-described embodiments of this application should not be construed as limiting this application. In particular, any changes, modifications, equivalent substitutions, and variations made by those skilled in the art within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A user-side multi-objective decision-making execution method, characterized in that, include: Operational status data of multiple user-side devices to be decided are collected to construct a decision state vector. The operational status data includes local measurement data of the multiple user-side devices and group perception information obtained through monitoring of the multiple user-side devices by the grid-side system. The local measurement data includes the available photovoltaic output and energy storage state of charge of each device, as well as the aggregated available adjustable capacity of the electric vehicle cluster, flexible load operation status, real-time local load, and local node voltage amplitude. The group perception information includes ancillary service demand instructions, system marginal carbon emission intensity, system electricity price, and critical feeder voltage over-limit warning indicators. The decision state vector is input into a preset Markov policy network, and the output is a decision action vector; wherein, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives; the action space constraints are constructed based on each data column in the decision state vector; the state transition incentives are constructed based on multiple preset optimization objectives; the multiple optimization objectives include operating cost optimization, carbon emission optimization, and command response optimization; The decision action vector is denormalized to obtain the original action vector. The original action vector is then pruned, and the pruning is performed on the action boundary constructed based on the action space constraints to obtain the actionable action vector. The actionable action vector is then sent to the multiple user-side devices so that the multiple user-side devices can execute decisions based on the actionable action vector.

2. The user-side multi-objective decision execution method according to claim 1, characterized in that, The process of collecting operational status data from multiple user-side devices to be decided upon, and constructing a decision status vector, specifically includes: Collect operational status data from the multiple user-side devices; By integrating the local measurement data from the multiple user-side devices and the group perception information, a decision state vector is constructed.

3. The user-side multi-objective decision execution method according to claim 2, characterized in that, The Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives, specifically including: Based on the device types of the multiple user-side devices and combined with the data columns of the decision state vector, action space constraints are constructed and determined; Based on the multiple optimization objectives and the decision state vector, state transition incentives are constructed and determined. Based on the decision state vector, the model parameters of the feedforward neural network are set, and combined with the action space constraints and the state transition incentives, the feedforward neural network is iteratively trained to obtain a Markov policy network; wherein, the decision state vector is obtained by fusing the local measurement data and the group perception information, and the decision state vector satisfies a preset fused state dimension; the feedforward neural network includes an input layer, which receives the decision state vector as input.

4. The user-side multi-objective decision execution method according to claim 3, characterized in that, The step of constructing and determining action space constraints based on the device types of the multiple user-side devices and in conjunction with the data columns of the decision state vector specifically includes: Based on the device types of the multiple user-side devices, the executable action vector of the Markov policy network is determined; wherein, the device types include distributed photovoltaic, battery energy storage, electric vehicles and flexible loads; the executable action vector includes a component representing the charging and discharging power of energy storage, a component representing the net regulation power of electric vehicles and a component representing the power that can be reduced by flexible loads; Based on each data column of the decision state vector, the action space of the Markov policy network is determined; Based on the executable action vectors and the action space, action space constraints are constructed.

5. The user-side multi-objective decision execution method according to claim 3, characterized in that, The process of constructing and determining state transition stimuli based on the multiple optimization objectives and the decision state vector specifically includes: Based on the operating cost optimization objective, and combining the system electricity price component and the instruction response component in the decision state vector, the operating cost optimization incentive is determined. Based on the carbon emission optimization objective, and combined with the system marginal carbon emission intensity component in the decision state vector, the carbon emission optimization incentive is determined. Based on the instruction response optimization objective, and in conjunction with the instruction response components in the decision state vector, the instruction response optimization incentive is determined. The normalized weighted sum of the operating cost optimization incentive, the carbon emission optimization incentive, and the instruction response optimization incentive is used as the state transition incentive of the Markov policy network.

6. The user-side multi-objective decision execution method according to claim 1, characterized in that, The step of inputting the decision state vector into a preset Markov policy network and outputting a decision action vector specifically includes: Based on the control modes of the multiple user-side devices, the output mode of the decision action vector of the Markov policy network is determined; wherein, the output mode includes the output vector length of the decision action vector and the output type of each component; The decision state vector is input into the Markov policy network to obtain the decision action vector under the output mode control.

7. A user-oriented multi-objective decision execution system, characterized in that, It includes a device data acquisition module, a decision action generation module, and a decision action execution module; The equipment data acquisition module is used to collect the operating status data of multiple user-side devices to be decided, and construct a decision state vector. The operating status data includes local measurement data of the multiple user-side devices and group perception information obtained through monitoring of the multiple user-side devices by the grid-side system. The local measurement data includes the available photovoltaic output and energy storage charge status of each device, as well as the aggregated available adjustable capacity of the electric vehicle cluster, flexible load operating status, real-time local load, and local node voltage amplitude. The group perception information includes ancillary service demand instructions, system marginal carbon emission intensity, system electricity price, and critical feeder voltage over-limit warning indicators. The decision action generation module is used to input the decision state vector into a preset Markov policy network and output a decision action vector; wherein, the Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives; the action space constraints are constructed based on each data column in the decision state vector; the state transition incentives are constructed based on multiple preset optimization objectives; the multiple optimization objectives include operating cost optimization, carbon emission optimization, and command response optimization; The decision action execution module is used to inverse normalize the decision action vector to obtain an original action vector, prune the original action vector by pruning the action boundary constructed based on the action space constraints to obtain a feasible action vector, and send the feasible action vector to the multiple user-side devices so that the multiple user-side devices can execute decisions based on the feasible action vector.

8. A user-oriented multi-objective decision execution system according to claim 7, characterized in that, The Markov policy network is obtained by iteratively training a feedforward neural network based on action space constraints and state transition incentives, specifically including: Based on the device types of the multiple user-side devices and combined with the data columns of the decision state vector, action space constraints are constructed and determined; Based on the multiple optimization objectives and the decision state vector, state transition incentives are constructed and determined. Based on the decision state vector, the model parameters of the feedforward neural network are set, and combined with the action space constraints and the state transition incentives, the feedforward neural network is iteratively trained to obtain a Markov policy network; wherein, the decision state vector is obtained by fusing local measurement data and group perception information, and the decision state vector satisfies a preset fused state dimension; the feedforward neural network includes an input layer, which receives the decision state vector as input.

9. A user-oriented multi-objective decision execution system according to claim 8, characterized in that, The step of constructing and determining action space constraints based on the device types of the multiple user-side devices and in conjunction with the data columns of the decision state vector specifically includes: Based on the device types of the multiple user-side devices, the executable action vector of the Markov policy network is determined; wherein, the device types include distributed photovoltaic, battery energy storage, electric vehicles and flexible loads; the executable action vector includes a component representing the charging and discharging power of energy storage, a component representing the net regulation power of electric vehicles and a component representing the power that can be reduced by flexible loads; Based on each data column of the decision state vector, the action space of the Markov policy network is determined; Based on the executable action vectors and the action space, action space constraints are constructed.

Citation Information

Patent Citations

  • Power distribution network operation optimization method and device, equipment and storage medium

    CN120728558A

  • Power grid energy storage demand response multi-agent reinforcement learning dynamic optimization method

    CN121367205A