Multi-agent-based receiving end power system optimization operation method and device

Through the optimized operation method of the receiving power system based on multi-agents, the problem that the existing technology is difficult to cope with complex situations in multiple regions and multiple nodes is solved, the stability and economics of the power system are improved, and the power scheduling is optimized and the grid operation cost is reduced.

CN119994903AActive Publication Date: 2025-05-13CHINA AGRI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510468427.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the complex situation of multiple regions and multiple nodes, and it is difficult to take into account both parameter selection and system optimization, which greatly affects the optimization effect.

Method used

The optimization operation method of the receiving power system based on multi-agents is adopted. By determining the circuit topology of the power system, the node and edge features are extracted, the region is divided using spectral clustering algorithm, the power system power balance model and the scenery prediction model are constructed, and the operation operation of the power system is optimized by combining the dual-depth deterministic strategy gradient optimization model.

Benefits of technology

Effectively respond to the complex situations of multiple regions and multiple nodes, improve the stability and economy of the power system, optimize power scheduling, and reduce power operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119994903A_ABST
    Figure CN119994903A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric power systems, in particular to a receiving end electric power system optimization operation method and device based on multiple agents, and the method comprises the steps: carrying out the region division of an electric power system, and enabling each region node in the region division result of the electric power system to meet a power balance constraint condition; establishing a wind and light prediction model to obtain segmented prediction data; based on the segmented prediction data and a pre-constructed thermal power generation and energy storage system model, establishing a dual-depth deterministic strategy gradient optimization model, and training a first depth deterministic strategy gradient agent on the basis of historical environment data to determine the optimal parameters of the agent, and enabling the second depth deterministic strategy gradient agent to interact with the environment according to the optimal parameters of the agent, and optimizing the operation of the power system. Therefore, the problems that in the prior art, the complex conditions of multiple areas and multiple nodes are difficult to effectively deal with, parameter selection and system optimization are difficult to consider at the same time, and the optimization effect is greatly influenced are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of power systems, and in particular to a method and device for optimizing the operation of a receiving-end power system based on multiple agents. Background Art

[0002] With the continuous expansion of the scale of power systems and the widespread application of renewable energy, the complexity and uncertainty of power systems have increased significantly. In order to effectively deal with the complexity and uncertainty of multi-regional nodes in power systems and improve the stability and economy of power systems, an efficient and accurate power balance optimization method for multi-regional nodes in power systems is needed to provide a scientific basis for power system operation.

[0003] In the existing technology system, traditional power system optimization methods often have difficulty in effectively dealing with complex situations in multiple regions and nodes, especially when considering the volatility of renewable energy sources such as wind power and photovoltaics and the charging and discharging strategies of energy storage systems. The limitations of traditional methods are more obvious. When deep deterministic strategy ladder algorithms are applied to power system optimization, they usually use a single optimization algorithm, which makes it difficult to take into account both parameter selection and system optimization at the same time, resulting in poor optimization results, which needs to be solved urgently. Summary of the invention

[0004] The present application provides a method and device for optimizing the operation of a receiving-end power system based on a multi-agent system, so as to solve the problems that the prior art is difficult to effectively cope with the complex situation of multiple regions and multiple nodes, and is difficult to take into account both parameter selection and system optimization at the same time, which greatly affects the optimization effect.

[0005] The first aspect of the present application provides a method for optimizing the operation of a receiving-end power system based on a multi-agent, comprising the following steps: determining a circuit topology structure corresponding to a target power system, extracting node features and edge features corresponding to the circuit topology structure, and dividing the target power system into regions based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system regional division result; based on a pre-constructed power system power balance model, making each regional node in the power system regional division result satisfy a preset power balance constraint condition, and obtaining wind speed data and irradiance data corresponding to the target power system, and performing segmented prediction processing based on a pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain the corresponding segmented power system. Prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; construct a dual-depth deterministic policy gradient optimization model based on the pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost; collect the power system historical environment data of the target power system, and train the first deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

[0006] Optionally, in one embodiment of the present application, the circuit topology structure corresponding to the target power system is determined, and the node features and edge features corresponding to the circuit topology structure are extracted, and the target power system is regionalized based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system regionalization result, including: determining the node set, edge set and time scale of the power transmission line power data of the circuit topology structure corresponding to the target power system; obtaining the node set and edge set of each layer of the circuit topology structure ... i and nodes j A set of neighboring nodes is obtained, and based on a preset learnable weight matrix, an activation function and the set of neighboring nodes, node features of each node in the circuit topology are extracted, wherein: i and j is a positive integer; based on the edge set and the preset multi-layer perceptron, the time scale of the power data of the transmission line, the node i and the node j The node features in each layer are concatenated to obtain the edge features of each edge in the circuit topology structure; the node features are calculated based on the node features and the preset Gaussian kernel function bandwidth parameters. iand the node j The similarity between them is calculated, and a corresponding similarity matrix is ​​constructed through the similarity, so as to perform regional division on the target power system according to the similarity matrix and the spectral clustering algorithm to obtain a regional division Laplace matrix; eigendecomposition is performed on the regional division Laplace matrix to obtain eigenvectors corresponding to the first q minimum eigenvalues, and K-means clustering is performed on the eigenvectors to generate the regional division result of the power system, wherein q is a positive integer.

[0007] Optionally, in one embodiment of the present application, the power balance model of the power system constructed in advance is used to make each regional node in the regional division result of the power system satisfy a preset power balance constraint, including: determining the node power flowing through each regional node in the regional division result of the power system per hour, and constructing a line power constraint based on the preset maximum power flowing through the node per hour, the minimum power flowing through the node per hour and the power flowing through the node per hour; determining the wind power generation power per hour, photovoltaic power generation power per hour, thermal power generation power per hour and energy storage system extraction power per hour corresponding to each regional node, and calculating the multi-energy total power generation per hour based on the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour and the energy storage system extraction power per hour; calculating the total line load per hour based on the node power flowing through the node per hour, and constructing the power balance constraint based on the multi-energy total power generation per hour and the total line load per hour, so that each regional node satisfies the power balance constraint.

[0008] Optionally, in one embodiment of the present application, the wind speed data and irradiance data corresponding to the target power system are obtained, and based on a pre-constructed wind-solar prediction model, and in combination with the wind speed data and the irradiance data, segmented prediction processing is performed to obtain corresponding segmented prediction data, and the corresponding wind power generation cost and photovoltaic power generation cost are calculated according to the segmented prediction data, including: calculating the mean and standard deviation of the wind speed data and the irradiance data, and calculating the wind speed preprocessing data and photovoltaic power generation preprocessing data corresponding to the wind speed data and the irradiance data according to the mean and the standard deviation; constructing the wind-solar prediction model based on a preset segmented dynamic trend decomposition prediction model, and performing segmented prediction on the wind speed preprocessing data and the photovoltaic power generation preprocessing data through the wind-solar prediction model to obtain the first k Forecast wind speed and m The predicted irradiance is: k and mis a positive integer; obtaining the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiance corresponding to the target power system, and based on the cut-in wind speed, the rated wind speed, the cut-out wind speed and the first k Forecast wind speed for the segment and calculate the k The wind power is predicted in sections, and the m The first segment is calculated based on the predicted irradiance, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance value and the maximum irradiance value. m The photovoltaic power generation power of the first section is respectively k The wind power segment prediction power and the m Calculation of photovoltaic power generation k Total wind power forecast within the hour and m The total photovoltaic power forecast within the hour and determine the corresponding k Actual wind power generation in the hour and m The actual photovoltaic power generation in the hour is based on the k The actual wind power generation capacity and k The wind power forecast total power in the hour is used to calculate the wind power forecast power error, and the wind power forecast power error is calculated by the wind power forecast total power in the hour. m The actual photovoltaic power generation in the hour and the m The photovoltaic power generation prediction power error is calculated based on the photovoltaic predicted total power within the hour; the wind power prediction power error and the photovoltaic power generation prediction power error are judged respectively whether they meet the preset error requirements, wherein when the wind power prediction power error and the photovoltaic power generation prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, power generation cost coefficient and the k The total wind power forecast within the hour is used to calculate the wind power generation cost, and the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and the m The photovoltaic power generation cost is calculated based on the total photovoltaic power forecast within the hour.

[0009] Optionally, in one embodiment of the present application, the pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost, construct a dual-depth deterministic policy gradient optimization model, including: calculating the hourly thermal power generation value of the target power system, and based on the hourly thermal power generation value, the preset hourly thermal power generation value and the hourly thermal power generation maximum value, constructing a thermal power output constraint; obtaining the first n Hourly thermal power generation and n- 1 hour thermal power generation power, and according to the n Hours of thermal power generation power, the n-The power change rate constraint of the thermal power unit is determined based on the thermal power generation power of one hour and the preset ramp rate; the thermal power generation operation and maintenance cost coefficient, the fuel cost coefficient, the startup cost coefficient and the startup state variable are determined, and the thermal power generation cost is calculated based on the thermal power generation operation and maintenance cost coefficient, the fuel cost coefficient, the startup cost coefficient, the startup state variable and the hourly thermal power generation value, and the emission constraint of the thermal power unit is constructed according to the preset emission coefficient and the maximum allowable emission amount; the thermal power generation model is constructed based on the thermal power generation cost, the emission constraint of the thermal power unit, the power change rate constraint of the thermal power unit and the above; the charging efficiency, discharge efficiency, charging power and discharge power of the target power system are obtained, and based on the charging efficiency, the discharge efficiency , the charging power and the discharging power, calculate the charging and discharging power of the energy storage system; based on the preset maximum charging power of the energy storage system, the maximum discharging power of the energy storage system, the charging efficiency and the discharging efficiency, determine the corresponding charging and discharging power constraints and charging and discharging efficiency constraints; calculate the capacity of the energy storage system through the charging and discharging power of the energy storage system, and construct the capacity constraint of the energy storage system according to the capacity of the energy storage system, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system; calculate the cost of the energy storage system based on the preset charging and discharging cost coefficient of the energy storage system, the operation and maintenance cost coefficient of the energy storage system and the charging and discharging power of the energy storage system, and construct the energy storage system model through the cost of the energy storage system, the capacity constraint of the energy storage system, the charging and discharging power constraint and the charging and discharging efficiency constraint.

[0010] Optionally, in one embodiment of the present application, the collecting of the power system historical environment data of the target power system, and training the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, including: collecting the power system historical environment data of the target power system, and determining the first state space, the first reward function and the first system historical cost corresponding to the first deep deterministic policy gradient agent; Determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, and construct a first action space through the parameters to be optimized, and train the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function and the first system historical cost, and in combination with the historical environmental data of the power system to obtain the agent target parameters; determine the second state space, the second action space, the second reward function and the second system cost corresponding to the second deep deterministic policy gradient agent, and optimize the operation of the target power system based on the second state space, the second action space, the second reward function and the second system cost, and in combination with the agent target parameters.

[0011] Optionally, in one embodiment of the present application, the k The mathematical expression of the wind power segment prediction power is:

[0012] in, Indicates the k Segment forecast wind speed; Indicates the k Segmental wind power prediction power; It represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; Indicates the temperature; Indicates the fan blade radius; Indicates the wind direction correction factor; Indicates the angle between wind direction and wind turbine orientation; , , represent the cut-in wind speed, the rated wind speed and the cut-out wind speed respectively.

[0013] The second aspect of the present application provides a receiving-end power system optimization operation device based on multi-agents, including: a regional division module, which is used to determine the circuit topology structure corresponding to the target power system, and extract the node features and edge features corresponding to the circuit topology structure, and divide the target power system into regions based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system regional division result; a segmented prediction module, which is used to make each regional node in the power system regional division result meet the preset power balance constraint conditions based on a pre-constructed power system power balance model, and obtain the wind speed data and irradiance data corresponding to the target power system, and perform segmented prediction processing based on the pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain the corresponding segmented Prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; a modeling module, which is used to construct a dual-depth deterministic policy gradient optimization model based on a pre-built thermal power generation model and an energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost; an optimization module, which is used to collect the power system historical environment data of the target power system, and train the first deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

[0014] Optionally, in one embodiment of the present application, the area division module includes: a first determination unit, used to determine the node set, edge set and time scale of the power data of the transmission line of the circuit topology structure corresponding to the target power system; an extraction unit, used to obtain the node set in each layer of the circuit topology structure; i and nodes j A set of neighboring nodes is obtained, and based on a preset learnable weight matrix, an activation function and the set of neighboring nodes, node features of each node in the circuit topology are extracted, wherein: i and j is a positive integer; a feature concatenation unit, used to concatenate the time scale of the power data of the transmission line, the node i and the node j The node features in each layer are subjected to feature concatenation operation to obtain edge features of each edge in the circuit topology structure; the first calculation unit is used to calculate the node features based on the node features and the preset Gaussian kernel function bandwidth parameters. i and the node jThe similarity between the two regions is calculated, and a corresponding similarity matrix is ​​constructed through the similarity, so as to perform regional division on the target power system according to the similarity matrix and the spectral clustering algorithm to obtain a regional division Laplace matrix; an eigendecomposition unit is used to perform eigendecomposition on the regional division Laplace matrix to obtain eigenvectors corresponding to the first q smallest eigenvalues, and perform K-means clustering on the eigenvectors to generate the regional division result of the power system, wherein q is a positive integer.

[0015] Optionally, in one embodiment of the present application, the segmented prediction module includes: a first construction unit, used to determine the node power flowing through each regional node per hour in the power system regional division result, and construct a line power constraint based on the preset maximum node power flowing through the node per hour, the minimum node power flowing through the node per hour and the node power flowing through the node per hour; a second determination unit, used to determine the wind power generation power per hour, photovoltaic power generation power per hour, thermal power generation power per hour and energy storage system extraction power per hour corresponding to each regional node, and calculate the multi-energy total power generation power per hour based on the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour and the energy storage system extraction power per hour; a second construction unit, used to calculate the total line load per hour based on the node power flowing through the node per hour, and construct the power balance constraint condition based on the multi-energy total power generation power per hour and the total line load per hour, so that each regional node satisfies the power balance constraint condition.

[0016] Optionally, in one embodiment of the present application, the segmented prediction module also includes: a second calculation unit, used to calculate the mean and standard deviation of the wind speed data and the irradiance data, and calculate the wind speed preprocessing data and photovoltaic power generation preprocessing data corresponding to the wind speed data and the irradiance data according to the mean and the standard deviation; a third construction unit, used to construct the wind-solar prediction model based on a preset segmented dynamic trend decomposition prediction model, and perform segmented prediction on the wind speed preprocessing data and the photovoltaic power generation preprocessing data through the wind-solar prediction model to obtain the first k Forecast wind speed and m The predicted irradiance is: k and m is a positive integer; a first acquisition unit, used to obtain the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiance corresponding to the target power system, and based on the cut-in wind speed, the rated wind speed, the cut-out wind speed and the first k Forecast wind speed for the segment and calculate the k The wind power is predicted in sections, and the mThe first segment is calculated based on the predicted irradiance, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance value and the maximum irradiance value. m The photovoltaic power generation power of the first segment; the third calculation unit is used to calculate the photovoltaic power generation power of the first segment respectively. k The wind power segment prediction power and the m Calculation of photovoltaic power generation k Total wind power forecast within the hour and m The total photovoltaic power forecast within the hour and determine the corresponding k Actual wind power generation in the hour and m The actual photovoltaic power generation in the hour is based on the k The actual wind power generation capacity and k The wind power forecast total power in the hour is used to calculate the wind power forecast power error, and the wind power forecast power error is calculated by the wind power forecast total power in the hour. m The actual photovoltaic power generation in the hour and the m The photovoltaic power generation prediction power error is calculated based on the photovoltaic predicted total power within the hour; a judgment unit is used to respectively judge whether the wind power prediction power error and the photovoltaic power generation prediction power error meet the preset error requirements, wherein when the wind power prediction power error and the photovoltaic power generation prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, power generation cost coefficient and the k The total wind power forecast within the hour is used to calculate the wind power generation cost, and the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and the m The photovoltaic power generation cost is calculated based on the total photovoltaic power forecast within the hour.

[0017] Optionally, in one embodiment of the present application, the modeling module includes: a fourth calculation unit, used to calculate the hourly thermal power generation value of the target power system, and to construct a thermal power output constraint based on the hourly thermal power generation value, a preset hourly thermal power generation minimum value and an hourly thermal power generation maximum value; a second acquisition unit, used to acquire the first n Hourly thermal power generation and n- 1 hour thermal power generation power, and according to the n Hours of thermal power generation power, the n-1 hour thermal power generation power and a preset climbing rate determine the power change rate constraint of the thermal power unit; a third determination unit, used to determine the thermal power operation and maintenance cost coefficient, fuel cost coefficient, startup cost coefficient and startup state variable, and calculate the thermal power generation cost based on the thermal power operation and maintenance cost coefficient, the fuel cost coefficient, the startup cost coefficient, the startup state variable and the hourly thermal power generation value, and construct the thermal power unit emission constraint according to the preset emission coefficient and the maximum allowable emission amount; an establishment unit, used to construct the thermal power generation model based on the thermal power generation cost, the thermal power unit emission constraint, the thermal power unit power change rate constraint and the said; a third acquisition unit, used to acquire the charging efficiency, discharge efficiency, charging power and discharge power of the target power system, and based on the charging efficiency, the discharge efficiency, The charging power and the discharging power are used to calculate the charging and discharging power of the energy storage system; a fourth determination unit is used to determine the corresponding charging and discharging power constraints and charging and discharging efficiency constraints based on the preset maximum charging power of the energy storage system, the maximum discharging power of the energy storage system, the charging efficiency and the discharging efficiency; a fourth construction unit is used to calculate the capacity of the energy storage system through the charging and discharging power of the energy storage system, and to construct the capacity constraint of the energy storage system according to the capacity of the energy storage system, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system; a fifth calculation unit is used to calculate the cost of the energy storage system based on the preset charging and discharging cost coefficient of the energy storage system, the operation and maintenance cost coefficient of the energy storage system and the charging and discharging power of the energy storage system, and to construct the energy storage system model through the cost of the energy storage system, the capacity constraint of the energy storage system, the charging and discharging power constraint and the charging and discharging efficiency constraint.

[0018] Optionally, in one embodiment of the present application, the optimization module includes: a collection unit, used to collect the power system historical environment data of the target power system, and determine the first state space, first reward function and first system historical cost corresponding to the first deep deterministic policy gradient agent; a training unit, used to determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, and construct a first action space through the parameters to be optimized, and train the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function and the first system historical cost, and in combination with the power system historical environment data, to obtain the agent target parameters; a fifth determination unit, used to determine the second state space, second action space, second reward function and second system cost corresponding to the second deep deterministic policy gradient agent, and optimize the operation of the target power system based on the second state space, the second action space, the second reward function and the second system cost, and in combination with the agent target parameters.

[0019] Optionally, in one embodiment of the present application, thek The mathematical expression of the wind power segment prediction power is:

[0020] in, Indicates the k Segment forecast wind speed; Indicates the k Segment-by-segment wind power prediction; It represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; Indicates the temperature; Indicates the fan blade radius; Indicates the wind direction correction factor; Indicates the angle between wind direction and wind turbine orientation; , , represent the cut-in wind speed, the rated wind speed and the cut-out wind speed respectively.

[0021] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-agent-based receiving-end power system optimization operation method as described in the above embodiment.

[0022] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned multi-agent-based receiving-end power system optimization operation method.

[0023] Therefore, the embodiments of the present application have the following beneficial effects: The embodiments of the present application can determine the circuit topology structure corresponding to the target power system, extract the node features and edge features corresponding to the circuit topology structure, and divide the target power system into regions based on the node features, edge features and a preset spectral clustering algorithm to generate a power system regional division result; based on a pre-constructed power system power balance model, each regional node in the power system regional division result satisfies the preset power balance constraint condition, and obtain the wind speed data and irradiance data corresponding to the target power system, and perform segmented prediction processing based on a pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain the corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; based on the pre-constructed The thermal power generation model and energy storage system model are constructed, and the dual-depth deterministic policy gradient optimization model is constructed in combination with the segmented prediction data, wind power generation cost and photovoltaic power generation cost; the power system historical environment data of the target power system is collected, and the first deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model is trained through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for the optimization of the power balance operation of multi-regional nodes in the power system, which is helpful to optimize power dispatching and reduce the operation cost of the power grid. As a result, the existing technology is difficult to effectively cope with the complex situation of multiple regions and multiple nodes, and it is difficult to take into account parameter selection and system optimization at the same time, which greatly affects the optimization effect.

[0024] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart of a receiving-end power system optimization operation method based on multi-agents provided according to an embodiment of the present application; Figure 2 A schematic diagram of the execution logic of a receiving-end power system optimization operation method based on multi-agents provided in one embodiment of the present application; Figure 3 This is an example diagram of a receiving-end power system optimization operation device based on multi-agents according to an embodiment of the present application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0026] Among them, 10-receiving-end power system optimization operation device based on multi-agent; 100-region division module, 200-segment prediction module, 300-modeling module, 400-optimization module; 401-memory, 402-processor, 403-communication interface. DETAILED DESCRIPTION

[0027] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0028] The following describes the receiving-end power system optimization operation method and device based on multi-agents according to the embodiments of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a receiving-end power system optimization operation method based on multi-agents. In this method, by determining the circuit topology structure corresponding to the target power system, and extracting the node features and edge features corresponding to the circuit topology structure, and based on the node features, edge features and a preset spectral clustering algorithm, the target power system is divided into regions to generate a power system regional division result; based on a pre-constructed power system power balance model, each regional node in the power system regional division result satisfies the preset power balance constraint condition, and obtains the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on the pre-constructed wind and solar prediction model and in combination with the wind speed data and irradiance data to obtain the corresponding segmented prediction data, and calculate the corresponding segmented prediction data according to the segmented prediction data. The corresponding wind power generation cost and photovoltaic power generation cost; based on the pre-built thermal power generation model and energy storage system model, and combined with the segmented prediction data, wind power generation cost and photovoltaic power generation cost, a dual-depth deterministic policy gradient optimization model is constructed; the power system historical environmental data of the target power system is collected, and the first deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model is trained through the power system historical environmental data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for the optimization of power system multi-region node power balance operation, helping to optimize power dispatching and reduce power grid operation costs. As a result, the existing technology is difficult to effectively cope with the complex situation of multiple regions and multiple nodes, and it is difficult to take into account parameter selection and system optimization at the same time, which greatly affects the optimization effect.

[0029] Specifically, Figure 1A flowchart of a receiving-end power system optimization operation method based on multi-agent provided in an embodiment of the present application.

[0030] like Figure 1 As shown, the receiving-end power system optimization operation method based on multi-agent includes the following steps: In step S101, a circuit topology structure corresponding to a target power system is determined, and node features and edge features corresponding to the circuit topology structure are extracted. Based on the node features, edge features and a preset spectral clustering algorithm, the target power system is divided into regions to generate a power system region division result.

[0031] The embodiment of the present application can first establish a multi-regional partition model of the power system, so as to extract the topological characteristics and power flow characteristics of the power system through the multi-regional partition model of the power system using a feature extraction method based on topological structure, and combine it with a spectral clustering algorithm to realize regional partition.

[0032] Optionally, in one embodiment of the present application, a circuit topology structure corresponding to the target power system is determined, and node features and edge features corresponding to the circuit topology structure are extracted, and the target power system is regionalized based on the node features, edge features and a preset spectral clustering algorithm to generate a power system regionalization result, including: determining a node set, an edge set and a time scale of power line power data of the circuit topology structure corresponding to the target power system; obtaining the node set in each layer of the circuit topology structure ... i and nodes j The node features of each node in the circuit topology are extracted based on the preset learnable weight matrix, activation function and neighbor node set, where: i and j is a positive integer; based on the edge set and the preset multi-layer perceptron, the time scale and node i and nodes j The node features in each layer are concatenated to obtain the edge features of each edge in the circuit topology. Based on the node features and the preset Gaussian kernel function bandwidth parameters, the node i and nodes j The similarity between them is calculated, and the corresponding similarity matrix is ​​constructed through the similarity, so as to divide the target power system into regions according to the similarity matrix and the spectral clustering algorithm to obtain the regional division Laplace matrix; the regional division Laplace matrix is ​​eigendecomposed to obtain the eigenvectors corresponding to the first q minimum eigenvalues, and the eigenvectors are clustered by K means to generate the power system regional division result, where q is a positive integer.

[0033] Specifically, the process of dividing the power system into multiple regions in the embodiment of the present application is as follows: 1. Data preprocessing: In the embodiment of the present application, the node set of the circuit topology structure is , the edge set is ; The power data of each transmission line is modeled on an hourly scale as ; The node feature matrix is , For the real number field Matrix; Constructing a weighted adjacency matrix , For the real number field Matrix, where Represents the power flow between nodes.

[0034] 2. Extract node features and edge features through a feature extraction method based on topological structure: Initialize node feature matrix and the adjacency matrix , update the node features and edge features respectively: 1) Node feature update, for each layer have: (1) in, Representation Node In the Characteristics of the layer; Representation Node In the Characteristics of the layer; Representation Node The set of neighboring nodes of Representation Node The set of neighboring nodes of is a learnable weight matrix; Represents the activation function.

[0035] 2) Edge feature update, for each edge: (2) in, Indicates that the edge Layer features; || represents feature concatenation; MLP represents multi-layer perceptron, Represents the power data of the transmission line.

[0036] Therefore, after the L-layer graph neural network, the node feature matrix can be obtained. ,in, For the real number field matrix, is the final feature dimension.

[0037] 3. Based on the extracted node features and edge features, the embodiment of the present application can use the spectral clustering algorithm to divide the power system into regions: (1) Compute nodes and nodes The similarity between: (3) in, Representation Node and nodes The similarity between and represents the extracted node features, Represents the bandwidth parameter of the Gaussian kernel function.

[0038] (2) Determine the similarity matrix based on the similarity obtained above ,in, For the real number field Matrix; Based on the similarity matrix, the embodiment of the present application can use the spectral clustering algorithm to divide the power system into regions, as described below: 1) Construct the region partition Laplace matrix: (4) in, is the degree matrix; is the Laplace matrix.

[0039] 2) Perform eigendecomposition on the region partition Laplace matrix to obtain the eigenvectors corresponding to the first q smallest eigenvalues, and Perform K-means clustering, where For the real number field Matrix, thus obtaining the final power system area division result.

[0040] Therefore, the embodiments of the present application realize regional division by extracting relevant features of the power system and combining the spectral clustering algorithm, thereby providing a reliable theoretical basis for the subsequent generation of wind and solar intelligent scenarios and the realization of power balance in the receiving power system.

[0041] In step S102, based on the pre-constructed power system power balance model, each regional node in the power system regional division result satisfies the preset power balance constraint conditions, and obtains the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on the pre-constructed wind and solar prediction model and in combination with the wind speed data and irradiance data to obtain the corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data.

[0042] Furthermore, the embodiments of the present application can also establish a power system power balance model to ensure that the load demand of each node is equal to the sum of the powers emitted by all lines connected to the node, satisfying the power balance constraint; thereafter, the embodiments of the present application can establish a wind and solar prediction model based on a segmented dynamic trend decomposition prediction model to pre-process the wind speed and photovoltaic data, and perform segmented prediction on the pre-processed wind speed and photovoltaic data, thereby ensuring that the predicted power error of wind power generation and photovoltaic power generation is within an allowable range.

[0043] Optionally, in one embodiment of the present application, based on a pre-constructed power system power balance model, each regional node in the power system regional division result satisfies a preset power balance constraint, including: determining the node power flowing through each regional node in the power system regional division result every hour, and constructing a line power constraint based on the preset maximum power flowing through the node every hour, the minimum power flowing through the node every hour, and the power flowing through the node every hour; determining the wind power generation power, photovoltaic power generation power, thermal power generation power, and energy storage system extraction power corresponding to each regional node every hour, and calculating the multi-energy total power generation per hour based on the wind power generation power, photovoltaic power generation power, thermal power generation power, and energy storage system extraction power per hour; calculating the total line load per hour based on the node power flowing through the node every hour, and constructing a power balance constraint based on the multi-energy total power generation per hour and the total line load per hour, so that each regional node satisfies the power balance constraint.

[0044] In the actual implementation process, the embodiment of the present application may be provided with a power system having regional nodes, and the power flowing through the nodes per hour is (From the node To Node ), the specific constraints are as follows: The line power constraint is: (5) in, is the maximum power flowing through the node within one hour; It is the minimum power flowing through the node within one hour.

[0045] The total line load per hour is: (6) in, is the total line load per hour.

[0046] The total power generated by multiple energy sources per hour is: (7) in, , , and They are wind power generation, photovoltaic power generation, thermal power generation and energy storage system extraction power per hour. It is the total power generated by multiple energy sources per hour.

[0047] Therefore, the power balance equation of the embodiment of the present application is: (8) Optionally, in one embodiment of the present application, wind speed data and irradiance data corresponding to the target power system are obtained, and segmented prediction processing is performed based on a pre-constructed wind-solar prediction model and in combination with the wind speed data and irradiance data to obtain corresponding segmented prediction data, and the corresponding wind power generation cost and photovoltaic power generation cost are calculated based on the segmented prediction data, including: calculating the mean and standard deviation of the wind speed data and the irradiance data, and calculating the wind speed preprocessing data and photovoltaic power generation preprocessing data corresponding to the wind speed data and the irradiance data according to the mean and the standard deviation; constructing a wind-solar prediction model based on a preset segmented dynamic trend decomposition prediction model, and performing segmented prediction on the wind speed preprocessing data and the photovoltaic power generation preprocessing data through the wind-solar prediction model to obtain the first k Forecast wind speed and m The predicted irradiance is: k and m is a positive integer; obtain the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiance corresponding to the target power system, and k Forecast wind speed for the segment and calculate the k The wind power is predicted in sections, and the m The irradiance, PV module efficiency, PV module area, minimum irradiance and maximum irradiance are calculated for the first segment. m The photovoltaic power generation power of the first k Wind power segment prediction power and m Calculation of photovoltaic power generation k Total wind power forecast within the hour and m Predict the total photovoltaic power within the hour and determine the corresponding power system k Actual wind power generation in the hour and m The actual photovoltaic power generation in the hour is based on k Actual wind power generation in the hour and k The wind power forecast total power in the hour is used to calculate the wind power forecast power error, and the wind power forecast total power is calculated by m The actual photovoltaic power generation and mThe photovoltaic power generation prediction power error is calculated based on the photovoltaic predicted total power within the hour; the wind power prediction power error and the photovoltaic power generation prediction power error are judged respectively whether they meet the preset error requirements. When the wind power prediction power error and the photovoltaic power generation prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, power generation cost coefficient and k The total wind power forecast within the hour is used to calculate the wind power generation cost, and the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and m Calculate the photovoltaic power generation cost based on the total photovoltaic power forecast within the hour.

[0048] In the specific implementation process, the embodiment of the present application can obtain the wind speed data and irradiance data corresponding to the power system (i.e., the historical data of photovoltaic power generation), and perform segmented prediction processing by combining the wind speed data and irradiance data through the wind-solar prediction model to obtain the corresponding segmented prediction data (including the segmented predicted wind power and photovoltaic power generation power), thereby using the segmented prediction data to calculate the corresponding wind power generation cost and photovoltaic power generation cost. The specific process is as follows: Forecasting wind power generation by segment: (1) Preprocess the wind speed data to make it meet the model input requirements (9) in, Preprocess data for wind speed; is the original wind speed data (i.e. wind speed data); is the mean of wind speed data; is the standard deviation of the wind speed data.

[0049] (2) The wind and solar prediction model using the segmented dynamic trend decomposition prediction model predicts the wind speed in each segment, as shown in the following formula: (10) in, and are the power inertia coefficient and the power fluctuation coefficient respectively; The predicted Wind speed segment; represents the lag operator; represents the lag operator d Power; represents white noise; q represents the sliding average order; represents the autoregressive order.

[0050] Optionally, in one embodiment of the present application, k The mathematical expression of the wind power segment prediction power is:

[0051] in, Indicates k Segment forecast wind speed; Indicates k Segmental wind power prediction power; It represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; Indicates the temperature; Indicates the fan blade radius; Indicates the wind direction correction factor; Indicates the angle between wind direction and wind turbine orientation; , , They represent the cut-in wind speed, rated wind speed and cut-out wind speed respectively.

[0052] It should be noted that based on formula (10), the embodiment of the present application can obtain The wind power generation power of the kth section (i.e. the predicted power of the kth section wind power) is: (11) in, Indicates k The predicted wind speed for the first wind speed); Indicates k The wind power prediction power of the first wind power segment predicted power); It represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; Indicates the temperature; Indicates the fan blade radius; Indicates the wind direction correction factor; Indicates the angle between wind direction and wind turbine orientation; , , They represent the cut-in wind speed, rated wind speed and cut-out wind speed respectively.

[0053] Therefore, through formula (11) we can get K Total wind power forecast within the hour for: (12) Furthermore, the embodiment of the present application can predict the total wind power within K hours based on Calculating wind power forecast error , as shown below: (13) in, for Actual wind power generation during the hour.

[0054] Afterwards, the embodiment of the present application can determine the wind power prediction power error according to the following formula: Is it within the allowable error range? (14) in, Forecast power error for wind power; is the maximum value of wind power prediction error.

[0055] Finally, in the wind power prediction error When the error is within the allowable range as shown in formula (14), the embodiment of the present application can calculate the wind power generation cost according to the predicted total wind power within K hours, as shown in the following formula: (15) in, For wind power generation costs; is the operation and maintenance cost coefficient; is the power generation cost coefficient.

[0056] Forecast of photovoltaic power generation by segment: It should be noted that, in the embodiment of the present application, the photovoltaic power generation power adopts a segmented prediction model, which is updated every hour. The specific process is as follows: (1) The historical data of photovoltaic power generation (i.e., irradiance data) is standardized to meet the model input requirements. The standardized processing expression is: (16) in, Preprocessing data for photovoltaic power generation; is the original irradiance data (i.e. irradiance data); represents the mean value of irradiance data; is the standard deviation of the irradiance data.

[0057] (2) The wind and solar prediction model using the segmented dynamic trend decomposition prediction model predicts the irradiance of each segment, as shown in the following formula: (17) in, and are the power inertia coefficient and the power fluctuation coefficient respectively; represents the lag operator; Indicates Segment predicted irradiance; represents white noise; Indicates the order.

[0058] From formula (17), we can get The photovoltaic power generation power of this section is: (18) in, Indicates The predicted power generation of photovoltaic power segment; Indicates the efficiency of photovoltaic modules; Indicates the area of ​​PV modules; and Respectively represent the minimum and maximum values ​​of irradiance; Indicates Segment predicted irradiance; Indicates the rated irradiation intensity; Indicates the solar altitude angle (radians); Indicates the solar azimuth (radians); Indicates the tilt angle of the photovoltaic panel (radians); Indicates the azimuth angle of the photovoltaic panel (radians); Indicates PV panel temperature; Indicates the reference temperature; Represents the temperature coefficient.

[0059] Therefore, the embodiment of the present application can be Calculation of photovoltaic power generation The total photovoltaic power predicted within an hour is as shown in the following formula: (19) in, Predict total power for PV segments; For the The predicted power generation of photovoltaic power station.

[0060] (3) Based on Calculate the photovoltaic power prediction error by using the total photovoltaic power prediction within the hour: (20) in, for Actual photovoltaic power generation in the hour; Predict power error for photovoltaic power generation.

[0061] The photovoltaic power generation prediction power error is used to determine the allowable range of photovoltaic prediction power error: (twenty one) in, Predicting power errors for photovoltaic power generation; The maximum value of the photovoltaic power generation prediction power error.

[0062] After that, the photovoltaic power generation prediction power error When the error is within the allowable range as shown in formula (21), the embodiment of the present application can calculate the photovoltaic power generation cost according to the photovoltaic segmented predicted total power, as shown in the following formula: (twenty two) in, is the cost of photovoltaic power generation; is the photovoltaic operation and maintenance cost coefficient; is the photovoltaic power generation cost coefficient.

[0063] Therefore, the embodiments of the present application ensure that the predicted power errors of wind power generation and photovoltaic power generation are within the allowable range by preprocessing and segmented prediction of wind speed and photovoltaic data, providing important theoretical guidance and basis for the construction of a dual-depth deterministic policy gradient optimization model.

[0064] In step S103, a dual-depth deterministic policy gradient optimization model is constructed based on the pre-built thermal power generation model and energy storage system model, and in combination with the segmented prediction data, wind power generation cost and photovoltaic power generation cost.

[0065] Furthermore, the embodiments of the present application also need to establish a thermal power generation model based on the output range, ramp rate and power generation cost constraints, and establish an energy storage system model through charging and discharging power, energy storage capacity and efficiency constraints. After that, the embodiments of the present application can use the thermal power generation model and the energy storage system model, and combine the segmented prediction data (including wind power segmented prediction power and photovoltaic power generation power, etc.), wind power generation cost and photovoltaic power generation cost to construct a parameter-sharing dual-depth deterministic policy gradient optimization model.

[0066] Optionally, in one embodiment of the present application, based on a pre-built thermal power generation model and an energy storage system model, and in combination with segmented forecast data, wind power generation costs, and photovoltaic power generation costs, a dual-depth deterministic policy gradient optimization model is constructed, including: calculating the hourly thermal power generation value of the target power system, and based on the hourly thermal power generation value, the preset hourly thermal power generation value, and the hourly thermal power generation maximum value, constructing a thermal power output constraint; obtaining the first n Hourly thermal power generation and n- 1 hour thermal power generation capacity, and according to the n Hourly thermal power generation, n-The power change rate constraint of the thermal power unit is determined based on the thermal power generation power of one hour and the preset ramp rate; the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, startup cost coefficient and startup state variable are determined, and the thermal power generation cost is calculated based on the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, startup cost coefficient, startup state variable and hourly thermal power generation value, and the emission constraint of the thermal power unit is constructed according to the preset emission coefficient and the maximum allowable emission amount; a thermal power generation model is constructed based on the thermal power generation cost, the emission constraint of the thermal power unit, the power change rate constraint of the thermal power unit and; the charging efficiency, discharge efficiency, charging power and discharge power of the target power system are obtained, and based on the charging efficiency, discharge efficiency, The charging power and discharging power are used to calculate the charging and discharging power of the energy storage system; based on the preset maximum charging power of the energy storage system, the maximum discharging power of the energy storage system, the charging efficiency and the discharging efficiency, the corresponding charging and discharging power constraints and charging and discharging efficiency constraints are determined; the energy storage system capacity is calculated through the charging and discharging power of the energy storage system, and the energy storage system capacity constraints are constructed according to the energy storage system capacity, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system; based on the preset charging and discharging cost coefficient of the energy storage system, the energy storage system operation and maintenance cost coefficient and the charging and discharging power of the energy storage system, the energy storage system cost is calculated, and the energy storage system model is constructed through the energy storage system cost, energy storage system capacity constraints, charging and discharging power constraints and charging and discharging efficiency constraints.

[0067] It should be noted that the process of establishing the thermal power generation model and the energy storage system model in the embodiment of the present application is as follows: 1. A thermal power generation model is established based on the output range, ramp rate and power generation cost constraints, as shown in the following formula: (twenty three) in, is the hourly thermal power generation value; It is the control input of thermal power unit; , , It is the coefficient of the thermal power output formula and can take different values ​​according to the operating conditions.

[0068] The thermal power output constraints are: (twenty four) in, It is the hourly thermal power generation value; It is the maximum value of thermal power generation per hour.

[0069] In the embodiment of the present application, the power change rate of the thermal power unit is limited by the ramp rate, as shown in the following formula: (25) in, Indicates Hours of thermal power generation, Indicates Thermal power generation capacity per hour; is the climbing rate; The time interval is one hour.

[0070] Afterwards, the embodiment of the present application can calculate the thermal power generation cost according to the hourly thermal power generation value: (26) in, The cost of thermal power generation; is the thermal power generation operation and maintenance cost coefficient; is the fuel cost coefficient; is the startup cost coefficient; is the startup state variable, which is shown in the following formula: (27) In addition, the emission constraints of the thermal power unit in the embodiment of the present application are: (28) (29) in, is the emission; , , is the emission factor; is the maximum allowable emission.

[0071] 2. Establish an energy storage system model based on charging and discharging power, energy storage capacity and efficiency constraints: In the embodiment of the present application, the charging and discharging power of the energy storage system is: (30) in, For charging efficiency; is the discharge efficiency; is the power during charging; is the power during discharge; The charging and discharging power of the energy storage system.

[0072] The charge and discharge power constraints are determined based on the above charge power and discharge power, as shown in the following formula: (31) (32) in, Indicates the maximum charging power of the energy storage system; Indicates the maximum discharge power of the energy storage system.

[0073] In addition, the charge and discharge efficiency constraints in the embodiments of the present application are: (33) (34) in, For charging efficiency; is the discharge efficiency.

[0074] Then, the embodiment of the present application can calculate the capacity of the energy storage system based on the charging and discharging power of the energy storage system, as shown in the following formula: (35) in, For energy storage systems Hour capacity; For energy storage systems Hourly capacity; The time interval is one hour.

[0075] The above energy storage system capacity is used to determine the energy storage system capacity constraint: (36) in, is the minimum capacity of the energy storage system; is the maximum capacity of the energy storage system.

[0076] Finally, the embodiment of the present application can calculate the cost of the energy storage system according to the charging and discharging power of the energy storage system, as shown in the following formula: (37) in, The cost of the energy storage system; is the charging and discharging cost coefficient of the energy storage system; is the operation and maintenance cost coefficient of the energy storage system.

[0077] Therefore, the embodiment of the present application provides data support for the construction of a dual-depth deterministic policy gradient optimization model by establishing a thermal power generation model and an energy storage system model, thereby ensuring the smooth realization of wind and solar intelligent scene generation and power balance in the receiving power system.

[0078] In step S104, the historical power system environment data of the target power system is collected, and the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model is trained through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

[0079] Afterwards, the embodiment of the present application can train a first deep deterministic policy gradient agent (i.e., the first deep deterministic policy gradient agent) through historical environmental data (i.e., the historical environmental data of the power system) to select the corresponding optimal parameters of the agent (i.e., the target parameters of the agent), and after obtaining the optimal parameters of the agent, use a second deep deterministic policy gradient agent (i.e., the second deep deterministic policy gradient agent) to interact with the environment, thereby optimizing the operation of the power system.

[0080] Therefore, the embodiments of the present application can effectively cope with the complexity and uncertainty of multi-regional nodes in the power system, and improve the stability and economy of the power system.

[0081] Optionally, in one embodiment of the present application, historical power system environment data of the target power system is collected, and the first deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model is trained through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, including: collecting the historical power system environment data of the target power system, and determining the first state space, the first reward function and the first system function corresponding to the first deep deterministic policy gradient agent. system historical cost; determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, and construct a first action space through the parameters to be optimized, and based on the first state space, the first action space, the first reward function and the first system historical cost, and in combination with the power system historical environmental data, train the first deep deterministic policy gradient agent to obtain the agent target parameters; determine the second state space, the second action space, the second reward function and the second system cost corresponding to the second deep deterministic policy gradient agent, and based on the second state space, the second action space, the second reward function and the second system cost, and in combination with the agent target parameters, optimize the operation of the target power system.

[0082] It should be noted that the embodiment of the present application can first collect the historical environmental data of the target power system (i.e., the receiving power system), and determine the first agent state space (i.e., the first state space), the first agent action space (i.e., the first action space), the first agent reward function (i.e., the first reward function) and the system cost (i.e., the first system cost) corresponding to the first deep deterministic policy gradient agent.

[0083] Specifically, the state space of the first agent is for: (38) in, For historical load data, For historical wind power, For historical wind and solar power generation forecast data, For historical thermal power output data, It is the charging and discharging data of the historical energy storage system.

[0084] The embodiment of the present application constructs the first agent action space by using the discount factor and the time step (i.e. the parameters to be optimized) : (39) in, is the discount factor; is the time step.

[0085] First agent reward function for: (40) System Cost for: (41) in, , , , They are the historical cost of wind power generation, the historical cost of photovoltaic power generation, the historical cost of thermal power generation and the historical cost of energy storage system.

[0086] Afterwards, the embodiment of the present application can train the first deep deterministic policy gradient agent through the historical environmental data of the power system, combined with the first agent state space, the first agent action space, the first agent reward function and the system cost, to obtain the agent target parameters, that is, the optimal discount factor is and the optimal time step is .

[0087] Furthermore, the embodiment of the present application can use the agent target parameter (i.e., the optimal discount factor) selected by the first deep deterministic policy gradient agent through the second deep deterministic policy gradient agent , the optimal time step ) interact with the environment to optimize power system operation.

[0088] Among them, the state space of the second agent (i.e., the second state space) for: (42) The action space of the second agent (i.e., the second action space) for: (43) The reward function of the second agent (i.e. the second reward function) for: (44) The system cost (i.e. the second system cost) is: (45) in, , , and The costs are wind power generation, photovoltaic power generation, thermal power generation and energy storage system costs.

[0089] It can be understood that the embodiments of the present application effectively optimize the balanced power of a complex power system containing multiple energy sources through the architecture of a power system multi-region division model, a power system power balance model, a wind and solar prediction model based on a segmented dynamic trend decomposition prediction model, a thermal power generation and energy storage system model, and a parameter-sharing dual deep deterministic policy gradient optimization model. It can not only effectively cope with the complexity and uncertainty of multi-regional nodes in the power system and improve the stability and economy of the power system, but also consider the selection of deep deterministic policy gradient agent parameters, so that the learning efficiency of the agent is improved and a higher reward function value is achieved; in addition, the embodiments of the present application provide a more reliable basis for the optimization of power balance operation of multi-regional nodes in the power system, which helps to optimize power dispatch and reduce the operating costs of the power grid.

[0090] The following further illustrates the execution logic of the multi-agent based receiving-end power system optimization operation method of the present application in combination with the accompanying drawings.

[0091] Figure 2 FIG. 1 is a schematic diagram of the execution logic of the multi-agent-based receiving-end power system optimization operation method of the present application. Figure 2 As shown, the execution steps of the receiving-end power system optimization operation method based on multi-agent of the present application are as follows: S201: Establish a multi-regional partition model for the power system, extract the topological features and power flow features of the power system through a feature extraction method based on topological structure, and realize regional partition in combination with spectral clustering algorithm; S202: Establishing a power balance model of the power system to ensure that the load demand of each node is equal to the sum of the power output of all lines connected to the node, and satisfying the power balance constraint condition; S203: Establish a wind and solar power prediction model based on a segmented dynamic trend decomposition prediction model, pre-process the wind speed and photovoltaic data, and perform segmented prediction to ensure that the predicted power errors of wind power generation and photovoltaic power generation are within an allowable range; S204: Establishing thermal power generation and energy storage system models, establishing a thermal power generation model based on output range, ramp rate, and power generation cost constraints, and establishing an energy storage system model based on charge and discharge power, energy storage capacity, and efficiency constraints; S205: Establish a dual deep deterministic policy gradient optimization model with parameter sharing. Based on historical environmental data, use the first deep deterministic policy gradient agent training to select the optimal parameters of the agent. After obtaining the optimal parameters, use the second deep deterministic policy gradient agent to interact with the environment to optimize the operation of the power system.

[0092] According to the receiving-end power system optimization operation method based on multi-agent proposed in the embodiment of the present application, by determining the circuit topology structure corresponding to the target power system, and extracting the node features and edge features corresponding to the circuit topology structure, and based on the node features, edge features and a preset spectral clustering algorithm, the target power system is divided into regions to generate a power system regional division result; based on a pre-constructed power system power balance model, each regional node in the power system regional division result satisfies the preset power balance constraint condition, and obtains the wind speed data and irradiance data corresponding to the target power system, and performs segmented prediction processing based on the pre-constructed wind and solar prediction model and in combination with the wind speed data and irradiance data to obtain the corresponding segmented prediction data, and calculates the corresponding wind power generation cost and according to the segmented prediction data. Photovoltaic power generation cost; based on the pre-built thermal power generation model and energy storage system model, and combined with the segmented prediction data, wind power generation cost and photovoltaic power generation cost, a dual-depth deterministic policy gradient optimization model is constructed; the power system historical environmental data of the target power system is collected, and the first deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model is trained through the power system historical environmental data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for the optimization of power balance operation of multi-regional nodes in the power system, which is helpful to optimize power dispatching and reduce power grid operation costs.

[0093] Secondly, the receiving-end power system optimization operation device based on multi-agents proposed according to the embodiment of the present application is described with reference to the accompanying drawings.

[0094] Figure 3 It is a block diagram of a receiving-end power system optimization operation device based on multi-agents according to an embodiment of the present application.

[0095] like Figure 3 As shown, the receiving-end power system optimization operation device 10 based on multi-agent includes: a region division module 100, a segment prediction module 200, a modeling module 300 and an optimization module 400.

[0096] Among them, the regional division module 100 is used to determine the circuit topology structure corresponding to the target power system, and extract the node features and edge features corresponding to the circuit topology structure, and divide the target power system into regions based on the node features, edge features and a preset spectral clustering algorithm to generate a power system regional division result.

[0097] The segmented prediction module 200 is used to make each regional node in the power system regional division result meet the preset power balance constraint conditions based on the pre-constructed power system power balance model, and obtain the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on the pre-constructed wind and solar prediction model and in combination with the wind speed data and irradiance data to obtain the corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data.

[0098] The modeling module 300 is used to construct a dual-depth deterministic policy gradient optimization model based on a pre-built thermal power generation model and an energy storage system model, and in combination with segmented prediction data, wind power generation cost, and photovoltaic power generation cost.

[0099] The optimization module 400 is used to collect the historical environmental data of the target power system, and train the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model through the historical environmental data of the power system to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

[0100] Optionally, in one embodiment of the present application, the region division module 100 includes: a first determination unit, an extraction unit, a feature concatenation unit, a first calculation unit and a feature decomposition unit.

[0101] The first determination unit is used to determine the node set, edge set and time scale of the power data of the transmission line of the circuit topology structure corresponding to the target power system.

[0102] Extraction unit, used to obtain nodes in each layer of the circuit topology i and nodes j The node features of each node in the circuit topology are extracted based on the preset learnable weight matrix, activation function and neighbor node set, where: i and j Is a positive integer.

[0103] Feature splicing unit, used to analyze the time scale and node of power transmission line data based on edge set and preset multi-layer perceptroni and nodes j The node features in each layer are concatenated to obtain the edge features of each edge in the circuit topology.

[0104] The first calculation unit is used to calculate the node based on the node characteristics and the preset Gaussian kernel function bandwidth parameter. i and nodes j The similarity between them is calculated, and the corresponding similarity matrix is ​​constructed through the similarity, so as to divide the target power system into regions according to the similarity matrix and spectral clustering algorithm to obtain the regional division Laplace matrix.

[0105] The eigendecomposition unit is used to perform eigendecomposition on the regional partition Laplace matrix to obtain the eigenvectors corresponding to the first q smallest eigenvalues, and perform K-means clustering on the eigenvectors to generate the power system regional partition result, wherein q is a positive integer.

[0106] Optionally, in one embodiment of the present application, the segment prediction module 200 includes: a first construction unit, a second determination unit and a second construction unit.

[0107] Among them, the first construction unit is used to determine the hourly node power flowing through each regional node in the power system regional division result, and construct the line power constraint based on the preset hourly maximum power flowing through the node, hourly minimum power flowing through the node and hourly power flowing through the node.

[0108] The second determination unit is used to determine the wind power generation power within each hour, the photovoltaic power generation power within each hour, the thermal power generation power within each hour and the energy storage system extraction power within each hour corresponding to each regional node, and calculate the total multi-energy power generation power within each hour based on the wind power generation power within each hour, the photovoltaic power generation power within each hour, the thermal power generation power within each hour and the energy storage system extraction power within each hour.

[0109] The second construction unit is used to calculate the total line load per hour based on the power flowing through the node per hour, and to construct a power balance constraint condition based on the total power generation power of multiple energy sources per hour and the total line load per hour, so that each regional node meets the power balance constraint condition.

[0110] Optionally, in one embodiment of the present application, the segment prediction module 200 further includes: a second calculation unit, a third construction unit, a first acquisition unit, a third calculation unit and a judgment unit.

[0111] The second calculation unit is used to calculate the mean and standard deviation of the wind speed data and the irradiance data, and calculate the wind speed preprocessing data and photovoltaic power generation preprocessing data corresponding to the wind speed data and the irradiance data according to the mean and the standard deviation.

[0112] The third construction unit is used to construct a wind and solar power prediction model based on a preset segmented dynamic trend decomposition prediction model, and perform segmented prediction on the wind speed preprocessing data and the photovoltaic power generation preprocessing data through the wind and solar power prediction model to obtain the first k Forecast wind speed and m The predicted irradiance is: k and m Is a positive integer.

[0113] The first acquisition unit is used to obtain the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiance corresponding to the target power system, and based on the cut-in wind speed, rated wind speed, cut-out wind speed and the first k Forecast wind speed for the segment and calculate the k The wind power is predicted in sections, and the m The irradiance, PV module efficiency, PV module area, minimum irradiance and maximum irradiance are calculated for the first segment. m The photovoltaic power generation capacity.

[0114] The third computing unit is used to respectively k Wind power segment prediction power and m Calculation of photovoltaic power generation k Total wind power forecast within the hour and m Predict the total photovoltaic power within the hour and determine the corresponding power system k Actual wind power generation in the hour and m The actual photovoltaic power generation in the hour is based on k Actual wind power generation in the hour and k The wind power forecast total power in the hour is used to calculate the wind power forecast power error, and the wind power forecast total power is calculated by m The actual photovoltaic power generation and m The photovoltaic power generation prediction error is calculated based on the total photovoltaic power prediction within the hour.

[0115] The judgment unit is used to judge whether the wind power prediction power error and the photovoltaic power generation prediction power error meet the preset error requirements, wherein when the wind power prediction power error and the photovoltaic power generation prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, power generation cost coefficient and k The total wind power forecast within the hour is used to calculate the wind power generation cost, and the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and m Calculate the photovoltaic power generation cost based on the total photovoltaic power forecast within the hour.

[0116] Optionally, in one embodiment of the present application, the modeling module 300 includes: a fourth calculation unit, a second acquisition unit, a third determination unit, a building unit, a third acquisition unit, a fourth determination unit, a fourth construction unit and a fifth calculation unit.

[0117] Among them, the fourth calculation unit is used to calculate the hourly thermal power generation value of the target power system, and construct the thermal power output constraint based on the hourly thermal power generation value, the preset hourly thermal power generation minimum value and the hourly thermal power generation maximum value.

[0118] The second acquisition unit is used to acquire the n Hourly thermal power generation and n- 1 hour thermal power generation capacity, and according to the n Hourly thermal power generation, n- The 1-hour thermal power generation capacity and the preset ramp rate determine the power change rate constraint of the thermal power unit.

[0119] The third determination unit is used to determine the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, startup cost coefficient and startup state variable, and calculate the thermal power generation cost based on the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, startup cost coefficient, startup state variable and hourly thermal power generation value, and construct the emission constraint of the thermal power unit according to the preset emission coefficient and the maximum allowable emission amount.

[0120] A unit is established to construct a thermal power generation model based on the thermal power generation cost, the emission constraints of the thermal power units, the power change rate constraints of the thermal power units, and.

[0121] The third acquisition unit is used to acquire the charging efficiency, discharging efficiency, charging power and discharging power of the target power system, and calculate the charging and discharging power of the energy storage system based on the charging efficiency, discharging efficiency, charging power and discharging power.

[0122] The fourth determining unit is used to determine corresponding charging and discharging power constraints and charging and discharging efficiency constraints based on preset maximum charging power of the energy storage system, maximum discharging power of the energy storage system, charging efficiency, and discharging efficiency.

[0123] The fourth construction unit is used to calculate the capacity of the energy storage system through the charging and discharging power of the energy storage system, and to construct the capacity constraint of the energy storage system according to the capacity of the energy storage system, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system.

[0124] The fifth calculation unit is used to calculate the cost of the energy storage system based on the preset energy storage system charging and discharging cost coefficient, the energy storage system operation and maintenance cost coefficient and the energy storage system charging and discharging power, and to construct an energy storage system model through the energy storage system cost, energy storage system capacity constraint, charging and discharging power constraint and charging and discharging efficiency constraint.

[0125] Optionally, in one embodiment of the present application, the optimization module 400 includes: a collection unit, a training unit and a fifth determination unit.

[0126] Among them, the collection unit is used to collect the historical environmental data of the target power system, and determine the first state space, the first reward function and the first system historical cost corresponding to the first deep deterministic policy gradient agent.

[0127] A training unit is used to determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, and construct a first action space through the parameters to be optimized, and train the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function and the first system historical cost, and in combination with the historical environmental data of the power system to obtain the agent target parameters.

[0128] The fifth determination unit is used to determine the second state space, the second action space, the second reward function and the second system cost corresponding to the second deep deterministic policy gradient agent, and optimize the operation of the target power system based on the second state space, the second action space, the second reward function and the second system cost, and in combination with the agent target parameters.

[0129] Optionally, in one embodiment of the present application, k The mathematical expression of the wind power segment prediction power is:

[0130] in, Indicates k Segment forecast wind speed; Indicates k Segmental wind power prediction power; It represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; Indicates the temperature; Indicates the fan blade radius; Indicates the wind direction correction factor; Indicates the angle between wind direction and wind turbine orientation; , , They represent the cut-in wind speed, rated wind speed and cut-out wind speed respectively.

[0131] It should be noted that the aforementioned explanation of the embodiment of the receiving-end power system optimization operation method based on multi-agent is also applicable to the receiving-end power system optimization operation device based on multi-agent of this embodiment, and will not be repeated here.

[0132] The receiving-end power system optimization operation device based on multi-agents proposed in the embodiment of the present application includes a regional division module 100, which is used to determine the circuit topology structure corresponding to the target power system, and extract the node features and edge features corresponding to the circuit topology structure, and divide the target power system into regions based on the node features, edge features and a preset spectral clustering algorithm to generate a power system regional division result; a segmented prediction module 200, which is used to make each regional node in the power system regional division result meet the preset power balance constraint condition based on a pre-constructed power system power balance model, and obtain the wind speed data and irradiance data corresponding to the target power system, and perform segmented prediction processing based on the pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain the corresponding segmented prediction data, and calculate the corresponding wind power generation cost according to the segmented prediction data. cost and photovoltaic power generation cost; a modeling module 300, which is used to construct a dual-depth deterministic policy gradient optimization model based on a pre-built thermal power generation model and an energy storage system model, and in combination with segmented prediction data, wind power generation cost and photovoltaic power generation cost; an optimization module 400, which is used to collect the power system historical environmental data of the target power system, and train the first deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model through the power system historical environmental data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual-depth deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for the optimization of power system multi-region node power balance operation, which is helpful to optimize power dispatching and reduce grid operation costs.

[0133] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include: Memory 401 , processor 402 , and a computer program stored in the memory 401 and executable on the processor 402 .

[0134] When the processor 402 executes the program, the receiving-end power system optimization operation method based on multi-agent provided in the above embodiment is implemented.

[0135] Furthermore, the electronic device further comprises: The communication interface 403 is used for communication between the memory 401 and the processor 402 .

[0136] The memory 401 is used to store computer programs that can be executed on the processor 402 .

[0137] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0138] If the memory 401, the processor 402 and the communication interface 403 are implemented independently, the communication interface 403, the memory 401 and the processor 402 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0139] Optionally, in a specific implementation, if the memory 401, the processor 402 and the communication interface 403 are integrated on a chip, the memory 401, the processor 402 and the communication interface 403 can communicate with each other through an internal interface.

[0140] The processor 402 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0141] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned multi-agent-based receiving-end power system optimization operation method.

[0142] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0143] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0144] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0145] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.

[0146] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0147] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0148] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0149] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A receiving-end power system optimization operation method based on multi-agent, characterized in that: The following steps are involved: Determine a circuit topology structure corresponding to a target power system, extract node features and edge features corresponding to the circuit topology structure, and divide the target power system into regions based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system region division result; Based on a pre-constructed power system power balance model, each regional node in the power system regional division result satisfies a preset power balance constraint condition, and obtains wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on a pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain corresponding segmented prediction data, and calculate corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; Based on the pre-built thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost, a dual-depth deterministic policy gradient optimization model is constructed; Collect historical power system environment data of the target power system, and train the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

2. The receiving-end power system optimization operation method based on multi-agent according to claim 1 is characterized in that: The determining of the circuit topology structure corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology structure, and performing regional division on the target power system based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system regional division result, includes: Determine a node set, an edge set and a time scale of power data of a transmission line of a circuit topology structure corresponding to the target power system; Get the nodes in each layer of the circuit topology i and nodes j A set of neighboring nodes is obtained, and based on a preset learnable weight matrix, an activation function and the set of neighboring nodes, node features of each node in the circuit topology are extracted, wherein: i and j is a positive integer; Based on the edge set and the preset multi-layer perceptron, the time scale of the power data of the transmission line, the node i and the node j Performing feature concatenation operation on node features in each layer to obtain edge features of each edge in the circuit topology structure; Based on the node characteristics and the preset Gaussian kernel function bandwidth parameters, the node i and the node j The similarity between the two targets is calculated, and a corresponding similarity matrix is ​​constructed through the similarity, so as to perform regional division on the target power system according to the similarity matrix and the spectral clustering algorithm, so as to obtain a regional division Laplace matrix; Performing eigendecomposition on the region partition Laplace matrix to obtain eigenvectors corresponding to the first q minimum eigenvalues, and performing K-means clustering on the eigenvectors to generate the power system region partition result, wherein q is a positive integer.

3. The receiving-end power system optimization operation method based on multi-agent according to claim 1 is characterized in that: The power system power balance model based on the pre-constructed one makes each regional node in the power system regional division result meet the preset power balance constraint condition, including: Determine the hourly node power flowing through each regional node in the power system regional division result, and construct a line power constraint based on the preset hourly maximum node power flowing through the node, the hourly minimum node power flowing through the node and the hourly node power flowing through the node; Determine the wind power generation power within each hour, the photovoltaic power generation power within each hour, the thermal power generation power within each hour, and the energy storage system extraction power within each hour corresponding to each regional node, and calculate the multi-energy total power generation power within each hour based on the wind power generation power within each hour, the photovoltaic power generation power within each hour, the thermal power generation power within each hour, and the energy storage system extraction power within each hour; The total line load per hour is calculated according to the power flowing through the node per hour, and the power balance constraint condition is constructed based on the total power generation power of multiple energy sources per hour and the total line load per hour, so that each regional node satisfies the power balance constraint condition.

4. The receiving-end power system optimization operation method based on multi-agent according to claim 3 is characterized in that: The step of acquiring wind speed data and irradiance data corresponding to the target power system, performing segmented prediction processing based on a pre-built wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain corresponding segmented prediction data, and calculating corresponding wind power generation costs and photovoltaic power generation costs according to the segmented prediction data, includes: Calculating the mean and standard deviation of the wind speed data and the irradiance data, and respectively calculating wind speed preprocessing data and photovoltaic power generation preprocessing data corresponding to the wind speed data and the irradiance data according to the mean and the standard deviation; Based on the preset segmented dynamic trend decomposition prediction model, the wind-solar prediction model is constructed, and the wind speed preprocessing data and the photovoltaic power generation preprocessing data are segmentedly predicted by the wind-solar prediction model to obtain the first k Forecast wind speed and m The predicted irradiance is: k and m is a positive integer; Obtain the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiance corresponding to the target power system, and calculate the cut-in wind speed, rated wind speed, cut-out wind speed and the first k Forecast wind speed for the segment, calculate the k The wind power is predicted in sections, and the m The first segment is calculated based on the predicted irradiance, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance value and the maximum irradiance value. m The photovoltaic power generation power of the segment; Respectively through the k The wind power segment prediction power and the m Calculation of photovoltaic power generation k Total wind power forecast within the hour and m The total photovoltaic power forecast within the hour and determine the corresponding k Actual wind power generation in the hour and m The actual photovoltaic power generation in the hour is based on the k The actual wind power generation capacity and k The wind power forecast total power in the hour is used to calculate the wind power forecast power error, and the wind power forecast power error is calculated by the wind power forecast total power in the hour. m The actual photovoltaic power generation in the hour and the m Calculate the photovoltaic power prediction error by using the total photovoltaic power prediction within the hour; respectively judging whether the wind power prediction power error and the photovoltaic power prediction power error meet the preset error requirements, wherein, when the wind power prediction power error and the photovoltaic power prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, the power generation cost coefficient and the k The total wind power forecast within the hour is used to calculate the wind power generation cost, and the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and the m The photovoltaic power generation cost is calculated based on the total photovoltaic power forecast within the hour.

5. The receiving-end power system optimization operation method based on multi-agent according to claim 1 is characterized in that: The dual-depth deterministic policy gradient optimization model is constructed based on the pre-built thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost, including: Calculating the hourly thermal power generation value of the target power system, and constructing a thermal power output constraint based on the hourly thermal power generation value, a preset hourly thermal power generation minimum value, and an hourly thermal power generation maximum value; Get the n Hourly thermal power generation and n- 1 hour thermal power generation power, and according to the n Hours of thermal power generation power, the n- The 1-hour thermal power generation power and the preset ramp rate determine the power change rate constraint of the thermal power unit; Determine a thermal power generation operation and maintenance cost coefficient, a fuel cost coefficient, a startup cost coefficient and a startup state variable, and calculate the thermal power generation cost based on the thermal power generation operation and maintenance cost coefficient, the fuel cost coefficient, the startup cost coefficient, the startup state variable and the hourly thermal power generation value, and construct an emission constraint for the thermal power unit according to a preset emission coefficient and a maximum allowable emission amount; Based on the thermal power generation cost, the emission constraint of the thermal power unit, the power change rate constraint of the thermal power unit and the thermal power generation model, constructing the thermal power generation model; Obtaining the charging efficiency, discharging efficiency, charging power and discharging power of the target power system, and calculating the charging and discharging power of the energy storage system based on the charging efficiency, the discharging efficiency, the charging power and the discharging power; Determine corresponding charge and discharge power constraints and charge and discharge efficiency constraints based on a preset maximum charge power of the energy storage system, a maximum discharge power of the energy storage system, the charge efficiency, and the discharge efficiency; Calculating the capacity of the energy storage system by the charge and discharge power of the energy storage system, and constructing the capacity constraint of the energy storage system according to the capacity of the energy storage system, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system; Based on the preset energy storage system charging and discharging cost coefficient, the energy storage system operation and maintenance cost coefficient and the energy storage system charging and discharging power, the energy storage system cost is calculated, and the energy storage system model is constructed through the energy storage system cost, the energy storage system capacity constraint, the charging and discharging power constraint and the charging and discharging efficiency constraint.

6. The receiving-end power system optimization operation method based on multi-agent according to claim 5 is characterized in that: The collecting of the power system historical environment data of the target power system, and training the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model by the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, includes: Collecting historical power system environment data of the target power system, and determining a first state space, a first reward function, and a first system historical cost corresponding to the first deep deterministic policy gradient agent; Determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, and construct a first action space through the parameters to be optimized, and train the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function and the first system historical cost, and in combination with the power system historical environment data, to obtain the agent target parameters; Determine a second state space, a second action space, a second reward function, and a second system cost corresponding to the second deep deterministic policy gradient agent, and optimize the operation of the target power system based on the second state space, the second action space, the second reward function, and the second system cost, in combination with the agent target parameters.

7. The receiving-end power system optimization operation method based on multi-agent according to claim 4 is characterized in that: The said k The mathematical expression of the wind power segment prediction power is: in, Indicates the k Segment forecast wind speed; Indicates the k Segment-by-segment wind power prediction; It represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; Indicates the temperature; Indicates the fan blade radius; Indicates the wind direction correction factor; Indicates the angle between wind direction and wind turbine orientation; , , represent the cut-in wind speed, the rated wind speed and the cut-out wind speed respectively.

8. A receiving-end power system optimization operation device based on multi-agent, characterized in that: include: A regional division module, used to determine a circuit topology structure corresponding to a target power system, extract node features and edge features corresponding to the circuit topology structure, and perform regional division on the target power system based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system regional division result; A segmented prediction module is used to make each regional node in the regional division result of the power system meet the preset power balance constraint condition based on the pre-constructed power balance model of the power system, and obtain the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on the pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain the corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; A modeling module, for constructing a dual-depth deterministic policy gradient optimization model based on a pre-built thermal power generation model and an energy storage system model, and in combination with the segmented prediction data, the wind power generation cost, and the photovoltaic power generation cost; An optimization module is used to collect historical power system environment data of the target power system, and train the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model through the historical power system environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-agent-based receiving-end power system optimization operation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the receiving-end power system optimization operation method based on multi-agent as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Energy storage scheduling decision-making method and system based on optimal wind-fire storage combined operation index

    CN116454891A

  • Wind-solar-thermal storage system capacity optimization configuration method considering thermal power flexible peak regulation transformation

    CN116799795A

  • Distribution network distributed energy storage optimization configuration method and system

    CN117154778A

  • Power distribution network regulation and control method and system

    CN119543188A