Multi-agent-based optimal operation method and device for receiving-end power system

Through the multi-agent method, the circuit topology and region division of the power system are optimized, combined with power balance and landscape prediction models, the problem of the existing technology being difficult to cope with the complex situation of multiple regions and multiple nodes is solved, and efficient optimization and stability improvement of the power system is achieved.

CN119994903BActive Publication Date: 2025-07-01CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510468427.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-01
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the complex situation of multiple regions and multiple nodes, and it is difficult to take into account both parameter selection and system optimization, which greatly affects the optimization effect.

Method used

The optimization operation method of the receiving power system based on multi-agents is adopted to optimize the operation operation of the power system by determining the circuit topology, region division, power balance model, scenery prediction model and dual-depth deterministic strategy gradient optimization model of the power system.

Benefits of technology

It realizes efficient power balance optimization for multi-region nodes, improves the stability and economy of the power system, and reduces the operating costs of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119994903B_ABST
    Figure CN119994903B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of power systems, and particularly to an optimized operation method and device for a receiving-end power system based on multi-agent. Among them, the method includes: dividing the power system into regions, and ensuring that each regional node in the power system regional division result satisfies the power balance constraint condition; establishing a wind and light prediction model to obtain segmented prediction data; based on the segmented prediction data and the pre-constructed thermal power generation and energy storage system models, establishing a double deep deterministic policy gradient optimization model, and training the first deep deterministic policy gradient agent on the basis of historical environmental data to determine the optimal parameters of the agent, so that the second deep deterministic policy gradient agent interacts with the environment according to the optimal parameters of the agent to optimize the operation of the power system. Thus, it solves the problems in the prior art that it is difficult to effectively handle complex situations of multiple regions and multiple nodes, and it is difficult to take into account parameter selection and system optimization at the same time, which greatly affects the optimization effect, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of power systems, and particularly to an optimal operation method and device for a receiving-end power system based on multi-agent systems. Background Art

[0002] With the continuous expansion of the scale of power systems and the widespread application of renewable energy, the complexity and uncertainty of power systems have increased significantly. In order to effectively address the complexity and uncertainty of multi-region nodes in power systems and improve the stability and economy of power systems, an efficient and accurate power balance optimization method for multi-region nodes in power systems is needed to provide a scientific basis for power system operation.

[0003] In the existing technology system, traditional power system optimization methods often have difficulty effectively dealing with complex situations of multi-regions and multi-nodes. Especially when considering the volatility of renewable energy such as wind power and photovoltaic power and the charge and discharge strategies of energy storage systems, the limitations of traditional methods are more obvious. When the deep deterministic policy gradient algorithm is applied to power system optimization, a single optimization algorithm is usually adopted, making it difficult to simultaneously consider parameter selection and system optimization, resulting in poor optimization effects, which urgently need to be solved. Summary of the Invention

[0004] This application provides an optimal operation method and device for a receiving-end power system based on multi-agent systems to solve the problems that the existing technology is difficult to effectively deal with complex situations of multi-regions and multi-nodes, and is difficult to simultaneously consider parameter selection and system optimization, greatly affecting the optimization effect.

[0005] The first aspect embodiment of the present application provides an optimal operation method for a receiving-end power system based on multi-agent, including the following steps: determining the circuit topology corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology, and based on the node features, the edge features and a preset spectral clustering algorithm, performing regional division on the target power system to generate a power system regional division result; based on a pre-constructed power system power balance model, enabling each regional node in the power system regional division result to satisfy a preset power balance constraint condition, and obtaining the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on a pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain corresponding segmented prediction data, and calculating the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; based on a pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost, constructing a double deep deterministic policy gradient optimization model; collecting the power system historical environment data of the target power system, and training the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

[0006] Optionally, in an embodiment of the present application, the determining the circuit topology corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology, and based on the node features, the edge features and a preset spectral clustering algorithm, performing regional division on the target power system to generate a power system regional division result includes: determining the node set, edge set and time scale of the transmission line power data of the circuit topology corresponding to the target power system; obtaining the neighboring node sets of the nodes i and nodes j in each layer of the circuit topology, and based on a preset learnable weight matrix, activation function and the neighboring node sets, extracting the node features of each node in the circuit topology, where i and j are positive integers; based on the edge set and a preset multi-layer perceptron, performing a feature splicing operation on the time scale of the transmission line power data, the nodes i and the nodes j in the node features of each layer to obtain the edge features of each edge in the circuit topology; based on the node features and a preset Gaussian kernel function bandwidth parameter, calculating the nodes iand the node j The similarity between them, and a corresponding similarity matrix is constructed based on the similarity, so as to perform regional division on the target power system according to the similarity matrix and the spectral clustering algorithm to obtain a regional division Laplacian matrix; perform eigenvalue decomposition on the regional division Laplacian matrix to obtain eigenvectors corresponding to the first q minimum eigenvalues, and perform K-means clustering on the eigenvectors to generate the power system regional division result, where q is a positive integer.

[0007] Optionally, in an embodiment of the present application, making each regional node in the power system regional division result satisfy a preset power balance constraint condition based on a pre-constructed power system power balance model includes: determining the power flowing through the node per hour of each regional node in the power system regional division result, and constructing a line power constraint based on the preset maximum power flowing through the node per hour, the minimum power flowing through the node per hour, and the power flowing through the node per hour; determining the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour corresponding to each regional node, and calculating the total multi-energy power generation per hour according to the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour; calculating the total line load per hour according to the power flowing through the node per hour, and constructing the power balance constraint condition based on the total multi-energy power generation per hour and the total line load per hour, so that each regional node satisfies the power balance constraint condition.

[0008] Optionally, in an embodiment of the present application, obtaining the wind speed data and irradiance data corresponding to the target power system, and based on a pre-constructed wind-solar prediction model, combining the wind speed data and the irradiance data to perform segmented prediction processing to obtain corresponding segmented prediction data, and calculating the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data, includes: calculating the mean and standard deviation of the wind speed data and the irradiance data, and respectively calculating the wind speed preprocessing data and the photovoltaic preprocessing data corresponding to the wind speed data and the irradiance data according to the mean and the standard deviation; constructing the wind-solar prediction model based on a preset segmented dynamic trend decomposition prediction model, and performing segmented prediction on the wind speed preprocessing data and the photovoltaic preprocessing data through the wind-solar prediction model to obtain the predicted wind speed of the k th segment and the predicted irradiance of the m th segment, where k and mis a positive integer; obtain the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance, and maximum irradiance corresponding to the target power system, and based on the cut-in wind speed, the rated wind speed, the cut-out wind speed, and the k segment predicted wind speed, calculate the k segment wind power segmented predicted power, and calculate the m segment photovoltaic power generation power through the m segment predicted irradiance, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance, and the maximum irradiance; respectively calculate the k hour total predicted wind power and the m hour total predicted photovoltaic power through the k hour wind power segmented predicted power and the m hour photovoltaic power generation power, and determine the k hour actual wind power generation power and the m hour actual photovoltaic power generation power corresponding to the target power system, so as to calculate the wind power prediction power error according to the k hour actual wind power generation power and the k hour total predicted wind power, and calculate the photovoltaic power generation prediction power error through the m hour actual photovoltaic power generation power and the m hour total predicted photovoltaic power; respectively judge whether the wind power prediction power error and the photovoltaic power generation prediction power error meet the preset error requirements, wherein, when the wind power prediction power error and the photovoltaic power generation prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, power generation cost coefficient, and the k hour total predicted wind power, calculate the wind power generation cost, and calculate the photovoltaic power generation cost by using the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient, and the m hour total predicted photovoltaic power.

[0009] Optionally, in an embodiment of the present application, based on the pre-constructed thermal power generation model and energy storage system model, and combining the segmented prediction data, the wind power generation cost, and the photovoltaic power generation cost, constructing a double deep deterministic policy gradient optimization model includes: calculating the hourly thermal power generation value of the target power system, and constructing a thermal power output constraint based on the hourly thermal power generation value, the preset hourly small thermal power generation value, and the hourly maximum thermal power generation; obtaining the n th hour thermal power generation power and the n- 1 hour thermal power generation power, and according to the n th hour thermal power generation power, the n-The power change rate constraint of the thermal power unit is determined by the thermal power generation power in 1 hour and the preset ramp rate; the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, start-up cost coefficient and start-up state variable are determined, and the thermal power generation cost is calculated based on the thermal power generation operation and maintenance cost coefficient, the fuel cost coefficient, the start-up cost coefficient, the start-up state variable and the hourly thermal power generation value, and the thermal power unit emission constraint is constructed according to the preset emission coefficient and the maximum allowable emission; based on the thermal power generation cost, the thermal power unit emission constraint, the thermal power unit power change rate constraint and the above, the thermal power generation model is constructed; the charging efficiency, discharging efficiency, charging power and discharging power of the target power system are obtained, and the charge-discharge power of the energy storage system is calculated based on the charging efficiency, the discharging efficiency, the charging power and the discharging power; based on the preset maximum charging power of the energy storage system, the maximum discharging power of the energy storage system, the charging efficiency and the discharging efficiency, the corresponding charge-discharge power constraint and charge-discharge efficiency constraint are determined; the energy storage system capacity is calculated through the charge-discharge power of the energy storage system, and the energy storage system capacity constraint is constructed according to the energy storage system capacity, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system; based on the preset charge-discharge cost coefficient of the energy storage system, the operation and maintenance cost coefficient of the energy storage system and the charge-discharge power of the energy storage system, the energy storage system cost is calculated, and the energy storage system model is constructed through the energy storage system cost, the energy storage system capacity constraint, the charge-discharge power constraint and the charge-discharge efficiency constraint.

[0010] Optionally, in an embodiment of the present application, collecting the historical environmental data of the target power system, and training the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model through the historical environmental data of the power system to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, including: collecting the historical environmental data of the target power system, and determining the first state space, the first reward function and the first system historical cost corresponding to the first deep deterministic policy gradient agent; determining the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, constructing a first action space through the parameters to be optimized, and training the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function and the first system historical cost, and combining the historical environmental data of the power system to obtain the agent target parameters; determining the second state space, the second action space, the second reward function and the second system cost corresponding to the second deep deterministic policy gradient agent, and optimizing the operation of the target power system based on the second state space, the second action space, the second reward function and the second system cost, and combining the agent target parameters.

[0011] Optionally, in an embodiment of the present application, the k mathematical expression of the segmented predicted power of the wind power in the

[0012]

[0013] wherein, represents the predicted wind speed in the k segment; represents the segmented predicted power of the wind power in the k segment; represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; represents the temperature; represents the radius of the fan blade; represents the wind direction correction coefficient; represents the included angle between the wind direction and the fan orientation; , , respectively represent the cut-in wind speed, the rated wind speed and the cut-out wind speed.

[0014] In the second aspect of the present application, an embodiment provides an optimized operation device for a receiving-end power system based on multi-agent, including: a region division module, configured to determine the circuit topology structure corresponding to the target power system, extract the node features and edge features corresponding to the circuit topology structure, and based on the node features, the edge features and a preset spectral clustering algorithm, divide the target power system into regions to generate a power system region division result; a segmented prediction module, configured to, based on a pre-constructed power system power balance model, make each regional node in the power system region division result satisfy a preset power balance constraint condition, and obtain the wind speed data and irradiance data corresponding to the target power system, and based on a pre-constructed wind and light prediction model, and in combination with the wind speed data and the irradiance data, perform segmented prediction processing to obtain corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; a modeling module, configured to, based on a pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost, construct a double deep deterministic policy gradient optimization model; an optimization module, configured to collect the power system historical environment data of the target power system, and train the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation operation of the target power system according to the agent target parameters.

[0015] Optionally, in an embodiment of the present application, the region division module includes: a first determination unit, configured to determine the node set, edge set and time scale of the transmission line power data of the circuit topology structure corresponding to the target power system; an extraction unit, configured to obtain the adjacent node sets of the nodes i and nodes j in each layer of the circuit topology structure, and based on a preset learnable weight matrix, activation function and the adjacent node sets, extract the node features of each node in the circuit topology structure, where i and j are positive integers; a feature splicing unit, configured to, based on the edge set and a preset multi-layer perceptron, perform a feature splicing operation on the time scale of the transmission line power data, the nodes i and the nodes j in the node features in each layer to obtain the edge features of each edge in the circuit topology structure; a first calculation unit, configured to, based on the node features and a preset Gaussian kernel function bandwidth parameter, calculate the nodes i and the nodes jThe similarity between them is calculated, and a corresponding similarity matrix is constructed based on the similarity. The target power system is partitioned according to the similarity matrix and the spectral clustering algorithm to obtain a regional partition Laplacian matrix. A feature decomposition unit is configured to perform feature decomposition on the regional partition Laplacian matrix to obtain eigenvectors corresponding to the first q minimum eigenvalues, and perform K-means clustering on the eigenvectors to generate the power system regional partition result, where q is a positive integer.

[0016] Optionally, in an embodiment of the present application, the segmented prediction module includes: a first construction unit configured to determine the power flowing through each node per hour in each region node of the power system regional partition result, and construct a line power constraint based on a preset maximum power flowing through the node per hour, a minimum power flowing through the node per hour, and the power flowing through the node per hour; a second determination unit configured to determine the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour corresponding to each region node, and calculate the total multi-energy power generation per hour according to the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour; a second construction unit configured to calculate the total line load per hour according to the power flowing through the node per hour, and construct the power balance constraint condition based on the total multi-energy power generation per hour and the total line load per hour, so that each region node satisfies the power balance constraint condition.

[0017] Optionally, in an embodiment of the present application, the segmented prediction module further includes: a second calculation unit configured to calculate the mean and standard deviation of the wind speed data and the irradiance data, and calculate the preprocessed wind speed data and the preprocessed photovoltaic power generation data corresponding to the wind speed data and the irradiance data respectively according to the mean and the standard deviation; a third construction unit configured to construct the wind-solar prediction model based on a preset segmented dynamic trend decomposition prediction model, and perform segmented prediction on the preprocessed wind speed data and the preprocessed photovoltaic power generation data through the wind-solar prediction model to obtain the predicted wind speed of the k th segment and the predicted irradiance of the m th segment, where k and m are positive integers; a first acquisition unit configured to acquire the cut-in wind speed, the rated wind speed, the cut-out wind speed, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance, and the maximum irradiance corresponding to the target power system, and calculate the segmented predicted wind power of the k th segment based on the cut-in wind speed, the rated wind speed, the cut-out wind speed, and the predicted wind speed of the k th segment, and pass the mCalculate the predicted irradiance, the efficiency of the photovoltaic module, the area of the photovoltaic module, the minimum irradiance and the maximum irradiance to calculate the m section of photovoltaic power generation; a third calculation unit for respectively passing through the k section of wind power segmented prediction power and the m section of photovoltaic power generation to calculate k the total predicted wind power within the hour and m the total predicted photovoltaic power within the hour, and determine the corresponding k actual wind power generation power within the hour and m actual photovoltaic power generation power within the hour, so as to calculate the wind power prediction power error according to the k actual wind power generation power within the hour and the k total predicted wind power within the hour, and calculate the photovoltaic power generation prediction power error through the m actual photovoltaic power generation power within the hour and the m total predicted photovoltaic power within the hour; a judgment unit for respectively judging whether the wind power prediction power error and the photovoltaic power generation prediction power error meet the preset error requirements, wherein, when the wind power prediction power error and the photovoltaic power generation prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, power generation cost coefficient and the k total predicted wind power within the hour, calculate the wind power generation cost, and use the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and the m total predicted photovoltaic power within the hour to calculate the photovoltaic power generation cost.

[0018] Optionally, in an embodiment of the present application, the modeling module includes: a fourth calculation unit for calculating the hourly thermal power generation value of the target power system, and constructing a thermal power output constraint based on the hourly thermal power generation value, the preset hourly small thermal power generation value and the hourly maximum thermal power generation value; a second acquisition unit for acquiring the n hour's thermal power generation power and the n- (n + 1)th hour's thermal power generation power, and according to the n hour's thermal power generation power, the n-The power change rate constraint of the thermal power unit is determined by the thermal power generation power in 1 hour and the preset ramp rate; a third determination unit is configured to determine the thermal power generation operation and maintenance cost coefficient, the fuel cost coefficient, the start-up cost coefficient and the start-up state variable, and calculate the thermal power generation cost based on the thermal power generation operation and maintenance cost coefficient, the fuel cost coefficient, the start-up cost coefficient, the start-up state variable and the hourly thermal power generation value, and construct the thermal power unit emission constraint according to the preset emission coefficient and the maximum allowable emission; a construction unit is configured to construct the thermal power generation model based on the thermal power generation cost, the thermal power unit emission constraint, the thermal power unit power change rate constraint and the ; a third acquisition unit is configured to acquire the charging efficiency, the discharging efficiency, the charging power and the discharging power of the target power system, and calculate the charge-discharge power of the energy storage system based on the charging efficiency, the discharging efficiency, the charging power and the discharging power; a fourth determination unit is configured to determine the corresponding charge-discharge power constraint and the charge-discharge efficiency constraint based on the preset maximum charging power of the energy storage system, the maximum discharging power of the energy storage system, the charging efficiency and the discharging efficiency; a fourth construction unit is configured to calculate the energy storage system capacity through the charge-discharge power of the energy storage system, and construct the energy storage system capacity constraint according to the energy storage system capacity, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system; a fifth calculation unit is configured to calculate the energy storage system cost based on the preset charge-discharge cost coefficient of the energy storage system, the operation and maintenance cost coefficient of the energy storage system and the charge-discharge power of the energy storage system, and construct the energy storage system model through the energy storage system cost, the energy storage system capacity constraint, the charge-discharge power constraint and the charge-discharge efficiency constraint.

[0019] Optionally, in an embodiment of the present application, the optimization module includes: a collection unit configured to collect the power system historical environment data of the target power system, and determine the first state space, the first reward function and the first system historical cost corresponding to the first deep deterministic policy gradient agent; a training unit configured to determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, construct a first action space through the parameters to be optimized, and train the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function and the first system historical cost, and in combination with the power system historical environment data, to obtain the target parameters of the agent; a fifth determination unit configured to determine the second state space, the second action space, the second reward function and the second system cost corresponding to the second deep deterministic policy gradient agent, and optimize the operation of the target power system based on the second state space, the second action space, the second reward function and the second system cost, and in combination with the target parameters of the agent.

[0020] Optionally, in an embodiment of the present application, thek The mathematical expression for the segmented predicted power of wind power is as follows:

[0021]

[0022] Wherein, represents the predicted wind speed of the k th segment; represents the segmented predicted power of wind power of the k th segment; represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; represents the air temperature; represents the radius of the wind turbine blade; represents the wind direction correction coefficient; represents the angle between the wind direction and the orientation of the wind turbine; , , respectively represent the cut-in wind speed, the rated wind speed and the cut-out wind speed mentioned above.

[0023] The third aspect of the embodiments of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the method for optimizing the operation of the receiving-end power system based on multi-agent as described in the above embodiments.

[0024] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the method for optimizing the operation of the receiving-end power system based on multi-agent as described above.

[0025] Therefore, the embodiments of the present application have the following beneficial effects:

[0026] Embodiments of the present application can generate a power system regional division result by determining the circuit topology corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology, and performing regional division on the target power system based on the node features, edge features, and a preset spectral clustering algorithm; based on the pre-constructed power system power balance model, each regional node in the power system regional division result satisfies the preset power balance constraint conditions, and the wind speed data and irradiance data corresponding to the target power system are obtained, so as to perform segmented prediction processing based on the pre-constructed wind-solar prediction model and in combination with the wind speed data and irradiance data to obtain corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; based on the pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, wind power generation cost, and photovoltaic power generation cost, a double deep deterministic policy gradient optimization model is constructed; the power system historical environment data of the target power system is collected, and the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model is trained through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for the optimal operation of the power system multi-regional node power balance, helping to optimize power dispatching, and reducing the grid operation cost. Thus, the problems in the prior art that it is difficult to effectively handle complex situations of multiple regions and multiple nodes, and it is difficult to take into account parameter selection and system optimization at the same time, greatly affecting the optimization effect, etc. are solved.

[0027] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0029] Figure 1 FIG. is a flowchart of a method for optimizing the operation of a receiving-end power system based on multi-agents according to an embodiment of the present application;

[0030] Figure 2 FIG. is an execution logic schematic diagram of a method for optimizing the operation of a receiving-end power system based on multi-agents provided by an embodiment of the present application;

[0031] Figure 3 FIG. is an example diagram of a device for optimizing the operation of a receiving-end power system based on multi-agents according to an embodiment of the present application;

[0032] Figure 4 Schematic structural diagram of the electronic device provided by the embodiment of the present application.

[0033] Among them, 10 - Multi-agent based receiving-end power system optimal operation device; 100 - Region division module, 200 - Segment prediction module, 300 - Modeling module, 400 - Optimization module; 401 - Memory, 402 - Processor, 403 - Communication interface. Specific implementation manners

[0034] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.

[0035] The multi-agent based receiving-end power system optimal operation method and device according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the problems mentioned in the above background art, the present application provides a multi-agent based receiving-end power system optimal operation method. In this method, by determining the circuit topology structure corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology structure, and based on the node features, edge features and a preset spectral clustering algorithm, the target power system is regionally divided to generate a power system regional division result; based on the pre-constructed power system power balance model, each regional node in the power system regional division result satisfies the preset power balance constraint condition, and the wind speed data and irradiance data corresponding to the target power system are obtained, so as to perform segment prediction processing based on the pre-constructed wind-solar prediction model and in combination with the wind speed data and irradiance data to obtain corresponding segment prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segment prediction data; based on the pre-constructed thermal power generation model and energy storage system model, and in combination with the segment prediction data, wind power generation cost and photovoltaic power generation cost, a double deep deterministic policy gradient optimization model is constructed; the power system historical environment data of the target power system is collected, and the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model is trained through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for the optimal operation of the power system multi-regional node power balance, helping to optimize power dispatching and reduce the grid operation cost. Thus, the problems in the prior art that it is difficult to effectively cope with the complex situations of multiple regions and multiple nodes, and it is difficult to take into account parameter selection and system optimization at the same time, greatly affecting the optimization effect, etc. are solved.

[0036] Specifically, Figure 1 FIG. 1 is a flowchart of an optimal operation method for a receiving-end power system based on multi-agent provided by an embodiment of the present application.

[0037] As Figure 1 shown, the optimal operation method for the receiving-end power system based on multi-agent includes the following steps:

[0038] In step S101, determine the circuit topology structure corresponding to the target power system, extract the node features and edge features corresponding to the circuit topology structure, and based on the node features, edge features, and a preset spectral clustering algorithm, perform regional division on the target power system to generate a power system regional division result.

[0039] In the embodiment of the present application, a multi-region division model of the power system can be established first, so as to use the feature extraction method based on the topology structure through the multi-region division model of the power system to extract the topology features and power flow features of the power system, and combine the spectral clustering algorithm to realize regional division.

[0040] Optionally, in an embodiment of the present application, determining the circuit topology structure corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology structure, and based on the node features, edge features, and a preset spectral clustering algorithm, performing regional division on the target power system to generate a power system regional division result includes: determining the node set, edge set, and time scale of the transmission line power data of the circuit topology structure corresponding to the target power system; obtaining the node i and node j adjacent node sets in each layer of the circuit topology structure, and based on a preset learnable weight matrix, activation function, and adjacent node sets, extracting the node features of each node in the circuit topology structure, where i and j are positive integers; based on the edge set and a preset multi-layer perceptron, performing a feature splicing operation on the time scale of the transmission line power data, node i and node j node features in each layer to obtain the edge features of each edge in the circuit topology structure; based on the node features and a preset Gaussian kernel function bandwidth parameter, calculating the similarity between node i and node j and constructing a corresponding similarity matrix through the similarity, so as to perform regional division on the target power system according to the similarity matrix and the spectral clustering algorithm to obtain a regional division Laplacian matrix; performing eigenvalue decomposition on the regional division Laplacian matrix to obtain the eigenvectors corresponding to the first q smallest eigenvalues, and performing K-means clustering on the eigenvectors to generate a power system regional division result, where q is a positive integer.

[0041] Specifically, the process of multi - area division of the power system in the embodiments of this application is as follows:

[0042] 1. Data pre - processing:

[0043] In the embodiments of this application, the node set of the circuit topology structure is , and the edge set is ; the power data of each transmission line is modeled on an hourly scale as ; the node feature matrix is , is a matrix over the real number field; a weighted adjacency matrix , is a matrix over the real number field, where represents the power flow between nodes.

[0044] 2. Extract node features and edge features through a feature extraction method based on the topological structure:

[0045] Initialize the node feature matrix and the adjacency matrix , and update the node features and edge features respectively:

[0046] 1) Node feature update. For each layer there is:

[0047] (1)

[0048] Among them, represents the feature of node at the th layer; represents the feature of node at the th layer; represents the set of adjacent nodes of node ; represents the set of adjacent nodes of node ; is a learnable weight matrix; represents the activation function.

[0049] 2) Edge feature update. For each edge there is:

[0050] (2)

[0051] Among them, represents the feature of the edge at the th layer; || represents feature concatenation; MLP represents a multi - layer perceptron, represents the power data of the transmission line.

[0052] Thus, after passing through the L-layer graph neural network in the embodiments of this application, a node feature matrix can be obtained , where is a matrix in the real number field , and is the final feature dimension

[0053] 3. Based on the extracted node features and edge features, the embodiments of this application can use the spectral clustering algorithm to divide the power system into regions:

[0054] (1) Calculate the similarity between node and node :

[0055] (3)

[0056] where represents the similarity between node and node ; and represent the extracted node features, represents the bandwidth parameter of the Gaussian kernel function

[0057] (2) Determine the similarity matrix according to the similarity obtained above, where is a matrix in the real number field ; Based on the similarity matrix, the embodiments of this application can use the spectral clustering algorithm to divide the power system into regions, which is specifically described as follows:

[0058] 1) Construct the Laplacian matrix for region division:

[0059] (4)

[0060] where is the degree matrix; is the Laplacian matrix

[0061] 2) Perform eigenvalue decomposition on the Laplacian matrix for region division to obtain the eigenvectors corresponding to the first q smallest eigenvalues, and perform K-means clustering on the eigenvectors , where is a matrix in the real number field to obtain the final power system region division result

[0062] Thus, the embodiments of this application extract the relevant features of the power system and combine the spectral clustering algorithm to achieve region division, providing a reliable theoretical basis for the subsequent generation of wind-solar intelligent scenarios and the realization of power balance in the receiving-end power system

[0063] In step S102, based on the pre-constructed power balance model of the power system, each regional node in the power system regional division result is made to satisfy the preset power balance constraint condition, and the wind speed data and irradiance data corresponding to the target power system are obtained, so as to perform segmented prediction processing based on the pre-constructed wind and light prediction model, in combination with the wind speed data and irradiance data, to obtain the corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data.

[0064] Furthermore, the embodiment of the present application can also establish a power balance model of the power system to ensure that the load demand of each node is equal to the sum of the power generated by all lines connected to the node, satisfying the power balance constraint condition; thereafter, the embodiment of the present application can establish a wind and light prediction model based on the segmented dynamic trend decomposition prediction model to preprocess the wind speed and photovoltaic data, and perform segmented prediction on the preprocessed wind speed and photovoltaic data, so as to ensure that the predicted power error of wind power generation and photovoltaic power generation is within the allowable range.

[0065] Optionally, in an embodiment of the present application, based on the pre-constructed power balance model of the power system, making each regional node in the power system regional division result satisfy the preset power balance constraint condition includes: determining the power flowing through the node per hour of each regional node in the power system regional division result, and constructing a line power constraint based on the preset maximum power flowing through the node per hour, the minimum power flowing through the node per hour, and the power flowing through the node per hour; determining the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour corresponding to each regional node, and calculating the total multi-energy power generation per hour according to the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour; calculating the total line load per hour according to the power flowing through the node per hour, and constructing a power balance constraint condition based on the total multi-energy power generation per hour and the total line load per hour, so that each regional node satisfies the power balance constraint condition.

[0066] In the actual execution process, the embodiment of the present application can assume that the power system has regional nodes, and the power flowing through the node per hour is (from node to node ), then the specific constraints are as follows:

[0067] The line power constraint is:

[0068] (5)

[0069] Among them, is the maximum power flowing through the node within one hour; It is the minimum power flowing through the node within one hour.

[0070] The total line load within each hour is:

[0071] (6)

[0072] Wherein, is the total line load within each hour.

[0073] The total power generation of multiple energy sources within each hour is:

[0074] (7)

[0075] Wherein, , , and are respectively the wind power generation power, photovoltaic power generation power, thermal power generation power and the power taken out by the energy storage system within each hour; is the total power generation of multiple energy sources within each hour.

[0076] Therefore, the power balance equation of the embodiment of the present application is:

[0077] (8)

[0078] Optionally, in an embodiment of the present application, wind speed data and irradiance data corresponding to the target power system are obtained, and based on a pre-constructed wind-solar prediction model, combined with the wind speed data and irradiance data, segmented prediction processing is performed to obtain corresponding segmented prediction data, and the corresponding wind power generation cost and photovoltaic power generation cost are calculated according to the segmented prediction data, including: calculating the mean and standard deviation of the wind speed data and irradiance data, and respectively calculating the preprocessed wind speed data and preprocessed photovoltaic power generation data corresponding to the wind speed data and irradiance data according to the mean and standard deviation; based on a preset segmented dynamic trend decomposition prediction model, constructing a wind-solar prediction model, and performing segmented prediction on the preprocessed wind speed data and preprocessed photovoltaic power generation data through the wind-solar prediction model to obtain the k th segment predicted wind speed and the m th segment predicted irradiance, wherein, k and m are positive integers; obtaining the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiance corresponding to the target power system, and based on the cut-in wind speed, rated wind speed, cut-out wind speed and the k th segment predicted wind speed, calculating the k th segment wind power segmented prediction power, and calculating the m th segment predicted irradiance, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiancem Segment photovoltaic power generation; respectively through the k Segment wind power segmented prediction power and the m Segment photovoltaic power generation power calculation k Total predicted wind power within hours and m Total predicted photovoltaic power within hours, and determine the corresponding k Actual wind power generation power within hours and m Actual photovoltaic power generation power within hours, in order to k Actual wind power generation power within hours and k Total predicted wind power within hours to calculate the wind power prediction error, and through m Actual photovoltaic power generation power within hours and m Total predicted photovoltaic power within hours to calculate the photovoltaic power generation prediction error; respectively determine whether the wind power prediction error and the photovoltaic power generation prediction error meet the preset error requirements, where, when the wind power prediction error and the photovoltaic power generation prediction error meet the error requirements, based on the preset operation and maintenance cost coefficient, power generation cost coefficient and k Total predicted wind power within hours, calculate the wind power generation cost, and use the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and m Total predicted photovoltaic power within hours to calculate the photovoltaic power generation cost.

[0079] In the specific implementation process, the embodiments of the present application can obtain the wind speed data and irradiance data (i.e., historical data of photovoltaic power generation) corresponding to the power system, and through the wind and light prediction model, combine the wind speed data and irradiance data to perform segmented prediction processing to obtain the corresponding segmented prediction data (including segmented predicted wind power and photovoltaic power generation power), so as to calculate the corresponding wind power generation cost and photovoltaic power generation cost using the segmented prediction data, and the specific process is as follows:

[0080] Segmented prediction of wind power generation:

[0081] (1) Preprocess the wind speed data to make it meet the model input requirements

[0082] (9)

[0083] Among them, is the preprocessed wind speed data; is the original wind speed data (i.e., wind speed data); is the mean of the wind speed data; is the standard deviation of the wind speed data.

[0084] (2) Use the wind and light prediction model of the segmented dynamic trend decomposition prediction model to predict the wind speed of each segment, as shown in the following formula:

[0085] (10)

[0086] Among them, and are the power inertia coefficient and the power fluctuation coefficient respectively; represents the predicted section wind speed; represents the lag operator; represents the d power of the lag operator; represents white noise; q represents the moving average order; represents the autoregressive order.

[0087] Optionally, in an embodiment of the present application, the mathematical expression of the segmented predicted power of wind power in the k section is:

[0088]

[0089] Among them, represents the predicted wind speed in the k section; represents the segmented predicted power of wind power in the k section; represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; represents the air temperature; represents the radius of the wind turbine blade; represents the wind direction correction coefficient; represents the angle between the wind direction and the orientation of the wind turbine; , , represent the cut-in wind speed, the rated wind speed and the cut-out wind speed respectively.

[0090] It should be noted that based on Equation (10), the segmented wind power generation power (i.e., the segmented predicted power of wind power in the section) in the embodiment of the present application is:

[0091] (11)

[0092] Among them, represents the predicted wind speed in the k section (i.e., the predicted section wind speed); represents the segmented predicted power of wind power in the k section (i.e., the segmented predicted power of wind power in the section); represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; represents the air temperature; represents the radius of the wind turbine blade; represents the wind direction correction coefficient; represents the angle between the wind direction and the orientation of the wind turbine; 、 、 represent the cut-in wind speed, rated wind speed, and cut-out wind speed respectively.

[0093] Therefore, through Equation (11), the total predicted wind power within K hours can be obtained as: is:

[0094] (12)

[0095] Furthermore, the embodiments of the present application can calculate the wind power prediction error based on the total predicted wind power within K hours, as shown in the following equation: is:

[0096] (13)

[0097] Wherein, is the actual wind power generation within

[0098] hours. After that, the embodiments of the present application can determine whether the wind power prediction error

[0099] (14)

[0100] is within the allowable error range according to the following equation: where is the maximum value of the wind power prediction error.

[0101] Finally, when the wind power prediction error is within the allowable error range as shown in Equation (14), the embodiments of the present application can calculate the wind power generation cost based on the total predicted wind power within

[0102] K hours, as shown in the following equation:

[0103] Wherein, is the wind power generation cost; is the operation and maintenance cost coefficient; is the power generation cost coefficient.

[0104] Segmented prediction of photovoltaic power generation:

[0105] It should be noted that in the embodiments of the present application, a segmented prediction model is adopted for the photovoltaic power generation, and it is updated hourly. The specific process is as follows:

[0106] (1) Standardize the historical data of photovoltaic power generation (i.e., irradiance data) to meet the model input requirements. The standardization processing expression is:

[0107] (16)

[0108] Where, is the preprocessed data of photovoltaic power generation; is the original irradiance data (i.e., irradiance data); represents the mean value of the irradiance data; is the standard deviation of the irradiance data.

[0109] (2) Use the wind-solar prediction model of the segmented dynamic trend decomposition prediction model to predict the irradiance of each segment, as shown in the following formula:

[0110] (17)

[0111] Where, and are the power inertia coefficient and the power fluctuation coefficient respectively; represents the lag operator; represents the -th segment predicted irradiance; represents white noise; represents the order.

[0112] From formula (17), the -th segment photovoltaic power generation power is:

[0113] (18)

[0114] Where, represents the -th segment predicted photovoltaic power generation power; represents the photovoltaic module efficiency; represents the area of the photovoltaic module; and represent the minimum value and the maximum value of the irradiance respectively; represents the -th segment predicted irradiance; represents the rated irradiance intensity; represents the solar altitude angle (in radians); represents the solar azimuth angle (in radians); represents the tilt angle of the photovoltaic panel (in radians); represents the azimuth angle of the photovoltaic panel (in radians); represents the The temperature of the photovoltaic panel section; Indicates the reference temperature; Indicates the temperature coefficient.

[0115] Therefore, the embodiment of the present application can calculate the total predicted photovoltaic power within the hour through the photovoltaic power generation of the section, as shown in the following formula:

[0116] (19)

[0117] Wherein, is the total predicted photovoltaic power for segmented prediction; is the section of the predicted photovoltaic power generation.

[0118] (3) Calculate the predicted photovoltaic power generation error based on the total predicted photovoltaic power within the hour:

[0119] (20)

[0120] Wherein, is the hour of the actual photovoltaic power generation; is the predicted photovoltaic power generation error.

[0121] Use the predicted photovoltaic power generation error to determine the allowable range of the predicted photovoltaic power error:

[0122] (21)

[0123] Wherein, is the predicted photovoltaic power generation error; is the maximum value of the predicted photovoltaic power generation error.

[0124] After that, when the predicted photovoltaic power generation error is within the allowable error range shown in formula (21), the embodiment of the present application can calculate the photovoltaic power generation cost based on the total predicted photovoltaic power for segmented prediction, as shown in the following formula:

[0125] (22)

[0126] Wherein, is the photovoltaic power generation cost; is the photovoltaic operation and maintenance cost coefficient; is the photovoltaic power generation cost coefficient.

[0127] Thus, the embodiments of the present application preprocess the wind speed and photovoltaic data and perform segmented prediction, thereby ensuring that the prediction power errors of wind power generation and photovoltaic power generation are within the allowable range, providing important theoretical guidance and basis for the construction of the double deep deterministic policy gradient optimization model.

[0128] In step S103, based on the pre-constructed thermal power generation model and energy storage system model, and combined with the segmented prediction data, wind power generation cost, and photovoltaic power generation cost, a double deep deterministic policy gradient optimization model is constructed.

[0129] Furthermore, the embodiments of the present application also need to establish a thermal power generation model based on output range, ramp rate, and power generation cost constraints, etc., and establish an energy storage system model through charge and discharge power, energy storage capacity, and efficiency constraints, etc. After that, the embodiments of the present application can construct a double deep deterministic policy gradient optimization model with shared parameters through the thermal power generation model and the energy storage system model, combined with segmented prediction data (including wind power segmented prediction power, photovoltaic power generation power, etc.), wind power generation cost, and photovoltaic power generation cost.

[0130] Optionally, in an embodiment of the present application, constructing a double deep deterministic policy gradient optimization model based on the pre-constructed thermal power generation model and energy storage system model, and combined with segmented prediction data, wind power generation cost, and photovoltaic power generation cost, includes: calculating the hourly thermal power generation value of the target power system, and constructing a thermal power output constraint based on the hourly thermal power generation value, the preset hourly small thermal power generation value, and the hourly maximum thermal power generation value; obtaining the thermal power generation power of the n th hour and the thermal power generation power of the n- (1)st hour, and according to the thermal power generation power of the n th hour, the thermal power generation power of the n-Determine the power change rate constraint of the thermal power unit based on the 1-hour thermal power generation power and the preset ramp rate; determine the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, start-up cost coefficient, and start-up state variable, and calculate the thermal power generation cost based on the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, start-up cost coefficient, start-up state variable, and the hourly thermal power generation value. And construct the thermal power unit emission constraint according to the preset emission coefficient and the maximum allowable emission; construct the thermal power generation model based on the thermal power generation cost, the thermal power unit emission constraint, and the thermal power unit power change rate constraint; obtain the charging efficiency, discharging efficiency, charging power, and discharging power of the target power system, and calculate the charge-discharge power of the energy storage system based on the charging efficiency, discharging efficiency, charging power, and discharging power; determine the corresponding charge-discharge power constraint and charge-discharge efficiency constraint based on the preset maximum charging power of the energy storage system, the maximum discharging power of the energy storage system, the charging efficiency, and the discharging efficiency; calculate the energy storage system capacity through the charge-discharge power of the energy storage system, and construct the energy storage system capacity constraint according to the energy storage system capacity, the preset minimum capacity of the energy storage system, and the maximum capacity of the energy storage system; calculate the energy storage system cost based on the preset charge-discharge cost coefficient of the energy storage system, the operation and maintenance cost coefficient of the energy storage system, and the charge-discharge power of the energy storage system, and construct the energy storage system model through the energy storage system cost, the energy storage system capacity constraint, the charge-discharge power constraint, and the charge-discharge efficiency constraint.

[0131] It should be noted that the process of establishing the thermal power generation model and the energy storage system model in the embodiments of this application is as follows:

[0132] 1. Establish a thermal power generation model based on the output range, ramp rate, and power generation cost constraint, as shown in the following formula:

[0133] (23)

[0134] Among them, is the hourly thermal power generation value; is the control input of the thermal power unit; , , are the coefficients of the thermal power output formula, and different values can be taken according to the operating conditions.

[0135] The thermal power output constraint is:

[0136] (24)

[0137] Among them, is the hourly thermal power generation value; is the maximum hourly thermal power generation value.

[0138] In the embodiments of this application, the power change rate of the thermal power unit is limited by the ramp rate, as shown in the following formula:

[0139] (25)

[0140] Wherein, represents the thermal power generation power in the th hour, represents the thermal power generation power in the th hour; is the ramp rate; is the time interval of one hour.

[0141] After that, the embodiment of the present application can calculate the thermal power generation cost according to the hourly thermal power generation value:

[0142] (26)

[0143] Wherein, is the thermal power generation cost; is the thermal power generation operation and maintenance cost coefficient; is the fuel cost coefficient; is the start-up cost coefficient; is the start-up state variable, and the start-up state variable is shown as follows:

[0144] (27)

[0145] In addition, the emission constraint of the thermal power unit in the embodiment of the present application is:

[0146] (28)

[0147] (29)

[0148] Wherein, is the emission amount; , , are the emission coefficients; is the maximum allowable emission amount.

[0149] 2. Establish a energy storage system model based on the charge and discharge power, energy storage capacity and efficiency constraints, etc.:

[0150] In the embodiment of the present application, the charge and discharge power of the energy storage system is:

[0151] (30)

[0152] Wherein, is the charging efficiency; is the discharging efficiency; is the power during charging; is the power during discharging; is the charge and discharge power of the energy storage system.

[0153] Determine the charge-discharge power constraint based on the power during charging and the power during discharging, as shown in the following formula:

[0154] (31)

[0155] (32)

[0156] Among them, represents the maximum charging power of the energy storage system; represents the maximum discharging power of the energy storage system.

[0157] In addition, the charge-discharge efficiency constraint in the embodiments of the present application is:

[0158] (33)

[0159] (34)

[0160] Among them, is the charging efficiency; is the discharging efficiency.

[0161] Subsequently, the embodiments of the present application can calculate the capacity of the energy storage system based on the charge-discharge power of the energy storage system, as shown in the following formula:

[0162] (35)

[0163] Among them, is the capacity of the energy storage system at ; is the capacity of the energy storage system at ; is the time interval of one hour.

[0164] Use the above energy storage system capacity to determine the energy storage system capacity constraint:

[0165] (36)

[0166] Among them, is the minimum capacity of the energy storage system; is the maximum capacity of the energy storage system.

[0167] Finally, the embodiments of the present application can calculate the cost of the energy storage system according to the charge-discharge power of the energy storage system, as shown in the following formula:

[0168] (37)

[0169] Among them, is the cost of the energy storage system; is the charge-discharge cost coefficient of the energy storage system; is the operation and maintenance cost coefficient of the energy storage system.

[0170] Therefore, the embodiments of the present application build a thermal power generation model and an energy storage system model, thereby providing data support for the construction of the double deep deterministic policy gradient optimization model, and ensuring the smooth realization of the wind-solar intelligent scenario generation and power balance of the receiving-end power system.

[0171] In step S104, the historical environmental data of the target power system is collected, and the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model is trained through the historical environmental data of the power system to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

[0172] After that, the embodiments of the present application can train the first deep deterministic policy gradient agent (i.e., the first deep deterministic policy gradient agent) through the historical environmental data (i.e., the historical environmental data of the power system) to select the corresponding optimal agent parameters (i.e., the agent target parameters), and after obtaining the optimal agent parameters, use the second deep deterministic policy gradient agent (i.e., the second deep deterministic policy gradient agent) to interact with the environment, thereby optimizing the operation of the power system.

[0173] Therefore, the embodiments of the present application can effectively cope with the complexity and uncertainty of multi-region nodes of the power system, and improve the stability and economy of the power system.

[0174] Optionally, in an embodiment of the present application, power system historical environment data of the target power system is collected, and the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model is trained through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, including: collecting the power system historical environment data of the target power system, and determining the first state space, the first reward function, and the first system historical cost corresponding to the first deep deterministic policy gradient agent; determining the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, constructing the first action space through the parameters to be optimized, and training the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function, and the first system historical cost, and combining the power system historical environment data to obtain the agent target parameters; determining the second state space, the second action space, the second reward function, and the second system cost corresponding to the second deep deterministic policy gradient agent, and optimizing the operation of the target power system based on the second state space, the second action space, the second reward function, and the second system cost, and combining the agent target parameters.

[0175] It should be noted that in the embodiment of the present application, first, the power system historical environment data of the target power system (i.e., the receiving-end power system) can be collected, and the first agent state space (i.e., the first state space), the first agent action space (i.e., the first action space), the first agent reward function (i.e., the first reward function), and the system cost (i.e., the first system cost) corresponding to the first deep deterministic policy gradient agent are determined.

[0176] Specifically, the above-mentioned first agent state space is:

[0177] (38)

[0178] Among them, is historical load data, is historical wind power, is historical wind and solar power generation prediction data, is historical thermal power output data, is the charge and discharge data of the historical energy storage system.

[0179] In the embodiment of the present application, the first agent action space is constructed through the discount factor and the time step (i.e., the parameters to be optimized) :

[0180] (39)

[0181] Among them, is the discount factor; is the time step.

[0182] The reward function of the first agent is:

[0183] (40)

[0184] The system cost is:

[0185] (41)

[0186] Among them, , , , are the historical cost of wind power generation, the historical cost of photovoltaic power generation, the historical cost of thermal power generation, and the historical cost of the energy storage system, respectively.

[0187] After that, the embodiment of the present application can train the first deep deterministic policy gradient agent through the historical environmental data of the power system, and in combination with the state space of the first agent, the action space of the first agent, the reward function of the first agent, and the system cost, to obtain the agent target parameters, that is, the optimal discount factor is and the optimal time step is .

[0188] Furthermore, the embodiment of the present application can use the agent target parameters (i.e., the optimal discount factor , the optimal time step ) selected by the first deep deterministic policy gradient agent through the second deep deterministic policy gradient agent to interact with the environment and optimize the operation of the power system.

[0189] Among them, the state space of the second agent (i.e., the second state space) is:

[0190] (42)

[0191] The action space of the second agent (i.e., the second action space) is:

[0192] (43)

[0193] The reward function of the second agent (i.e., the second reward function) is:

[0194] (44)

[0195] The system cost (i.e., the second system cost) is as follows:

[0196] (45)

[0197] Among them, 、 、 and are the wind power generation cost, photovoltaic power generation cost, thermal power generation cost, and energy storage system cost.

[0198] It can be understood that the embodiments of the present application, through the architecture of the power system multi-region division model, power system power balance model, wind-solar prediction model based on the piecewise dynamic trend decomposition prediction model, thermal power generation and energy storage system model, and the dual deep deterministic policy gradient optimization model with parameter sharing, effectively optimize the balanced power of the complex power system with multiple energy sources. It can not only effectively cope with the complexity and uncertainty of the multi-region nodes of the power system, improve the stability and economy of the power system, but also consider the selection of the parameters of the deep deterministic policy gradient agent, which improves the learning efficiency of the agent and reaches a higher reward function value. In addition, the embodiments of the present application provide a more reliable basis for the optimal operation of the power system multi-region node power balance, which helps to optimize power dispatching and reduces the grid operation cost.

[0199] The following further explains the execution logic of the multi-agent-based receiving-end power system optimal operation method of the present application by combining the drawings.

[0200] Figure 2 is a schematic diagram of the execution logic of the multi-agent-based receiving-end power system optimal operation method of the present application. As Figure 2 shown, the execution steps of the multi-agent-based receiving-end power system optimal operation method of the present application are as follows:

[0201] S201: Establish a power system multi-region division model, extract the topological features and power flow features of the power system through a feature extraction method based on the topological structure, and implement region division in combination with the spectral clustering algorithm;

[0202] S202: Establish a power system power balance model to ensure that the load demand of each node is equal to the sum of the power output from all lines connected to the node, satisfying the power balance constraint conditions;

[0203] S203: Establish a wind-solar prediction model based on the piecewise dynamic trend decomposition prediction model, preprocess the wind speed and photovoltaic data, and perform piecewise prediction to ensure that the predicted power errors of wind power generation and photovoltaic power generation are within the allowable range;

[0204] S204: Establish a model for thermal power generation and energy storage system. Based on constraints such as output range, ramp rate, and generation cost, establish a thermal power generation model. Based on constraints such as charge-discharge power, energy storage capacity, and efficiency, establish an energy storage system model;

[0205] S205: Establish a dual deep deterministic policy gradient optimization model with parameter sharing. Based on historical environmental data, use the first deep deterministic policy gradient agent to train and select the optimal parameters of the agent. After obtaining the optimal parameters, use the second deep deterministic policy gradient agent to interact with the environment to optimize the operation of the power system.

[0206] According to the method for optimizing the operation of a receiving-end power system based on multi-agent proposed in the embodiments of the present application, by determining the circuit topology structure corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology structure, and based on the node features, edge features, and a preset spectral clustering algorithm, dividing the target power system into regions to generate a power system regional division result; based on the pre-constructed power system power balance model, making each regional node in the power system regional division result satisfy the preset power balance constraint conditions, and obtaining the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on the pre-constructed wind power and photovoltaic power prediction models in combination with the wind speed data and irradiance data to obtain corresponding segmented prediction data, and calculating the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; based on the pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, wind power generation cost, and photovoltaic power generation cost, construct a dual deep deterministic policy gradient optimization model; collect the power system historical environmental data of the target power system, and train the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model through the power system historical environmental data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for optimizing the power balance operation of multi-region nodes of the power system, helping to optimize power dispatching, and reducing the operating cost of the power grid.

[0207] Secondly, describe the device for optimizing the operation of a receiving-end power system based on multi-agent proposed in the embodiments of the present application with reference to the accompanying drawings.

[0208] Figure 3 is a block diagram of the device for optimizing the operation of a receiving-end power system based on multi-agent according to the embodiments of the present application.

[0209] As Figure 3As shown in the figure, the multi-agent-based optimal operation device 10 of the receiving-end power system includes: a regional division module 100, a segmented prediction module 200, a modeling module 300, and an optimization module 400.

[0210] Among them, the regional division module 100 is used to determine the circuit topology structure corresponding to the target power system, extract the node features and edge features corresponding to the circuit topology structure, and based on the node features, edge features, and a preset spectral clustering algorithm, perform regional division on the target power system to generate a regional division result of the power system.

[0211] The segmented prediction module 200 is used to, based on a pre-constructed power system power balance model, make each regional node in the regional division result of the power system satisfy a preset power balance constraint condition, and obtain the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on a pre-constructed wind-solar prediction model and in combination with the wind speed data and irradiance data to obtain corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data.

[0212] The modeling module 300 is used to, based on a pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, wind power generation cost, and photovoltaic power generation cost, construct a double deep deterministic policy gradient optimization model.

[0213] The optimization module 400 is used to collect the historical environment data of the power system of the target power system, and train the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model through the historical environment data of the power system to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation operation of the target power system according to the agent target parameters.

[0214] Optionally, in an embodiment of the present application, the regional division module 100 includes: a first determination unit, an extraction unit, a feature splicing unit, a first calculation unit, and a feature decomposition unit.

[0215] Among them, the first determination unit is used to determine the node set, edge set, and time scale of the transmission line power data of the circuit topology structure corresponding to the target power system.

[0216] The extraction unit is used to obtain the adjacent node sets of the nodes i and the nodes j in each layer of the circuit topology structure, and based on a preset learnable weight matrix, activation function, and adjacent node sets, extract the node features of each node in the circuit topology structure, where i and j are positive integers.

[0217] A feature splicing unit, configured to perform feature splicing operations on the time scale, nodes of the transmission line power data, and node i and node j feature in each layer to obtain the edge features of each edge in the circuit topology.

[0218] A first calculation unit, configured to calculate the similarity between nodes i and node j based on the node features and a preset Gaussian kernel function bandwidth parameter, and construct a corresponding similarity matrix through the similarity, so as to perform regional division on the target power system according to the similarity matrix and the spectral clustering algorithm to obtain a regional division Laplacian matrix.

[0219] A feature decomposition unit, configured to perform feature decomposition on the regional division Laplacian matrix to obtain the eigenvectors corresponding to the first q minimum eigenvalues, and perform K-means clustering on the eigenvectors to generate a power system regional division result, where q is a positive integer.

[0220] Optionally, in an embodiment of the present application, the segmented prediction module 200 includes: a first construction unit, a second determination unit, and a second construction unit.

[0221] Among them, the first construction unit is configured to determine the power flowing through each node per hour in each region node of the power system regional division result, and construct a line power constraint based on the preset maximum power flowing through the node per hour, the minimum power flowing through the node per hour, and the power flowing through the node per hour.

[0222] The second determination unit is configured to determine the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour corresponding to each region node, and calculate the total multi-energy power generation per hour according to the wind power generation power per hour, the photovoltaic power generation power per hour, the thermal power generation power per hour, and the power taken out by the energy storage system per hour.

[0223] The second construction unit is configured to calculate the total line load per hour according to the power flowing through the node per hour, and construct a power balance constraint condition based on the total multi-energy power generation per hour and the total line load per hour, so that each region node satisfies the power balance constraint condition.

[0224] Optionally, in an embodiment of the present application, the segmented prediction module 200 further includes: a second calculation unit, a third construction unit, a first acquisition unit, a third calculation unit, and a judgment unit.

[0225] Among them, the second calculation unit is used to calculate the mean and standard deviation of the wind speed data and the irradiance data, and calculate the preprocessed wind speed data and the preprocessed photovoltaic power generation data corresponding to the wind speed data and the irradiance data respectively according to the mean and the standard deviation.

[0226] The third construction unit is used to construct a wind-solar prediction model based on a preset piecewise dynamic trend decomposition prediction model, and perform piecewise prediction on the preprocessed wind speed data and the preprocessed photovoltaic power generation data through the wind-solar prediction model to obtain the k -th segment predicted wind speed and the m -th segment predicted irradiance, where k and m are positive integers.

[0227] The first acquisition unit is used to acquire the cut-in wind speed, the rated wind speed, the cut-out wind speed, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance and the maximum irradiance corresponding to the target power system, and calculate the k -th segment wind power segmented prediction power based on the cut-in wind speed, the rated wind speed, the cut-out wind speed and the k -th segment predicted wind speed, and calculate the m -th segment photovoltaic power generation power through the m -th segment predicted irradiance, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance and the maximum irradiance.

[0228] The third calculation unit is used to calculate the total predicted wind power within k hours and the total predicted photovoltaic power within m hours respectively through the k -th segment wind power segmented prediction power and the m -th segment photovoltaic power generation power, and determine the actual wind power generation power within k hours and the actual photovoltaic power generation power within m hours corresponding to the target power system, so as to calculate the wind power prediction power error according to the actual wind power generation power within k hours and the total predicted wind power within k hours, and calculate the photovoltaic power generation prediction power error through the actual photovoltaic power generation power within m hours and the total predicted photovoltaic power within m hours.

[0229] The judgment unit is used to judge whether the wind power prediction power error and the photovoltaic power generation prediction power error meet the preset error requirements respectively. Among them, when the wind power prediction power error and the photovoltaic power generation prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, the power generation cost coefficient and the k total predicted wind power within mCalculate the photovoltaic power generation cost based on the total predicted photovoltaic power within hours.

[0230] Optionally, in an embodiment of the present application, the modeling module 300 includes: a fourth calculation unit, a second acquisition unit, a third determination unit, a construction unit, a third acquisition unit, a fourth determination unit, a fourth construction unit, and a fifth calculation unit.

[0231] Among them, the fourth calculation unit is used to calculate the hourly thermal power generation value of the target power system, and based on the hourly thermal power generation value, the preset hourly small thermal power generation value, and the maximum hourly thermal power generation value, construct a thermal power output constraint.

[0232] The second acquisition unit is used to acquire the n hourly thermal power generation power of the n- th hour and the n hourly thermal power generation power of the n- (t + 1)th hour, and determine the power change rate constraint of the thermal power unit according to the hourly thermal power generation power of the

[0233] th hour, the

[0234] hourly thermal power generation power of the

[0235] (t + 1)th hour, and the preset ramp rate.

[0236] The third determination unit is used to determine the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, start-up cost coefficient, and start-up state variable, calculate the thermal power generation cost based on the thermal power generation operation and maintenance cost coefficient, fuel cost coefficient, start-up cost coefficient, start-up state variable, and hourly thermal power generation value, and construct a thermal power unit emission constraint according to the preset emission coefficient and maximum allowable emission.

[0237] The construction unit is used to construct a thermal power generation model based on the thermal power generation cost, thermal power unit emission constraint, thermal power unit power change rate constraint, and

[0238] The fifth calculation unit is configured to calculate the energy storage system cost based on a preset charge-discharge cost coefficient of the energy storage system, an operation and maintenance cost coefficient of the energy storage system, and the charge-discharge power of the energy storage system, and construct an energy storage system model through the energy storage system cost, the energy storage system capacity constraint, the charge-discharge power constraint, and the charge-discharge efficiency constraint.

[0239] Optionally, in an embodiment of the present application, the optimization module 400 includes: a collection unit, a training unit, and a fifth determination unit.

[0240] Among them, the collection unit is configured to collect the historical environmental data of the target power system, and determine the first state space, the first reward function, and the first system historical cost corresponding to the first deep deterministic policy gradient agent.

[0241] The training unit is configured to determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, construct the first action space through the parameters to be optimized, and train the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function, and the first system historical cost, and in combination with the historical environmental data of the power system, so as to obtain the target parameters of the agent.

[0242] The fifth determination unit is configured to determine the second state space, the second action space, the second reward function, and the second system cost corresponding to the second deep deterministic policy gradient agent, and optimize the operation of the target power system based on the second state space, the second action space, the second reward function, and the second system cost, and in combination with the target parameters of the agent.

[0243] Optionally, in an embodiment of the present application, the k mathematical expression of the segmented predicted power of the wind power in the

[0244]

[0245] Among them, represents the predicted wind speed of the k segment; represents the segmented predicted power of the wind power in the k segment; represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; represents the temperature; represents the radius of the fan blade; represents the wind direction correction coefficient; represents the included angle between the wind direction and the fan orientation; , , respectively represent the cut-in wind speed, the rated wind speed, and the cut-out wind speed.

[0246] It should be noted that the foregoing explanation of the embodiment of the method for optimizing the operation of the receiving-end power system based on multi-agent also applies to the device for optimizing the operation of the receiving-end power system based on multi-agent in this embodiment, and will not be elaborated here.

[0247] The device for optimizing the operation of the receiving-end power system based on multi-agent proposed according to the embodiment of the present application includes a region division module 100, which is used to determine the circuit topology corresponding to the target power system, extract the node features and edge features corresponding to the circuit topology, and based on the node features, edge features and a preset spectral clustering algorithm, divide the target power system into regions to generate a power system region division result; a segmented prediction module 200, which is used to make each region node in the power system region division result satisfy a preset power balance constraint condition based on a pre-constructed power system power balance model, and obtain the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on a pre-constructed wind and light prediction model and in combination with the wind speed data and irradiance data to obtain corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; a modeling module 300, which is used to construct a double deep deterministic policy gradient optimization model based on a pre-constructed thermal power generation model and energy storage system model, and in combination with the segmented prediction data, wind power generation cost and photovoltaic power generation cost; an optimization module 400, which is used to collect the historical environment data of the power system of the target power system, and train the first deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model through the historical environment data of the power system to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the double deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, thereby providing a reliable technical basis for the optimization of the power balance operation of multi-region nodes of the power system, helping to optimize power dispatching and reduce the grid operation cost.

[0248] Figure 4 The structural schematic diagram of the electronic device provided for the embodiment of the present application. The electronic device may include:

[0249] A memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.

[0250] When the processor 402 executes the program, it implements the method for optimizing the operation of the receiving-end power system based on multi-agent provided in the foregoing embodiment.

[0251] Further, the electronic device further includes:

[0252] A communication interface 403, which is used for communication between the memory 401 and the processor 402.

[0253] A memory 401 for storing a computer program that can run on a processor 402.

[0254] The memory 401 may include a high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.

[0255] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0256] Optionally, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other through an internal interface.

[0257] The processor 402 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0258] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method for optimizing the operation of a receiving-end power system based on multi-agent is implemented.

[0259] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0260] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0261] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.

[0262] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0263] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0264] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0265] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0266] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A receiving-end power system optimization operation method based on multi-agent, characterized in that: The following steps are involved: Determine a circuit topology structure corresponding to a target power system, extract node features and edge features corresponding to the circuit topology structure, and divide the target power system into regions based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system region division result; Based on a pre-constructed power system power balance model, each regional node in the power system regional division result satisfies a preset power balance constraint condition, and obtains wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on a pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain corresponding segmented prediction data, and calculate corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; Based on the pre-built thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost, a dual-depth deterministic policy gradient optimization model is constructed; Collect historical power system environment data of the target power system, and train the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model through the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

2. The receiving-end power system optimization operation method based on multi-agent according to claim 1 is characterized in that: The determining of the circuit topology structure corresponding to the target power system, extracting the node features and edge features corresponding to the circuit topology structure, and performing regional division on the target power system based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system regional division result, includes: Determine a node set, an edge set and a time scale of power data of a transmission line of a circuit topology structure corresponding to the target power system; Get the nodes in each layer of the circuit topology i and nodes j A set of neighboring nodes is obtained, and based on a preset learnable weight matrix, an activation function and the set of neighboring nodes, node features of each node in the circuit topology are extracted, wherein: i and j is a positive integer; Based on the edge set and the preset multi-layer perceptron, the time scale of the power data of the transmission line, the node i and the node j Performing feature concatenation operation on node features in each layer to obtain edge features of each edge in the circuit topology structure; Based on the node characteristics and the preset Gaussian kernel function bandwidth parameters, the node i and the node j The similarity between the two targets is calculated, and a corresponding similarity matrix is ​​constructed through the similarity, so as to perform regional division on the target power system according to the similarity matrix and the spectral clustering algorithm, so as to obtain a regional division Laplace matrix; Performing eigendecomposition on the region partition Laplace matrix to obtain eigenvectors corresponding to the first q minimum eigenvalues, and performing K-means clustering on the eigenvectors to generate the power system region partition result, wherein q is a positive integer.

3. The receiving-end power system optimization operation method based on multi-agent according to claim 1 is characterized in that: The power system power balance model based on the pre-constructed one makes each regional node in the power system regional division result meet the preset power balance constraint condition, including: Determine the hourly node power flowing through each regional node in the power system regional division result, and construct a line power constraint based on the preset hourly maximum node power flowing through the node, the hourly minimum node power flowing through the node and the hourly node power flowing through the node; Determine the wind power generation power within each hour, the photovoltaic power generation power within each hour, the thermal power generation power within each hour, and the energy storage system extraction power within each hour corresponding to each regional node, and calculate the multi-energy total power generation power within each hour based on the wind power generation power within each hour, the photovoltaic power generation power within each hour, the thermal power generation power within each hour, and the energy storage system extraction power within each hour; The total line load per hour is calculated according to the power flowing through the node per hour, and the power balance constraint condition is constructed based on the total power generation power of multiple energy sources per hour and the total line load per hour, so that each regional node satisfies the power balance constraint condition.

4. The receiving-end power system optimization operation method based on multi-agent according to claim 3 is characterized in that: The step of acquiring wind speed data and irradiance data corresponding to the target power system, performing segmented prediction processing based on a pre-built wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain corresponding segmented prediction data, and calculating corresponding wind power generation costs and photovoltaic power generation costs according to the segmented prediction data, includes: Calculating the mean and standard deviation of the wind speed data and the irradiance data, and respectively calculating wind speed preprocessing data and photovoltaic power generation preprocessing data corresponding to the wind speed data and the irradiance data according to the mean and the standard deviation; Based on the preset segmented dynamic trend decomposition prediction model, the wind-solar prediction model is constructed, and the wind speed preprocessing data and the photovoltaic power generation preprocessing data are segmentedly predicted by the wind-solar prediction model to obtain the first k Forecast wind speed and m The predicted irradiance is: k and m is a positive integer; Obtain the cut-in wind speed, rated wind speed, cut-out wind speed, photovoltaic module efficiency, photovoltaic module area, minimum irradiance and maximum irradiance corresponding to the target power system, and calculate the cut-in wind speed, rated wind speed, cut-out wind speed and the first k Forecast wind speed for the segment and calculate the k The wind power is predicted in sections, and the m The first segment is calculated based on the predicted irradiance, the photovoltaic module efficiency, the photovoltaic module area, the minimum irradiance value and the maximum irradiance value. m The photovoltaic power generation power of the segment; Respectively through the k The wind power segment prediction power and the m Calculation of photovoltaic power generation k Total wind power forecast within the hour and m The total photovoltaic power forecast within the hour and determine the corresponding k Actual wind power generation in the hour and m The actual photovoltaic power generation in the hour is based on the k The actual wind power generation capacity and k The wind power forecast total power in the hour is used to calculate the wind power forecast power error, and the wind power forecast power error is calculated by the wind power forecast total power in the hour. m The actual photovoltaic power generation in the hour and the m Calculate the photovoltaic power prediction error by using the total photovoltaic power prediction within the hour; respectively judging whether the wind power prediction power error and the photovoltaic power prediction power error meet the preset error requirements, wherein, when the wind power prediction power error and the photovoltaic power prediction power error meet the error requirements, based on the preset operation and maintenance cost coefficient, the power generation cost coefficient and the k The total wind power forecast within the hour is used to calculate the wind power generation cost, and the preset photovoltaic operation and maintenance cost coefficient, photovoltaic power generation cost coefficient and the m The photovoltaic power generation cost is calculated based on the total photovoltaic power forecast within the hour.

5. The receiving-end power system optimization operation method based on multi-agent according to claim 1 is characterized in that: The dual-depth deterministic policy gradient optimization model is constructed based on the pre-built thermal power generation model and energy storage system model, and in combination with the segmented prediction data, the wind power generation cost and the photovoltaic power generation cost, including: Calculating the hourly thermal power generation value of the target power system, and constructing a thermal power output constraint based on the hourly thermal power generation value, a preset hourly thermal power generation minimum value, and an hourly thermal power generation maximum value; Get the n Hourly thermal power generation and n- 1 hour thermal power generation power, and according to the n Hours of thermal power generation power, the n- The 1-hour thermal power generation power and the preset ramp rate determine the power change rate constraint of the thermal power unit; Determine a thermal power generation operation and maintenance cost coefficient, a fuel cost coefficient, a startup cost coefficient and a startup state variable, and calculate the thermal power generation cost based on the thermal power generation operation and maintenance cost coefficient, the fuel cost coefficient, the startup cost coefficient, the startup state variable and the hourly thermal power generation value, and construct an emission constraint for the thermal power unit according to a preset emission coefficient and a maximum allowable emission amount; Based on the thermal power generation cost, the emission constraint of the thermal power unit, the power change rate constraint of the thermal power unit and the thermal power generation model, constructing the thermal power generation model; Obtaining the charging efficiency, discharging efficiency, charging power and discharging power of the target power system, and calculating the charging and discharging power of the energy storage system based on the charging efficiency, the discharging efficiency, the charging power and the discharging power; Determine corresponding charge and discharge power constraints and charge and discharge efficiency constraints based on a preset maximum charge power of the energy storage system, a maximum discharge power of the energy storage system, the charge efficiency, and the discharge efficiency; Calculating the capacity of the energy storage system by the charge and discharge power of the energy storage system, and constructing the capacity constraint of the energy storage system according to the capacity of the energy storage system, the preset minimum capacity of the energy storage system and the maximum capacity of the energy storage system; Based on the preset energy storage system charging and discharging cost coefficient, the energy storage system operation and maintenance cost coefficient and the energy storage system charging and discharging power, the energy storage system cost is calculated, and the energy storage system model is constructed through the energy storage system cost, the energy storage system capacity constraint, the charging and discharging power constraint and the charging and discharging efficiency constraint.

6. The receiving-end power system optimization operation method based on multi-agent according to claim 5 is characterized in that: The collecting of the power system historical environment data of the target power system, and training the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model by the power system historical environment data to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters, includes: Collecting historical power system environment data of the target power system, and determining a first state space, a first reward function, and a first system historical cost corresponding to the first deep deterministic policy gradient agent; Determine the parameters to be optimized corresponding to the first deep deterministic policy gradient agent, and construct a first action space through the parameters to be optimized, and train the first deep deterministic policy gradient agent based on the first state space, the first action space, the first reward function and the first system historical cost, and in combination with the power system historical environment data, to obtain the agent target parameters; Determine a second state space, a second action space, a second reward function, and a second system cost corresponding to the second deep deterministic policy gradient agent, and optimize the operation of the target power system based on the second state space, the second action space, the second reward function, and the second system cost, in combination with the agent target parameters.

7. The receiving-end power system optimization operation method based on multi-agent according to claim 4 is characterized in that: The said k The mathematical expression of the wind power segment prediction power is: in, Indicates the k Segment forecast wind speed; Indicates the k Segmental wind power prediction power; It represents the wind energy conversion efficiency; represents the atmospheric pressure, represents the gas constant; Indicates the temperature; Indicates the fan blade radius; Indicates the wind direction correction factor; Indicates the angle between wind direction and wind turbine orientation; , , represent the cut-in wind speed, the rated wind speed and the cut-out wind speed respectively.

8. A receiving-end power system optimization operation device based on multi-agent, characterized in that: include: A regional division module, used to determine a circuit topology structure corresponding to a target power system, extract node features and edge features corresponding to the circuit topology structure, and perform regional division on the target power system based on the node features, the edge features and a preset spectral clustering algorithm to generate a power system regional division result; A segmented prediction module is used to make each regional node in the regional division result of the power system meet the preset power balance constraint condition based on the pre-constructed power balance model of the power system, and obtain the wind speed data and irradiance data corresponding to the target power system, so as to perform segmented prediction processing based on the pre-constructed wind-solar prediction model and in combination with the wind speed data and the irradiance data to obtain the corresponding segmented prediction data, and calculate the corresponding wind power generation cost and photovoltaic power generation cost according to the segmented prediction data; A modeling module, for constructing a dual-depth deterministic policy gradient optimization model based on a pre-built thermal power generation model and an energy storage system model, and in combination with the segmented prediction data, the wind power generation cost, and the photovoltaic power generation cost; An optimization module is used to collect the historical environmental data of the target power system, and train the first deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model through the historical environmental data of the power system to determine the agent target parameters corresponding to the first deep deterministic policy gradient agent, so that the second deep deterministic policy gradient agent in the dual deep deterministic policy gradient optimization model optimizes the operation of the target power system according to the agent target parameters.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-agent-based receiving-end power system optimization operation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the receiving-end power system optimization operation method based on multi-agent as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Energy storage scheduling decision-making method and system based on optimal wind-fire storage combined operation index

    CN116454891A

  • Power distribution network regulation and control method and system

    CN119543188A