Active cell site power equipment regulation method based on federated learning and power system
By using federated learning, the server and the agent collaborate in training and clustering, which solves the shortcomings of traditional distribution area control mode in energy management and low-carbon optimization, and realizes efficient and stable control and low-carbon operation of power equipment.
Patent Information
- Application Number
- CN202411666944.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Traditional distribution area control models lack intelligent means to handle complex energy trading, carbon quota management, and power equipment operation, making it difficult to achieve efficient energy management and low-carbon optimization.
By employing a federated learning approach, collaborative training and clustering between the server and the agent are used to integrate and optimize power equipment control strategies, thereby improving learning efficiency and stability.
It has improved the operational stability of the power grid system and the efficiency of power equipment regulation, reduced energy waste and carbon emissions, and achieved efficient low-carbon regulation.
Smart Images

Figure CN119154520B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of electric power engineering, artificial intelligence and federated learning, and more particularly, to a method for regulating active power distribution area equipment based on federated learning and a power system. BACKGROUND
[0002] In the power system, with the rapid development of distributed energy and smart grid, the power equipment in the power distribution area of the power system gradually has the ability of intelligent regulation and self-optimization. However, with the real-time collection of power operation data and the demand for intelligent regulation, the traditional power distribution area regulation mode lacks sufficient intelligent means in dealing with complex energy transactions, carbon quota management, power equipment operation and other tasks, and it is difficult to realize efficient energy management and low-carbon optimization of the power system. SUMMARY
[0003] Therefore, the present application provides a method for regulating active power distribution area equipment based on federated learning and a power system.
[0004] One aspect of the present application provides a method for regulating active power distribution area equipment based on federated learning, applied to a power grid system, the power grid system comprising a server and a plurality of agent ends, the method comprising: performing a first stage training operation on an agent end model by the agent end to obtain first agent end model parameters, wherein the agent end model is used to process the operation state of a target active power distribution area in the power grid system to obtain power equipment actions for regulating power equipment in the target active power distribution area; performing parameter aggregation on the first agent end model parameters of each of the plurality of agent ends by the server to obtain global model parameters; performing a second stage training operation on a global agent model of each of the plurality of agent ends by the agent end to obtain second agent end model parameters of the agent end, the global agent model being determined based on the global model parameters; clustering the second agent end model parameters of each of the plurality of agent ends by the server to obtain a plurality of center model parameters; determining target agent end model parameters from the plurality of center model parameters by the agent end; and regulating the power equipment related to the agent end to perform power equipment actions by the agent end using a target agent end model determined based on the target agent end model parameters.
[0005] Another aspect of the present application provides a power grid system, comprising: a server and a plurality of agent terminals; the agent terminal is configured to perform a first stage training operation on an agent terminal model to obtain a first agent terminal model parameter, wherein the agent terminal model is used to obtain a power equipment action for regulating power equipment in a target active area in the power grid system by processing the area operation state of the target active area; the server is configured to aggregate the first agent terminal model parameters of each of the plurality of agent terminals to obtain global model parameters; the plurality of agent terminals are configured to perform a second stage training operation on a global agent model of each of the plurality of agent terminals to obtain a second agent terminal model parameter of the agent terminal, and the global agent model is determined based on the global model parameters; the server is configured to cluster the second agent terminal model parameters of each of the plurality of agent terminals to obtain a plurality of center model parameters; the agent terminal is configured to determine a target agent terminal model parameter from the plurality of center model parameters; and the target agent terminal model determined based on the target agent terminal model parameter is used to regulate the power equipment related to the agent terminal to perform the power equipment action.
[0006] According to the embodiment of the present application, by training the local agent terminal model of each agent terminal based on the federated learning mechanism, the first agent terminal model parameter can preliminarily learn the power equipment regulation strategy of the agent terminal locally. By aggregating the first agent terminal model parameters of each of the plurality of agent terminals, the global model parameters obtained can integrate the power equipment regulation strategies of the plurality of agent terminals, and each agent terminal can accelerate the learning efficiency and speed of the power equipment regulation strategy by training the global agent model. The server clusters the second agent terminal model parameters to obtain a plurality of center model parameters representing the centers of different types of power equipment regulation strategies, and then the agent terminal determines the target agent terminal model parameter from the plurality of center model parameters, so that the target agent terminal model integrates the power equipment regulation strategies of the plurality of agent terminals of the same type, and the agent terminal can regulate the power equipment to perform the power equipment action based on the target agent terminal model more efficiently and stably, thereby improving the overall operation stability of the power grid system. BRIEF DESCRIPTION OF DRAWINGS
[0007] The above and other objects, features and advantages of the present application will become more apparent from the following description of the embodiments of the present application taken with reference to the accompanying drawings, in which:
[0008] Figure 1 An exemplary system architecture of a power system to which the active area power equipment regulation method based on federated learning can be applied according to an embodiment of the present application is shown;
[0009] Figure 2A flow chart of the active cell site power equipment regulation method based on federated learning according to an embodiment of the present application is shown.
[0010] Figure 3 A flow chart of the target agent end model parameter determined by the agent end from the plurality of center model parameters according to an embodiment of the present application is shown.
[0011] Figure 4 A flow chart of the active cell site power equipment regulation method based on federated learning according to another embodiment of the present application is shown.
[0012] Figure 5 An effect comparison chart of the active cell site power equipment regulation method based on federated learning according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0013] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It should be understood, however, that the description which follows is merely exemplary and is not intended to limit the scope of the application. In the following detailed description of embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that one or more embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.
[0014] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "includes" and tautological expressions thereof, means the inclusion of but not limited to.
[0015] All terms used herein including technical and scientific terms have the same meanings as commonly understood by one of ordinary skill in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning that is consistent with the context of the specification, and not be interpreted in an idealized or overly formal way.
[0016] In the case where expressions similar to "at least one of A, B, and C, and the like" are used, it should be generally construed that the meaning is interpreted as the meaning commonly understood by one of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0017] In the embodiments of the present application, the collection, updating, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information) comply with the relevant legal regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data, and to maintain user personal information security, network security and national security.
[0018] In the embodiments of the present application, the authorization or consent of the user is obtained before the user's personal information is acquired or collected.
[0019] The inventor found that with the rapid development of artificial intelligence and blockchain technology, device control methods based on federated learning are gradually attracting attention. The active power supply area power device regulation and control method based on federated learning provided in the embodiments of the present application can realize cross-power supply area data collaborative calculation and power device regulation and control optimization.
[0020] The embodiments of the present application provide a kind of active power supply area power device regulation and control method based on federated learning and power system. The active power supply area power device regulation and control method based on federated learning is applied to power grid system, and power grid system includes server and multiple intelligent agent ends, and active power supply area power device regulation and control method includes: intelligent agent end carries out first stage training operation to intelligent agent end model, obtains first intelligent agent end model parameter, wherein, intelligent agent end model is used to obtain the power device action for regulating and controlling the power device in target active power supply area by processing the power supply area operating state of target active power supply area in power grid system;Server carries out parameter aggregation to the first intelligent agent end model parameter of multiple intelligent agent ends respectively, obtains global model parameter;Multiple intelligent agent ends carry out second stage training operation to respective global intelligent agent model, obtain the second intelligent agent end model parameter of intelligent agent end, and global intelligent agent model is determined based on global model parameter;Server carries out clustering to the second intelligent agent end model parameter of multiple intelligent agent ends respectively, obtains multiple center model parameters;Intelligent agent end determines target intelligent agent end model parameter from multiple center model parameters;And intelligent agent end uses target intelligent agent end model determined based on target intelligent agent end model parameter, regulates and controls the power device related to intelligent agent end to execute power device action.
[0021] According to an embodiment of the present application, the first agent end model parameters can be preliminarily learned to the power equipment regulation strategy at the agent end locally by training the local agent end model of each agent end based on a federated learning mechanism. The global model parameters obtained by aggregating the first agent end model parameters of each agent end can fuse the power equipment regulation strategies of multiple agent ends, so that each agent end can accelerate the learning efficiency and speed of the power equipment regulation strategy by training the global agent model. The server can make the multiple center model parameters represent the center of different types of power equipment regulation strategies by clustering the second agent end model parameters, so that the target agent end model can be determined from the multiple center model parameters by the agent end, and the target agent end model can fuse the power equipment regulation strategies of multiple agent end models of the same type, so that the agent end can regulate the power equipment to perform the power equipment action based on the target agent end model more efficiently and stably, and improve the overall operation stability of the power grid system.
[0022] Figure 1 An exemplary system architecture of a power system to which the active power supply area power equipment regulation method based on federated learning can be applied according to an embodiment of the present application is shown. It should be noted that, Figure 1 The system architecture shown is only an example of a system architecture to which the embodiments of the present application can be applied, to help those skilled in the art understand the technical content of the present application, but does not mean that the embodiments of the present application cannot be used in other devices, systems, environments or scenarios.
[0023] As Figure 1 The system architecture of the power system 100 according to this embodiment can include a first agent end 101, a second agent end 102, a third agent end 103 and a server 104, as shown.
[0024] The user can use the first agent end 101, the second agent end 102, the third agent end 103 to interact with the server 104 through the network 104 to receive or send messages, etc. The first agent end 101, the second agent end 102, the third agent end 103 can be servers, computers or other devices for performing data processing tasks. The server 104 can be any type of server such as a cloud server, a cluster server, etc. The server 104 can serve as a server to interact with the first agent end 101, the second agent end 102, the third agent end 103 for data interaction. The first agent end 101, the second agent end 102, the third agent end 103 can respectively regulate the power equipment of the first target active power supply area, the second target active power supply area and the third target active power supply area. The power equipment can include power generation equipment such as gas generators and wind turbines, or can also include transformers, secondary equipment and other types of power equipment.
[0025] It should be understood that Figure 1 The number of servers, agent ends, and target active substations in the system is only illustrative. Any number of servers, agent ends, and target active substations can be provided as needed.
[0026] Figure 2 A flowchart of a method for active substation power equipment regulation based on federated learning according to an embodiment of the application is shown.
[0027] As Figure 2 shown, the method includes operations S210-S260.
[0028] At operation S210, the agent end performs a first-stage training operation on the agent end model to obtain first agent end model parameters, wherein the agent end model is used to process substation operating states of a target active substation in a power grid system to obtain power equipment actions for regulating power equipment in the target active substation.
[0029] According to an embodiment of the application, the application is applied to a power grid system, which includes a server and a plurality of agent ends. The server can include a server, a distributed server, or any other type of device capable of performing data cloud functions. The agent end can include a computer, a local server, or any other type of device.
[0030] In one example, the agent end can be an edge device of the power grid system, for example, the agent end can be an intelligent fusion terminal in an active substation. The edge device can be communicatively connected to terminal devices such as smart meter devices in the power grid system to form edge devices of the power grid system.
[0031] In one example, the agent end model can be constructed based on a reinforcement learning model, and the agent end can perform the first-stage training operation based on a deep deterministic policy gradient (DDPG) algorithm.
[0032] According to an embodiment of the application, the target active substation can be associated with the agent end, and the same agent end can be associated with one or more target active substations to adjust the power equipment of the one or more target active substations to perform power equipment actions.
[0033] According to an embodiment of the application, the substation operating state can include operating states of power equipment in the target active substation, user-side power load, power line operating state data, and the like. The power equipment action can include regulation information for controlling operating parameters of the power equipment, such as power regulation information for regulating power generation parameters of a power generator, and the like.
[0034] At operation S220, the server aggregates the first agent-end model parameters of the plurality of agent ends to obtain global model parameters.
[0035] At operation S230, the plurality of agent ends perform second-stage training operations on the respective global agent models to obtain second agent-end model parameters of the agent ends.
[0036] According to an embodiment of the present application, the global agent model is determined based on the global model parameters. The server can send the global model parameters to the plurality of agent ends, and each agent end can migrate the global model parameters to the current first agent-end model to obtain the respective global agent model of the plurality of agent ends.
[0037] According to an embodiment of the present application, the first agent-end model parameters of the plurality of agent ends can be processed based on any type of algorithm, so that the global model parameters obtained can learn the respective control strategies of the plurality of first agent-end model parameters about the control of the power equipment by the agent end locally. In this way, the plurality of agent ends can perform the second-stage training operations on the respective global agent models to speed up the model training speed of the agent end locally and save the computing cost and energy consumption required for training the model.
[0038] In one example, the global agent model can be constructed based on a reinforcement learning model, and the agent end can perform the second-stage training operation based on a deep deterministic policy gradient (DDPG) algorithm.
[0039] At operation S240, the server clusters the second agent-end model parameters of the plurality of agent ends to obtain a plurality of center model parameters.
[0040] According to an embodiment of the present application, the second agent-end model parameters of the plurality of agent ends can be processed based on a clustering algorithm to obtain a plurality of clustering clusters. The same clustering cluster can include second agent-end model parameters representing the same type, and the clustering center of the clustering cluster can be a center model parameter.
[0041] At operation S250, the agent end determines target agent-end model parameters from the plurality of center model parameters.
[0042] According to an embodiment of the present application, the target agent-end model parameter can be a model parameter satisfying the power equipment regulation and control requirement of the agent end. The agent end can determine a candidate agent model parameter from the plurality of center model parameters, perform model parameter migration on the current second agent model based on the candidate agent model parameter, and obtain a candidate agent model. If the candidate power equipment action output by the candidate agent model can satisfy the regulation and control requirement, the candidate agent model can be determined as the target agent model.
[0043] It should be noted that the embodiments of the present application do not limit the specific way of determining the target agent-end model parameter from the plurality of center model parameters, as long as the agent end can satisfy the power equipment regulation and control requirement of the target active area.
[0044] In operation S260, the agent end regulates the power equipment related to the agent end to perform the power equipment action by using the target agent-end model determined based on the target agent-end model parameter.
[0045] According to an embodiment of the present application, the agent end can use the target agent-end model to process the area operation state of the target active area to obtain the power equipment action. The power equipment can operate based on the equipment regulation and control information represented by the power equipment action, so as to improve the operation efficiency and stability of the power equipment in each target active area in the power system.
[0046] According to an embodiment of the present application, the global model parameter and the center model parameter are ciphertext data encrypted by the server based on the homomorphic encryption algorithm, and the first agent-end model parameter and the second agent-end model parameter are ciphertext data encrypted by the agent end based on the homomorphic encryption algorithm.
[0047] According to an embodiment of the present application, the server and the agent end can perform data encryption based on the homomorphic encryption algorithm agreed upon, so that the model parameters (such as the global model parameter and the center model parameter, and the first agent-end model parameter and the second agent-end model parameter) transmitted between the server and the agent end can be ciphertext parameters, thereby improving the information interaction security.
[0048] In one example, the server aggregates the first agent-end model parameters of the plurality of agent ends to obtain the global model parameter, which can include: the plurality of agent ends homomorphically encrypt the first agent-end model parameters respectively, and send the first agent-end model parameters with ciphertext data attributes obtained to the server. The server performs weighted tie-breaking calculation on the first agent-end model parameters with ciphertext data attributes to obtain the global model parameter after parameter aggregation.
[0049] (1);
[0050] In formula (1), denotes the first agent model parameter with the ciphertext data attribute of the i-th agent end, and Enc[] denotes a homomorphic encryption algorithm. denotes the global model parameter with the ciphertext data attribute. Ω denotes the number of agent ends of a specified type.
[0051] In one example, the service end can broadcast the global model parameter with the ciphertext data attribute to the plurality of agent ends, the agent ends can perform homomorphic decryption on the global model parameter with the ciphertext data attribute based on a homomorphic decryption algorithm to obtain plaintext information of the global model parameter. Model parameter migration is performed based on the plaintext information of the global model parameter to obtain the global agent model.
[0052] According to an embodiment of the present application, the agent end model includes an action network and an evaluation network. The agent end model can be a model constructed based on a reinforcement learning algorithm, and the action network (or actor network) and the evaluation network (or critic network) can be constructed based on a deep learning algorithm. The action network and the evaluation network can be isomorphic models with the same or similar network structure.
[0053] According to an embodiment of the present application, the first-stage training operation can include the following operations.
[0054] The action network is used to process the t-th sample substation operating state to obtain a t-th sample power equipment action. The evaluation network is used to process the t-th sample substation operating state and the t-th sample power equipment action to obtain a t-th target evaluation value. The action network is used to process a t+1-th sample substation operating state to obtain a t+1-th sample power equipment action. The evaluation network is used to process the t+1-th sample substation operating state and the t+1-th sample power equipment action to obtain a t+1-th target evaluation value. The t-th target evaluation value and the t+1-th target evaluation value are processed based on a loss function minimization to obtain a policy loss value. The model parameters of the current agent end model are updated based on the policy loss value until a training condition is met.
[0055] In one example, the i-th agent end of the target active substation obtains a t-th sample substation operating state, and processes the t-th sample substation operating state s i,t The action network of the i-th agent end is input, and the t-th sample power equipment action a i,t The t-th target evaluation value can be a Q value obtained based on a Q function, and the t-th target evaluation value is denoted as The t+1-th sample substation operating state s i,t+1 The action network of the i-th agent end is input, and the t+1-th sample power equipment action a i,t+1 The t+1-th initial evaluation value is denoted as , the t+1 target evaluation value is expressed as . Wherein, r represents the reward function value, which can be obtained from the reward function value corresponding to the t sample station area running state in the four tuple of the sample experience pool. Denotes a preset coefficient. The policy loss value is obtained by processing the t target evaluation value and the t+1 target evaluation value using the minimum loss function. The model parameters of the action network are updated by processing the measurement loss value based on the stochastic gradient descent algorithm. Wherein the model parameters of the action network can be updated based on the value function and the policy function of the stochastic gradient descent algorithm guided by the temporal difference (TD) signal. γ represents a preset coefficient of the reward function.
[0056] According to an embodiment of the present application, the server aggregates the first agent end model parameters of each of the plurality of agent ends, including: determining the aggregation weights of the plurality of first agent end model parameters based on the sample data amounts of each of the plurality of agent ends; and performing weighted calculation on the plurality of first agent end model parameters based on the aggregation weights of the plurality of first agent end model parameters to obtain the global model parameters.
[0057] According to an embodiment of the present application, the sample data amount can be the data amount of the sample station area running state used to train the first agent end model parameters, and the larger the data amount scale of the sample data amount, the better the learning effect of the first agent model parameters. Based on the data amount scale represented by the sample data amount of each of the plurality of agent ends, the aggregation weights of the first agent end model parameters of the plurality of agent ends are allocated, so that the first agent end model parameters corresponding to the sample data amount with a larger data amount scale are allocated a larger aggregation weight. Therefore, the aggregation weights can be used to perform weighted calculation on the plurality of first agent end model parameters, so that the global model parameters can learn the regulation and control ability of the power equipment with strong regulation and control ability, so as to improve the accuracy and adaptability of the global model parameters for power equipment action prediction, and improve the training efficiency and training effect of the subsequent agent terminal device in the second stage training operation.
[0058] According to an embodiment of the present application, the agent end determines the target agent end model parameter from the plurality of center model parameters, including each agent end performing the operations shown in Figure 3 .
[0059] Figure 3 A flowchart showing the agent end determining the target agent end model parameter from the plurality of center model parameters according to an embodiment of the present application is shown.
[0060] As shown in Figure 3 , determining the target agent end model parameter from the plurality of center model parameters can include operations S310-S340.
[0061] At operation S310, the global agent model is updated based on the target central model parameter in the plurality of central model parameters to obtain an initial central agent end model.
[0062] At operation S320, the sample area operation state is processed based on the initial central agent end model to obtain an experience power equipment action.
[0063] At operation S330, the experience power equipment action is processed based on a loss function to obtain an experience loss value.
[0064] At operation S340, in response to the experience loss value satisfying a preset loss threshold condition, the target central model parameter is determined as a target agent end model parameter.
[0065] According to an embodiment of the present application, updating the global agent model based on the target central model parameter in the plurality of central model parameters can include performing model parameter transfer learning on the global agent model based on the target central model parameter to obtain the initial central agent end model.
[0066] According to an embodiment of the present application, processing the sample area operation state based on the initial central agent end model can include processing the sample area operation state of the second stage training operation by using the initial central agent end model to output the experience power equipment action. The experience loss value can be obtained by processing the experience power equipment action and label data corresponding to the experience power equipment action by using the loss function.
[0067] According to an embodiment of the present application, the experience loss value satisfying the preset loss threshold condition can include the experience loss value being less than or equal to a preset loss value threshold. Alternatively, it can also include a gradient value determined based on the experience loss value converging in a preset time threshold. Alternatively, it can also include the gradient value determined based on the experience loss value being less than or equal to a preset gradient value threshold. The embodiments of the present application do not limit the design manner of the preset loss threshold condition.
[0068] According to an embodiment of the present application, after the service end clusters the second intelligent end model parameters, it verifies whether the target center model parameters meet the power equipment regulation and control requirements of the intelligent end in the target active area based on the verification loss function through the intelligent end in the target active area, realizes local secondary clustering analysis of the multiple center model parameters obtained after preliminary clustering by the intelligent ends, and avoids heavy calculation overhead and communication burden caused by clustering model parameters of each two of the multiple intelligent ends in the federated learning process. Through the double-end clustering analysis of the second intelligent end model parameters by the service end and the intelligent end, the intelligent end dynamically selects the multiple center model parameters, timely corrects the clustering results of the center model parameters of the service end, accelerates the learning efficiency and training efficiency of the model for the multiple intelligent ends with the same type of power equipment regulation and control mode, and improves the prediction accuracy and robustness of the target intelligent end model.
[0069] In one example, the empirical loss value corresponding to the center model parameters of the clustering cluster c j may be represented based on the following formula (2).
[0070] (2).
[0071] In formula (2), argminL() represents the empirical loss function, represents the center model parameters of the clustering cluster c j , is the tth sample data of the ith intelligent end, and the tth sample data can include the tth sample area operating state and the corresponding sample power equipment action, etc.
[0072] In one example, the multiple intelligent ends can each perform the cooperative training operation based on the following operations.
[0073] The intelligent end of the target active area performs model initialization, sets the total communication round of the initialization training period as Nc, the starting round of the service end clustering operation as Ne=Nc+1, and sets the learning rate, the number of intelligent ends i.
[0074] In the first-stage training operation, each agent end locally performs the first-stage and second-stage training operations, for example, Nc rounds of reinforcement learning training can be specified to obtain the second agent end model parameters. Before the Ne training starts, the second agent end model parameters encrypted based on the homomorphic encryption algorithm are sent to the server. The server clusters the second agent end model parameters of each agent end based on a clustering algorithm to obtain multiple clusters. The cluster center of the cluster is the center model parameter. The agent end verifies the target center model parameter based on the experience loss function and sample data. If the experience loss value meets the preset loss threshold condition, the target center model parameter is determined as the target agent end model parameter. At the same time, the multiple agent ends corresponding to the target center model parameter can be divided into the target cluster c j of the target center model parameter.
[0075] According to an embodiment of the present application, the agent end uses the target agent end model determined based on the target agent end model parameter to regulate the power equipment related to the agent end to perform the power equipment action, which includes the following operations performed by the agent end.
[0076] The target agent end model is used to process the tth district operation state of the tth time period to obtain the tth power equipment action; the power equipment related to the agent end is controlled to perform the tth power equipment action to obtain the t+1th district operation state of the t+1th time period; the tth regulation reward information is obtained by processing the tth district operation state and the tth power equipment action based on the reward function; the target agent end model parameter of the current target agent end model is optimized based on the tth regulation reward information to obtain a new target agent end model; the t+1th power equipment action is obtained by processing the t+1th district operation state based on the new target agent end model; and the power equipment related to the agent end is controlled to perform the t+1th power equipment action.
[0077] According to an embodiment of the present application, the tth district operation state includes at least one of the following state information generated by the target active district in the tth time period: thermal energy storage device output, electric energy storage device power storage information, user side load, electric energy transaction information, and new energy device output.
[0078] According to an embodiment of the present application, the power storage information can include the power stored by the electric energy storage device. The electric energy transaction information can include the electric energy price, but is not limited thereto, and can also include the carbon transaction price, the gas purchase price, and the like. The new energy device output can include the photovoltaic energy device output, the wind power generation device output, and the like.
[0079] According to an embodiment of the present application, the tth power equipment action includes at least one of fuel cell operation power regulation information, gas turbine operation power regulation information, and gas boiler operation power regulation information.
[0080] In one example, the power regulation process of the ith target active area in the tth period can be understood as a time sequence Markov decision process. The tth area operation state s i,t Input the target intelligent agent end model to obtain the tth power equipment action a i,t The intelligent agent end regulates the power equipment of the ith target active area to execute the tth power equipment action by executing the tth power equipment action a i,t , and obtains the t+1th area operation state s i,t+1 .
[0081] The tth area operation state s i,t may be represented based on the following formula (3).
[0082] (3);
[0083] Wherein, p i,t represents a purchased electricity price signal, . Wherein, (p ) is a direct purchase power price signal, (p ) is a carbon quota transaction price signal, is a gas purchase price signal.
[0084] The carbon quota transaction price signal and the gas purchase price signal. represents a user-side load, represents a user-side power consumption, represents a photovoltaic device output, represents a wind power generation device output, represents an electric energy storage (EES) device power storage information, represents a thermal energy storage (TES) device output. Ω R represents a residential target active area, Ω C represents a commercial target active area, and Ω I represents an industrial target active area.
[0085] The tth power equipment action a i,t may be represented based on the following formula (4).
[0086] (4)
[0087] wherein, represent the proportion of the operating power output of the fuel cell (FC) device, the microturbine (MT) device and the gas boiler (GB) in the total operating power output, and a value of 0 indicates no output, represents the electric energy storage (EES) / thermal energy storage (TES) energy storage stage, and the output value is negative, which represents the energy release stage, and the output value is positive.
[0088] In the present example, the training and updating of the target agent end model at the agent end can be implemented based on a double-delay deep deterministic policy gradient (DDPG) to realize the generation of a new district operating state by interacting with the real-time environment of the target active district and the optimization of the power equipment action policy to obtain an optimal economic low-carbon regulation and control strategy.
[0089] The optimization goal of each agent end is to reduce the economic cost and carbon dioxide emissions brought by regulation and control operation and carbon quota trading as much as possible within a certain time scale, and therefore the determination of the regulation and control reward information can include the electricity-carbon trading price, the gas cost, the power equipment operation cost and the heat supply and demand balance constraint penalty term in the tthtarget active district. Unlike the trading constraints of net electricity demand, the active district agent will be punished by formula (7) if it cannot meet the heat supply and demand balance under the condition of complying with the component constraints, so as to realize the optimization of the target agent end model parameter of the current target agent end model.
[0090] In one example, the tthregulation and control reward information can be represented based on formula (5).
[0091] (5);
[0092] wherein, r i,t is the tthregulation and control reward information, is the power equipment operation cost, is the gas trading cost; is the net electricity demand. The gas trading cost and the heat supply and demand balance reward information are represented based on formulas (6) and (7) as follows
[0093] (6);
[0094] (7)
[0095] in, Information on rewards for balancing heat energy supply and demand. For the output of the combined heat and power (CHP) generator, To provide power to the gas-fired boiler, ζ Q and ζ S V represents the penalty coefficient for operational constraints. i,t This refers to the amount of gas purchased from external sources. and To directly purchase power quantity signals. and This is a signal for carbon trading volume. As a signal for gas purchase prices, To directly purchase electricity price signals, This is a signal for carbon quota trading prices.
[0096] In one example, the improved mean of absolute differences of pairwise cosine similarity (IMADC) algorithm can be used to perform cluster analysis on the parameters of multiple second-agent end models to address the distance clustering defects caused by high-dimensional computation.
[0097] For example, it can be based on formula (8)
[0098] (8);
[0099] Where k represents the number of model gradient update vectors, which can be used to update the gradient vectors corresponding to the model parameters at the second agent end. Perform truncated singular value decomposition along the neural network dimension to obtain k principal component update vectors v. α(n,v) represents the cosine similarity between the nth second agent end model parameter and the principal component update vector v, and α(m,v) represents the cosine similarity between the mth second agent end model parameter and the principal component update vector v.
[0100] α(n, v) can be determined based on the following formula (9).
[0101] (9);
[0102] The cosine similarity between the second agent model parameters and the principal component update vector at multiple agent ends can be determined by formula (5), and then the similarity matrix is obtained, which is expressed by formula (10).
[0103] (10)
[0104] The matrix elements of the similarity matrix S can be used to obtain clusters of second-agent model parameters from multiple agents based on the IMADC metric and a clustering algorithm. , where c i Let be the i-th cluster. The cluster center of the cluster can be represented as the center model parameter.
[0105] According to embodiments of the present invention, multiple associated agent ends are associated with the same target center model parameters, and the multiple agent ends include associated agent ends. It can be understood that the multiple associated agent ends may correspond to the same cluster, and the multiple associated agent ends can be understood as being the same agent end cluster, and the multiple associated agent ends of the same agent end cluster have similar control strategies.
[0106] Figure 4 A flowchart of a method for controlling active transformer substation power equipment based on federated learning according to another embodiment of the present invention is shown.
[0107] like Figure 4 As shown, the active distribution area power equipment control method based on federated learning also includes performing operations S410~S430 during the preset training phase.
[0108] During operation S410, the associated intelligent agent transmits the associated target intelligent agent model parameters to the server.
[0109] During S420 operation, the server aggregates the parameters of multiple currently associated target agent end-models to obtain the associated global model parameters; and
[0110] When operating S430, the associated intelligent agent uses a new target intelligent agent model determined based on the parameters of the associated global model to regulate the power equipment in the target active transformer area to perform power equipment actions.
[0111] According to an embodiment of the present invention, before the preset training phase, the associated agent end can perform training operations on their respective target agent end models based on a reinforcement learning mechanism to obtain associated target agent end model parameters related to the preset training phase.
[0112] According to an embodiment of the present invention, the associated target intelligent agent end model parameters can be ciphertext data encrypted based on a homomorphic encryption algorithm.
[0113] According to an embodiment of the present application, the service end aggregates the parameters of the plurality of current associated target intelligent agent end model parameters, which can realize the aggregation of the power equipment regulation strategies learned by the plurality of associated intelligent agent end models with similar regulation strategy modes, so as to facilitate the collaborative training of the plurality of associated intelligent agent ends in the same intelligent agent end cluster, and enable the associated global model parameters to accurately represent the diversified power equipment regulation strategies. Thus, the new target intelligent agent model is given to regulate the power equipment of the target active substation to perform the power equipment action, which can realize the iterative update of the power equipment regulation strategy of the target active substation, and improve the overall operation efficiency and operation stability of the distribution network.
[0114] Figure 5 An effect comparison diagram of the active substation power equipment regulation method based on federated learning provided by an embodiment of the present application is shown.
[0115] As Figure 5 shown, the active substation power equipment regulation method based on federated learning provided by an embodiment of the present application performs power equipment regulation by using three regulation methods, i.e., the advantage actor-critic (A2C) reinforcement learning algorithm, the master-slave game algorithm, and the deep Q-network (DQN) algorithm. The target active substation operation income improvement result and the power equipment execution efficiency improvement result both show obvious advantages, which indicates that the active substation power equipment regulation method based on federated learning provided by an embodiment of the present application can effectively optimize the operation cost of the power equipment in the active substation, and reduce energy waste and carbon emissions. Through federated learning and collaborative training of the intelligent agent ends of the plurality of target active substations, the device privacy protection and external attack problems can also be better addressed, and efficient economic and low-carbon regulation strategies can be implemented. In addition, through dynamic clustering and the optimized federated reinforcement learning mechanism, collaborative optimization can be realized among the intelligent agents of the plurality of active substations, and the solution speed is accelerated.
[0116] Based on the active substation power equipment regulation method based on federated learning provided by the above embodiment, an embodiment of the present application further provides a power grid system,
[0117] The power grid system provided by an embodiment of the present application comprises a service end and a plurality of intelligent agent ends.
[0118] According to an embodiment of the present application, the intelligent agent end is configured to perform a first-stage training operation on the intelligent agent end model to obtain first intelligent agent end model parameters, wherein the intelligent agent end model is used to obtain a power equipment action for regulating the power equipment in the target active substation by processing the substation operation state of the target active substation in the power grid system.
[0119] According to an embodiment of the present application, the server is configured to perform parameter aggregation on the first agent-end model parameters of the plurality of agent ends respectively, to obtain global model parameters.
[0120] According to an embodiment of the present application, the plurality of agent ends are configured to perform a second-stage training operation on the global agent model respectively to obtain second agent-end model parameters of the agent ends, the global agent model being determined based on the global model parameters.
[0121] According to an embodiment of the present application, the server is configured to perform clustering on the second agent-end model parameters of the plurality of agent ends respectively, to obtain a plurality of center model parameters.
[0122] According to an embodiment of the present application, the agent end is configured to determine target agent-end model parameters from the plurality of center model parameters; and to control the power equipment associated with the agent end to perform the power equipment action by using a target agent-end model determined based on the target agent-end model parameters.
[0123] The power grid system provided by the embodiments of the present application can be used to perform the active power equipment regulation method based on federated learning provided by the embodiments of the present application. The active power equipment regulation method based on federated learning provided by the embodiments of the present application can be applied to the power grid system provided by the embodiments of the present application, which will not be described herein.
[0124] The flowcharts and block diagrams in the drawings illustrate the possible architectural, functional, and operational architectures of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the figures. For example, two blocks that are shown in succession can actually be executed substantially concurrently, or they can be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or they can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways, even if such combinations or integrations are not expressly noted in the present application. In particular, features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.
[0125] The embodiments of the application have been described. However, these embodiments are merely for illustration and are not intended to limit the scope of the application. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Various alternatives and modifications to the embodiments described herein will be apparent to those skilled in the art in view of the foregoing without departing from the scope of the application.
Claims
1. A method for controlling active distribution transformer power equipment based on federated learning, characterized in that, The method is applied to a power grid system including a server and a plurality of agent terminals, and comprises the following steps: The agent terminal performs a first-stage training operation on an agent terminal model to obtain first agent terminal model parameters, wherein the agent terminal model is used to process the operating state of a target active power supply area in the power grid system to obtain power device actions for regulating power devices in the target active power supply area; The server aggregates the first agent terminal model parameters of the plurality of agent terminals to obtain global model parameters; The plurality of agent terminals perform a second-stage training operation on a global agent model based on the global model parameters to obtain second agent terminal model parameters of the agent terminals; The server clusters the second agent terminal model parameters of the plurality of agent terminals based on a clustering algorithm to obtain a plurality of center model parameters; The agent terminal performs the following operations: updating the global agent model based on a target center model parameter in the plurality of center model parameters to obtain an initial center agent terminal model; processing sample power supply area operating states based on the initial center agent terminal model to obtain experience power device actions; processing the experience power device actions based on a loss function to obtain an experience loss value; determining the target center model parameter as a target agent terminal model parameter in response to the experience loss value satisfying a preset loss threshold condition; The agent terminal regulates power devices associated with the agent terminal to perform power device actions using a target agent terminal model determined based on the target agent terminal model parameter; wherein a plurality of associated agent terminals are associated with the same target center model parameter, the plurality of associated agent terminals are clustered as a same agent terminal cluster, and the plurality of agent terminals include the associated agent terminals, The method further comprises: In a preset training stage, the associated agent terminals transmit associated target agent terminal model parameters of associated target agent terminal models to the server; The server aggregates a plurality of current associated target agent terminal model parameters to obtain associated global model parameters; and The associated agent terminals regulate power devices of the target active power supply area to perform power device actions using a new target agent model determined based on the associated global model parameters; wherein, before the preset training stage, the associated agent terminals perform training operations on respective target agent terminal models based on a reinforcement learning mechanism to obtain associated target agent terminal model parameters related to the preset training stage.
2. The active feeder power device regulation method of claim 1, wherein, The agent terminal regulates power devices associated with the agent terminal to perform power device actions using a target agent terminal model determined based on the target agent terminal model parameter, which comprises the following operations: The agent terminal performs the following operations: processing the tth power supply area operating state in the tth time period using the target agent terminal model to obtain the tth power device action; control the power equipment related to the agent end to execute the tth power equipment action, to obtain the t+1th substation operation state in the t+1th period; process the tth substation operation state and the tth power equipment action based on the reward function to obtain tth control reward information; optimize the target agent end model parameter of the current target agent end model based on the tth control reward information to obtain a new target agent end model; process the t+1th substation operation state based on the new target agent end model to obtain the t+1th power equipment action; and control the power equipment related to the agent end to execute the t+1th power equipment action.
3. The active feeder power device regulation method of claim 2, wherein, The tth substation operation state includes at least one of the following state information generated by the target active substation in the tth period: thermal energy storage device output, electric energy storage device power storage information, user side load, electric energy transaction information, and new energy device output.
4. The active feeder power device regulation method of claim 2, wherein, The tth power equipment action includes at least one of the following: fuel cell operation power regulation information, gas turbine operation power regulation information, and gas boiler operation power regulation information.
5. The active utility regulation method for a utility metering area according to claim 1, wherein, The agent end model includes an action network and an evaluation network. The first stage training operation includes: processing the tth sample substation operation state using the action network to obtain the tth sample power equipment action; processing the tth sample substation operation state and the tth sample power equipment action using the evaluation network to obtain the tth target evaluation value; processing the t+1th sample substation operation state using the action network to obtain the t+1th sample power equipment action; processing the t+1th sample substation operation state and the t+1th sample power equipment action using the evaluation network to obtain the t+1th target evaluation value; processing the tth target evaluation value and the t+1th target evaluation value based on a minimum loss function to obtain a policy loss value; updating the model parameter of the current agent end model based on the policy loss value until the training condition is met.
6. The active utility regulation method for a utility metering area according to claim 1, wherein, The global model parameter and the center model parameter are ciphertext data encrypted by the service end based on a homomorphic encryption algorithm, and the first agent end model parameter and the second agent end model parameter are ciphertext data encrypted by the agent end based on the homomorphic encryption algorithm.
7. The active utility regulation method for a utility metering area according to claim 1, wherein, The service end parameter aggregation of the first agent end model parameter of each of the plurality of agent ends includes: determining the aggregation weight of the plurality of first agent end model parameters based on the sample data amount of each of the plurality of agent ends; and performing weighted calculation on the plurality of first agent end model parameters based on the aggregation weight of the plurality of first agent end model parameters to obtain the global model parameter.
8. A power grid system, characterized by It includes: a service end and a plurality of agent ends; The agent end is configured to perform a first stage training operation on an agent end model to obtain a first agent end model parameter, wherein the agent end model is used to obtain a power equipment action for regulating power equipment in a target active substation in the power grid system by processing the substation operation state of the target active substation. The server is configured to aggregate the first agent-end model parameters of the multiple agent ends to obtain global model parameters; The multiple agent ends are configured to perform second-stage training operations on the global agent model based on the global model parameters to obtain second agent-end model parameters of the agent ends; The server is configured to cluster the second agent-end model parameters of the multiple agent ends based on a clustering algorithm to obtain multiple center model parameters; The agent end is configured to determine target agent-end model parameters from the multiple center model parameters, and to use a target agent-end model determined based on the target agent-end model parameters to regulate power equipment associated with the agent end to perform power equipment actions; The agent end is further configured to: update the global agent model based on target center model parameters in the multiple center model parameters to obtain an initial center agent-end model; process sample regional operating states based on the initial center agent-end model to obtain experience power equipment actions; process the experience power equipment actions based on a loss function to obtain an experience loss value; and determine the target center model parameters as the target agent-end model parameters in response to the experience loss value satisfying a preset loss threshold condition; wherein multiple associated agent ends are associated with the same target center model parameters, the multiple associated agent ends form a same agent-end clustering cluster, and the multiple agent ends include the associated agent ends, The power grid system is further configured to: In a preset training phase, the associated agent ends transmit associated target agent-end model parameters of associated target agent-end models to the server; The server aggregates the multiple current associated target agent-end model parameters to obtain associated global model parameters; and The associated agent ends use a new target agent model determined based on the associated global model parameters to regulate power equipment of the target active regional area to perform power equipment actions; wherein, before the preset training phase, the associated agent ends perform training operations on the respective target agent-end models based on a reinforcement learning mechanism to obtain associated target agent-end model parameters related to the preset training phase.
Citation Information
Patent Citations
Federal learning-based model training method and clustering analyzer
CN115510936A
Comprehensive energy system optimal scheduling method and system based on federal reinforcement learning
CN117151308A