Multi-data center microgrid cooperative operation method based on improved federal reinforcement learning

By employing an improved federated reinforcement learning approach, combined with counterfactual action evaluation, a dual Critic network, and an adaptive differential privacy protection mechanism, the data privacy and heterogeneity issues in multi-microgrid collaborative scheduling are addressed. This enables an efficient and stable energy management strategy, enhancing the system's collaborative optimization efficiency and privacy protection capabilities.

CN121749216BActive Publication Date: 2026-05-19INST OF ELECTRICAL ENG CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF ELECTRICAL ENG CHINESE ACAD OF SCI
Filing Date
2026-02-27
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing federated reinforcement learning methods suffer from problems such as data privacy leakage risks, low aggregation efficiency, and unstable model performance in the collaborative optimization scheduling of multiple microgrids. In particular, it is difficult to achieve efficient and fair collaborative scheduling in the collaborative learning among heterogeneous microgrids.

Method used

An improved federated reinforcement learning approach is adopted, which optimizes the energy management strategy of multi-data center microgrids by combining counterfactual action evaluation, dual Critic networks and contribution awareness mechanism with adaptive differential privacy protection and personalized model fusion mechanism, so as to achieve efficient collaborative operation under privacy protection.

Benefits of technology

It improves the efficiency and economy of collaborative optimization of multi-microgrid systems, enhances the adaptability and stability of the model, achieves a dynamic balance between privacy protection and model performance, and improves the accuracy of decision-making and the overall operational stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121749216B_ABST
    Figure CN121749216B_ABST
Patent Text Reader

Abstract

The application discloses a multi-data center micro-grid cooperative operation method based on improved federal reinforcement learning and belongs to the technical field of power system operation control. The method comprises the following steps: constructing a multi-agent reinforcement learning environment required for training of energy management strategies of each micro-grid, setting a state space, an action space and a reward function; improving local reinforcement learning training by using counterfactual action, double Critic network and delayed update strategy; improving the federal learning aggregation process based on a contribution degree perception mechanism and a FedAdam algorithm; combining adaptive differential privacy protection and an individualized model fusion mechanism, dynamically adjusting the privacy protection strength and differentially issuing global model parameters to guide strategy updating of each micro-grid. While protecting data privacy, the application effectively solves the client heterogeneity problem in multi-micro-grid cooperative scheduling and improves the economy and stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system operation and control technology, specifically relating to a multi-data center microgrid collaborative operation method based on improved federated reinforcement learning. Background Technology

[0002] With the increasing penetration of renewable energy and the accelerated digitalization of power systems, microgrids, as autonomous units integrating distributed energy, energy storage, and flexible loads, have become a key carrier for improving regional power supply reliability and energy utilization efficiency. Interconnecting multiple geographically proximate microgrids to form a multi-microgrid system, through internal energy sharing and coordinated dispatch, can further achieve the goals of peak shaving and valley filling, reducing operating costs, and enhancing overall resilience.

[0003] However, coordinated optimization and scheduling of multiple microgrids faces severe challenges. First, each microgrid belongs to a different operator, and their internal load characteristics, resource composition, and operational objectives differ significantly, exhibiting high heterogeneity. Traditional centralized optimization methods require the aggregation of all raw operational data, posing a significant risk of data privacy breaches and making it difficult to gain the trust and cooperation of the various operators. Second, the system state and market environment are dynamically changing, and scheduling decisions are characterized by high dimensionality, continuity, and strong uncertainty. Traditional mathematical programming methods are difficult to solve in real time, while heuristic algorithms are prone to getting trapped in local optima.

[0004] To address these issues, existing technologies have proposed distributed artificial intelligence solutions. Reinforcement learning (RL) has been introduced into the field of microgrid scheduling due to its excellent sequential decision-making and self-learning capabilities; however, its centralized training framework still requires the collection of data across the entire domain, and privacy risks remain. To address this, federated learning (FL), as a privacy-preserving distributed machine learning paradigm, is used in conjunction with FL. By training locally on the terminal and only uploading model parameters to the server for aggregation, collaborative learning can theoretically be achieved while protecting data privacy. For example, existing patents (such as Chinese Patent Publication No. CN120566600A) propose a multi-microgrid control method based on federated hierarchical reinforcement learning, which optimizes intra-cluster scheduling through dynamic clustering and game theory. However, existing federated reinforcement learning methods still have significant drawbacks: First, during federated aggregation, simple averaging or fixed weights are typically used, failing to fully consider the different contributions of different microgrid agents to the global model due to differences in data quality and environmental conditions. This leads to low aggregation efficiency, slow convergence, and even damage to model performance due to "client drift." Secondly, privacy protection mechanisms are often static or excessive. When adding noise for differential privacy protection, they fail to adaptively balance with the model training state and aggregation effect, which can easily cause unnecessary loss of model performance.

[0005] Therefore, designing an efficient and fair federated aggregation mechanism while strictly protecting data privacy, and enhancing the adaptability of the global model to heterogeneous clients, has become a core technical challenge that urgently needs to be addressed to improve the collaborative scheduling performance of multiple microgrids. Summary of the Invention

[0006] To address the aforementioned technical issues, this invention provides a multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning. This method uses improved reinforcement learning to evaluate the actual impact of each agent's actions on the collaborative operation of the multi-datacenter microgrid, and uses this evaluation to correct subsequent local policy update objectives, guiding each datacenter microgrid to adopt more reasonable energy management methods. Simultaneously, the improved federated learning evaluates the actual performance of each agent through a contribution-aware mechanism and combines adaptive momentum for global model updates, improving the collaborative energy management performance of the multi-datacenter microgrid system. Furthermore, the adaptive differential privacy protection and personalized model fusion mechanism employed in this invention can adaptively adjust the privacy protection strength based on the current federated aggregation effect and distribute global model parameters differentially to guide the energy management strategy updates of each datacenter microgrid, thereby enhancing the adaptability of the global model to agents in different datacenter microgrids.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning includes:

[0009] Step 1: Construct the multi-agent reinforcement learning environment required for training the energy management strategies of microgrids in each data center;

[0010] Step 2: Improve the reinforcement learning algorithm by adopting counterfactual action evaluation, dual Critic network and delayed update strategy, and train the energy management strategy of each data center microgrid locally based on the improved reinforcement learning algorithm;

[0011] Step 3: Improve the federated learning algorithm by adopting contribution awareness and FedAdam aggregation mechanism, and perform cloud aggregation and joint training on the local model parameters uploaded from each data center based on the improved federated learning algorithm;

[0012] Step 4: Based on the joint training results in the cloud, integrate the adaptive differential privacy protection mechanism and the personalized model fusion mechanism to distribute global model parameters to the microgrid agents in each data center in a differentiated manner, so as to guide the local energy management strategy update of each data center microgrid.

[0013] Furthermore, in step 1:

[0014] A state space is defined for each data center microgrid agent, comprising a global state and a local state. The global state includes the current time step and the market transaction prices of electrical and thermal energy. The local state includes the energy storage charge status, electrical / thermal load demand, and wind and solar power output.

[0015] An action space is defined for each data center microgrid agent, which includes control instructions for energy storage charging and discharging, micro-gas turbine output, boiler output, and electricity and heat pricing strategies in the peer-to-peer energy market;

[0016] Define a reward function for each data center microgrid agent that aims to maximize the net operating benefit of the microgrid.

[0017] Furthermore, the counterfactual action evaluation in step 2 specifically involves: calculating the difference between the real reward obtained by the agent in the current state from performing the current action and the counterfactual reward that could be obtained by performing an alternative action, to obtain the counterfactual causal effect; and mapping the causal effect to the temperature parameter through the Sigmoid function to obtain causal weights, which are used to correct the update target of the agent's policy network.

[0018] Furthermore, in step 2, a dual Critic network is used, specifically: each data center microgrid agent deploys two independent Critic value networks. When calculating the time-series difference target, the smaller of the estimated values ​​output by the two independent Critic value networks is taken, and the updated target is calculated based on the smaller value.

[0019] Furthermore, the contribution perception mechanism in step 3 specifically involves: calculating a local model update similarity index based on the directional similarity between the local model parameter update increment of each agent and the global average update increment; calculating a task performance index through nonlinear mapping based on the relative relationship between the cumulative reward of each agent and the average cumulative reward; and calculating and normalizing the contribution weight of each agent by combining the similarity index and the task performance index.

[0020] Furthermore, the FedAdam aggregation mechanism in step 3 specifically involves: weighting the local model update increments from each agent based on the contribution weights to obtain the global model increment; updating the first-order momentum and second-order momentum of the federated server using the global model increment; and calculating the new round of global model parameters based on the updated first-order momentum and second-order momentum and the server-side learning rate.

[0021] Furthermore, the adaptive differential privacy protection mechanism in step 4 specifically involves: dynamically adjusting the standard deviation of the noise injected into the global model parameters based on the average reward value after multiple rounds of federated aggregation; reducing the noise intensity to optimize the global model performance when the average reward value increases; and increasing the noise intensity to promote exploration when the average reward value stabilizes or decreases.

[0022] Furthermore, the personalized model fusion mechanism in step 4 is as follows: for each data center microgrid agent, the global model parameters processed by differential privacy and its local model parameters after the previous round of training are weighted and fused according to the preset personalization rate hyperparameter to generate personalized local model parameters suitable for the agent.

[0023] In a second aspect, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned multi-datacenter microgrid cooperative operation method based on improved federated reinforcement learning.

[0024] Thirdly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned multi-datacenter microgrid cooperative operation method based on improved federated reinforcement learning.

[0025] The beneficial effects of this invention are as follows:

[0026] Improved Collaborative Efficiency and Fairness: By introducing a contribution-aware mechanism jointly determined by model update similarity and task performance, and combining it with FedAdam optimizer for federated aggregation, the differentiated contributions of different microgrids to the global model can be accurately quantified and rewarded. This effectively solves the "uneven contribution" problem caused by the traditional federated averaging method ignoring client heterogeneity, accelerates model convergence speed, and significantly improves the overall collaborative optimization efficiency and economy of multi-microgrid systems.

[0027] Privacy-Performance Dynamic Balancing: An adaptive differential privacy protection mechanism is employed, which dynamically adjusts the intensity of added noise based on the real-time performance (such as average reward) of federated training. Noise is reduced in the early stages of model training or when performance is poor to promote learning, and noise is increased to enhance privacy after performance stabilizes. This achieves an intelligent dynamic balance between privacy protection strength and model performance, overcoming the shortcomings of excessive performance loss or insufficient protection under a fixed privacy budget.

[0028] Enhanced Model Personalization and Stability: Through a personalized model fusion mechanism, local model parameters are differentially integrated when the global model is distributed. This ensures that the strategies acquired by each microgrid absorb global collaborative knowledge while retaining their own operational characteristics. This significantly reduces strategy oscillations caused by direct overlay of model parameters, enhancing the adaptability of scheduling strategies to the local environment and the stability of system operation.

[0029] Improved accuracy of decision guidance: By introducing counterfactual action evaluation into local reinforcement learning training, each microgrid agent can more accurately understand the actual causal impact of its own actions on global rewards, thereby guiding it to learn energy management strategies that are more conducive to the overall system's coordination and optimizing decision quality from the source. Attached Figure Description

[0030] Figure 1 This is a flowchart of the multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning, as described in this invention.

[0031] Figure 2 This is a diagram illustrating a multi-data center microgrid application in an embodiment of the present invention.

[0032] Figure 3 This is a schematic diagram of wind power and photovoltaic data used in embodiments of the present invention;

[0033] Figure 4 This is a schematic diagram of electrical load and thermal load data used in an embodiment of the present invention;

[0034] Figure 5 A comparison chart of the total reward convergence curves under different algorithms;

[0035] Figure 6 A statistical comparison chart of reward values ​​after convergence for different algorithms. Detailed Implementation

[0036] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0037] like Figure 1 As shown, this invention provides a method for cooperative operation of multi-datacenter microgrids based on improved federated reinforcement learning, which mainly includes the following steps:

[0038] Step 1: Construct the multi-agent reinforcement learning environment required for training the energy management strategies of microgrids in each data center;

[0039] Step 2: Improve the reinforcement learning algorithm by adopting counterfactual action evaluation, dual Critic network and delayed update strategy, and train the energy management strategy of each data center microgrid locally based on the improved reinforcement learning algorithm;

[0040] Step 3: Improve the federated learning algorithm by adopting contribution awareness and FedAdam aggregation mechanism, and perform cloud aggregation and joint training on the local model parameters uploaded from each data center based on the improved federated learning algorithm;

[0041] Step 4: Based on the joint training results in the cloud, integrate the adaptive differential privacy protection mechanism and the personalized model fusion mechanism to distribute global model parameters to the microgrid agents in each data center in a differentiated manner, so as to guide the local energy management strategy update of each data center microgrid.

[0042] Furthermore, in step 1, the multi-agent reinforcement learning environment required for training the energy management strategies of microgrids in each data center is constructed, and the specific steps are as follows:

[0043] Step 1-1: Set up the state space:

[0044] Each data center microgrid is considered as an intelligent agent, and the observable state space of the i-th data center microgrid is... for:

[0045] (1)

[0046] In equation (1), This refers to the overall global state of the system. This represents the local state of the i-th data center microgrid.

[0047] The global state is environmental information shared by all data center microgrids, including the current time step. Electricity purchase price Electricity sales price Heat energy purchase price and the price of selling heat energy The global state is shown in the following equation:

[0048] (2)

[0049] Local status refers to the internal information of each data center microgrid, including the energy storage state of charge. Electricity load demand Heat load demand Wind power output and photovoltaic power output The local state is shown in the following formula:

[0050] (3)

[0051] Step 1-2: Set up the motion space:

[0052] The action space of the i-th data center microgrid is the set of operations it can execute, as shown in the following equation:

[0053] (4)

[0054] In equation (4): For energy storage charging and discharging control signals; This is the output control signal for the micro gas turbine; This is the output control signal for the hot boiler; Pricing strategies for the P2P electricity market; Pricing strategies for the P2P heat market.

[0055] Steps 1-3: Set the reward function:

[0056] To guide each data center microgrid to maximize its own economic benefits, the reward function is designed as the net revenue of the i-th data center microgrid at time t, as shown below:

[0057] (5)

[0058] In equation (5), Let be the net revenue of the i-th data center microgrid at time t. Let t be the total revenue of the i-th data center microgrid; Let be the total cost of the i-th data center microgrid at time t.

[0059] Total Revenue The composition is as follows:

[0060] (6)

[0061] (7)

[0062] (8)

[0063] (9)

[0064] In equations (6)-(9), The revenue generated by the i-th data center microgrid from providing electricity and heat to the local area; The revenue generated by the i-th data center microgrid from selling electricity and heat to the P2P energy market; The revenue generated by the i-th data center microgrid from the sale of electricity and heat to the main grid; , These are the prices for supplying electricity and heat to the local area, respectively. , These refer to the power supplied to the local area, specifically electrical and thermal energy. , These refer to the prices of electricity and heat traded in the P2P energy market, respectively. , These refer to the power of electricity and heat traded in the P2P energy market, respectively. , These are the prices for selling electricity and heat to the main grid, respectively. , These represent the power of electricity and heat sold to the main grid, respectively.

[0065] Total cost The composition is as follows:

[0066] (10)

[0067] (11)

[0068] (12)

[0069] (13)

[0070] In equations (10)-(13), The energy production cost of the i-th data center microgrid includes the operating costs of wind power, photovoltaics, micro gas turbines, and thermal boilers; The cost of purchasing electricity and heat from the P2P energy market for the i-th data center microgrid; The cost of purchasing electricity and heat from the main grid for the i-th data center microgrid. , , , These are the unit energy production costs for wind power, photovoltaic power, micro gas turbines, and thermal boilers, respectively. , , , These represent the power generated by wind power, photovoltaic power, micro gas turbines, and thermal boilers, respectively. , These refer to the power of electricity and heat purchased in the P2P energy market, respectively. , These represent the power of electrical and thermal energy purchased from the main network, respectively.

[0071] Further, step 2 includes:

[0072] Step 2-1: Reset the training environment and obtain the initialization status.

[0073] The training environment is restored to its pre-set initial configuration, thus eliminating the impact of the previous training or interaction process. After the environment is reset, observable states such as electricity trading prices, heat trading prices, energy storage state of charge, electricity load demand, heat load demand, wind power output, and photovoltaic power output are acquired. Simultaneously, the Actor and Critic network parameters are initialized for use in subsequent reinforcement learning training.

[0074] Step 2-2: Input the current state into the Actor network to obtain the current action.

[0075] The current environmental state is input into the Actor network for forward inference computation. The Actor network extracts features and performs nonlinear mapping on the state information based on the trained parameters, thereby outputting a continuous control strategy that satisfies action constraints. This strategy can provide a basis for subsequent environmental interactions and state evolution.

[0076] Steps 2-3: Perform actions in the environment to obtain the next state and reward.

[0077] The control policy output by the Actor network is applied to the training environment. The training environment evolves and updates its internal state based on the system dynamics model and operational constraints, thereby obtaining the next state after the action is executed. Then, the total system reward after the action is executed is calculated according to the reward function set in steps 1-3, thus forming a complete interactive process of "state-action-reward-next state". This process can provide a basis for subsequent policy evaluation and parameter updates.

[0078] Steps 2-4: Store information such as status, actions, and rewards into the experience replay pool.

[0079] The "state-action-reward-next state" sequence obtained during the agent's interaction with the environment is encapsulated as experience samples and stored in a priority experience replay pool. The priority experience replay pool assigns sampling priorities to experience samples based on their importance, ensuring that key samples with a greater impact on training are selected with higher probability, thereby improving the training efficiency of reinforcement learning.

[0080] Steps 2-5: Construct counterfactual actions to train the Actor and Critic networks through experience replay.

[0081] Draw M sample data from the priority experience replay pool and calculate the update target of the Actor network. :

[0082] (14)

[0083] In equation (14): This represents the observable current state of the i-th data center microgrid; For the current action that can be performed by the i-th data center microgrid; For Actor networks; , These are the first and second decentralized Critic networks, respectively.

[0084] By constructing counterfactual causal weights, the Actor network update objective is revised to guide the data center microgrid operation strategy towards a more rational update direction. The revised Actor network update objective is: :

[0085] (15)

[0086] In equation (15), E represents the mathematical expectation symbol. The causal weight of the current action on the system reward is used to guide the policy objective appropriately. Its calculation process is as follows:

[0087] (16)

[0088] In equation (16): The Sigmoid function is used to map counterfactual causal effects to... ; The temperature parameter is used to adjust the sensitivity of the causal weights to counterfactual causal effects. The counterfactual causal effect value is calculated as follows:

[0089] (17)

[0090] In equation (17): The reward value is based on the actual action. Counterfactual actions taken; The reward value for counterfactual actions.

[0091] Based on the M sample data extracted above, calculate the update target of the Critic network. :

[0092] (18)

[0093] In equation (18): The target Q-value output by the Critic target network is calculated as follows:

[0094] (19)

[0095] In equation (19): The instantaneous reward for the i-th data center microgrid at time t; As a discount factor, , These are the target Q values ​​output by the first and second Critic networks, respectively. For the target Actor network; , These are the first Critic target network and the second Critic target network, respectively.

[0096] Both the Acror and Critic networks update their network parameters using gradient descent.

[0097] (20)

[0098] (twenty one)

[0099] In equations (20)-(21): is the learning rate of the Actor network; The learning rate of the Critic network. This is the gradient operator.

[0100] Once the federation interval is reached, proceed to step 3, where local model parameters are uploaded to conduct collaborative training with other data center microgrids on the federation server. Otherwise, repeat steps 2-1 to 2-5 until the maximum number of training iterations is reached.

[0101] Further, step 3 includes:

[0102] Step 3-1: Upload the local model parameters of all agents.

[0103] After completing local training, each agent uploads its updated Acror and Critic network parameters to the federated server. The federated server then aggregates the Acror and Critic network parameters (i.e., local model parameters) of each agent to obtain the training results of all agents during distributed execution, thus achieving knowledge sharing and collaborative optimization while protecting privacy.

[0104] Step 3-2: Calculate the contribution weight of each agent.

[0105] The contribution weight of the i-th agent is determined by two metrics: model update similarity and task performance. Model update similarity quantifies the consistency between the i-th agent's model update direction and the global update direction; its calculation process is as follows:

[0106] (twenty two)

[0107] In equation (22): This represents the model increment for the i-th agent; For local model parameters; These are the parameters of the previous round of global model.

[0108] The average model increment for all agents is:

[0109] (twenty three)

[0110] In equation (23): This represents the average model increment; The number of agents participating in the federated aggregation.

[0111] Cosine similarity between the i-th agent and the average model increment for:

[0112] (twenty four)

[0113] (25)

[0114] In equations (24)-(25): To update the similarity for the model, it needs to be mapped to Interval.

[0115] Task performance is used to quantify the local model performance of the i-th agent in the current state, and its calculation process is as follows:

[0116] (26)

[0117] In equation (26): The task performance of the i-th agent; This is the cumulative reward for the i-th agent; The average cumulative reward for all agents; This is a nonlinear mapping function, and its calculation process is as follows:

[0118] (27)

[0119] In the formula: The input value is the nonlinear mapping function; This is the output value of the nonlinear mapping function.

[0120] The contribution weight of the i-th agent is calculated as follows:

[0121] (28)

[0122] (29)

[0123] In equations (28)-(29): The contribution weights before normalization; This represents the normalized contribution weight.

[0124] Step 3-3: Obtain global model parameters using the Fedadadam mechanism.

[0125] Calculate the global model increment based on contribution weights:

[0126] (30)

[0127] Update first-order momentum and second-order momentum:

[0128] (31)

[0129] (32)

[0130] In equations (31)-(32): , These represent the first and second momentum of the FedAdam algorithm at time t, respectively. , These represent the first and second momentum of the FedAdam algorithm at time t-1, respectively. , These represent different momentum decay coefficients.

[0131] Get global model parameters:

[0132] (33)

[0133] In equation (33): The server-side learning rate; It is a very small constant.

[0134] Further, step 4 includes:

[0135] Step 4-1: Add adaptive differential privacy noise to the global model parameters.

[0136] To prevent privacy attacks during the collaborative operation of multi-datacenter microgrids, this invention introduces an adaptive differential privacy protection mechanism on the federated server side. Specifically, adjustable noise is added to the global model parameters before they are distributed to each datacenter microgrid.

[0137] (34)

[0138] In equation (34): The global model parameters after adding adjustable noise; This represents the noise standard deviation, used to control the strength of differential privacy protection. A larger one... It can provide stronger privacy protection, but it will also reduce the effectiveness of federal aggregation.

[0139] The noise intensity can be dynamically adjusted based on the federated aggregation effect to achieve a dynamic balance between privacy protection and model performance. The specific adjustment process is as follows:

[0140] (35)

[0141] In equation (35): Let be the noise standard deviation at time t+1; Let be the noise standard deviation at time t; Noise attenuation rate; This represents the average reward over the most recent k rounds.

[0142] Step 4-2: Use a personalized fusion mechanism to correct the global model parameters.

[0143] Considering the heterogeneity among the microgrids of various data centers, this invention modifies the global model parameters to retain a certain degree of local characteristics, thereby reducing policy oscillations caused by the distribution of model parameters. The specific calculation process is as follows:

[0144] (36)

[0145] In equation (36): The personalization rate hyperparameter is used to weightedly fuse and generate a local model.

[0146] Step 4-3: Distribute global model parameters to all agents.

[0147] The global model parameters are distributed to the microgrid agents in each data center, and it is determined whether the maximum number of iterations has been reached. If the maximum number of iterations has not been reached, the process returns to step 2 to guide the next round of reinforcement learning training; when the maximum number of iterations has been reached, the process ends.

[0148] To verify the effectiveness of the proposed method in the collaborative operation of multi-data center microgrids, this invention... Figure 2 The analysis is conducted using a multi-datacenter microgrid example. This example consists of three datacenter microgrids, each configured with different capacities of photovoltaic, wind power, micro-gas turbines, thermal boilers, and energy storage devices. Each datacenter microgrid can trade electricity and heat through a P2P energy market. When the multi-datacenter microgrid system experiences energy surplus or shortage, it can also interact with external networks for energy exchange. Wind power and photovoltaic data are analyzed as follows: Figure 3 As shown, the electrical load and thermal load data are as follows: Figure 4The values ​​shown are all measured values ​​from a region in eastern my country. The parameter settings for each data center microgrid are shown in Table 1, and the parameter settings for the method proposed in this invention are shown in Table 2.

[0149] Table 1

[0150]

[0151] Table 2

[0152]

[0153] To accurately evaluate the performance advantages of the proposed method, the present invention sets up the following four schemes for comparative analysis.

[0154] Option 1: Use the Independent-MADDPG (I-MADDPG) algorithm to solve the problem. Each agent is trained independently based on local information and can only indirectly influence each other through shared environment.

[0155] Solution 2: The Counterfactual Independent-MADDPG (CI-MADDPG) algorithm is used for solving the problem. Based on Solution 1, counterfactual actions, a dual-critic network, and a delayed policy update mechanism are introduced.

[0156] Option 3: Use the Federal-MARL (F-MARL) algorithm to solve the problem. Based on Option 2, a federated learning mechanism is introduced to achieve the co-evolution of multi-agent policies through average parameter aggregation.

[0157] Solution 4: The proposed Contribution-Aware Federated-MARL (CAF-MARL) algorithm is used for solving the problem. Based on Solution 3, mechanisms such as contribution awareness, Fedadadam aggregation, and adaptive differential privacy protection are introduced.

[0158] Figure 5The graph shows the convergence curves of the total reward under different algorithms. The solid line represents the moving average reward value over the last 10 rounds, and the shaded area represents the fluctuation range of the reward value. As can be seen from the graph, the I-MADDPG algorithm relies entirely on local information for agent training, thus easily getting trapped in local optima in a multi-datacenter microgrid environment. The CI-MADDPG algorithm corrects the local policy gradient based on counterfactual actions, guiding the agent to update the policy in a more reasonable direction. However, the local decision layer struggles to perceive the global optimization direction of the system, resulting in poor collaborative optimization among multi-datacenter microgrids. The F-MARL algorithm achieves collaborative operation among multi-datacenter microgrids through average federated aggregation. However, the federated server struggles to distinguish the performance differences of each agent during actual training, making it susceptible to interference from low-quality samples and leading to slow convergence. In contrast, the proposed CAF-MARL algorithm exhibits the best overall performance, mainly due to the introduced contribution awareness and FedAdam aggregation mechanism. The contribution awareness mechanism prioritizes agent parameters that align with the global model direction and have recently performed well, thereby reducing the risk of local overfitting in the early training stages. The FedAdam aggregation mechanism can achieve faster convergence speed and operational stability by adaptively smoothing the federated aggregation process with momentum.

[0159] The reward statistics after convergence of different algorithms are as follows: Figure 6 As shown. From Figure 6 As can be seen, the CI-MADDPG algorithm can guide the agent to update gradients in the correct policy direction, thus achieving a better average reward value after convergence compared to the I-MADDPG algorithm. Furthermore, the CI-MADDPG algorithm avoids oscillations caused by Q-value overestimation and frequent policy updates, resulting in lower reward value volatility after convergence compared to the I-MADDPG algorithm. The F-MARL algorithm employs an average federated aggregation strategy, effectively improving the collaborative operation capability among multi-data center microgrids, thereby further increasing the average reward value after convergence. In contrast, the proposed CAF-MARL algorithm exhibits the largest average reward value and the smallest reward value volatility after convergence, primarily due to the introduced FedAdam aggregation and personalized fusion mechanism. This mechanism can adaptively smooth the aggregation process of federated learning through first and second-order momentum, while also considering the differentiated needs of different data center microgrids, effectively addressing policy oscillations caused by model parameter aggregation.

[0160] In a second aspect, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned multi-datacenter microgrid cooperative operation method based on improved federated reinforcement learning.

[0161] Thirdly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned multi-datacenter microgrid cooperative operation method based on improved federated reinforcement learning.

[0162] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-datacenter microgrid cooperative operation method based on improved federated reinforcement learning, characterized in that, include: Step 1: Construct the multi-agent reinforcement learning environment required for training the energy management strategies of microgrids in each data center; Step 2: Improve the reinforcement learning algorithm by adopting counterfactual action evaluation, dual Critic network and delayed update strategy, and train the energy management strategy of each data center microgrid locally based on the improved reinforcement learning algorithm; Step 3: Improve the federated learning algorithm using contribution awareness and FedAdam aggregation mechanisms, and perform cloud aggregation and joint training on the local model parameters uploaded from each data center based on the improved federated learning algorithm; the contribution awareness mechanism specifically involves: calculating a local model update similarity index based on the directional similarity between the local model parameter update increment of each agent and the global average update increment; calculating a task performance index through nonlinear mapping based on the relative relationship between the cumulative reward of each agent and the average cumulative reward; and calculating and normalizing the contribution weight of each agent by combining the similarity index and the task performance index. The FedAdam aggregation mechanism is as follows: the local model update increments from each agent are weighted based on the contribution weights to obtain the global model increment; the first-order momentum and second-order momentum of the federated server are updated using the global model increment; and the new round of global model parameters are calculated based on the updated first-order momentum and second-order momentum and the server-side learning rate. Step 4: Based on the joint training results in the cloud, integrate the adaptive differential privacy protection mechanism and the personalized model fusion mechanism to distribute global model parameters to the microgrid agents in each data center in a differentiated manner, so as to guide the local energy management strategy update of each data center microgrid.

2. The multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning according to claim 1, characterized in that, In step 1: A state space is defined for each data center microgrid agent, comprising a global state and a local state. The global state includes the current time step and the market transaction prices of electrical and thermal energy. The local state includes the energy storage charge status, electrical / thermal load demand, and wind and solar power output. Define an action space for each data center microgrid agent, the action space including control instructions for energy storage charging and discharging, micro-gas turbine output, thermal boiler output, and electricity and heat pricing strategies in the peer-to-peer energy market; Define a reward function for each data center microgrid agent that aims to maximize the net operating benefit of the microgrid.

3. The multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning according to claim 1, characterized in that, The counterfactual action evaluation in step 2 specifically involves: calculating the difference between the real reward obtained by the agent in the current state from performing the current action and the counterfactual reward that could be obtained by performing an alternative action, to obtain the counterfactual causal effect; The causal effects are then mapped to causal weights using a Sigmoid function and temperature parameters, which are used to correct the update target of the agent policy network.

4. The multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning according to claim 1, characterized in that, In step 2, a dual Critic network is used. Specifically, each data center microgrid agent deploys two independent Critic value networks. When calculating the time-series difference target, the smaller of the output values ​​of the two independent Critic value networks is taken, and the updated target is calculated based on the smaller value.

5. The multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning according to claim 1, characterized in that, The adaptive differential privacy protection mechanism in step 4 is as follows: dynamically adjust the standard deviation of the noise injected into the global model parameters according to the average reward value after multiple rounds of federated aggregation; reduce the noise intensity to optimize the global model performance when the average reward value increases; and increase the noise intensity to promote exploration when the average reward value is stable or decreases.

6. The multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning according to claim 1, characterized in that, The personalized model fusion mechanism in step 4 is as follows: for each data center microgrid agent, the global model parameters processed by differential privacy and the local model parameters after the previous round of training are weighted and fused according to the preset personalization rate hyperparameter to generate personalized local model parameters suitable for the agent.

7. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when one or more programs are executed by the one or more processors, the one or more processors implement the multi-datacenter microgrid cooperative operation method based on improved federated reinforcement learning as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, enable the processor to implement the multi-datacenter microgrid collaborative operation method based on improved federated reinforcement learning as described in any one of claims 1-6.