A comprehensive energy intelligent agent collaborative fault-tolerant group control method for improving energy system resilience

By dividing the integrated energy intelligent agent into multiple intelligent agent subsystems and constructing a leader-follower model, combined with Q-learning algorithm and evolutionary game theory, the problem of insufficient grid resilience is solved, enabling the grid to respond quickly and adaptively under equipment failure, thereby improving the grid's resilience and regulation efficiency.

CN121036079BActive Publication Date: 2026-03-24STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing integrated energy intelligence systems lack sufficient regulation capabilities to support grid resilience, especially in terms of group regulation and control, which makes it difficult for the grid to achieve rapid response and adaptive adjustment when facing complex and ever-changing external environments.

Method used

The integrated energy intelligent agent is divided into multiple intelligent agent subsystems. A leader-follower model is constructed and the Q-learning algorithm is used in conjunction with evolutionary game theory to achieve intra-group autonomy and inter-group coordination. The collaborative fault-tolerant control module maintains consistency under equipment failure and optimizes scheduling planning.

Benefits of technology

It has improved the adaptability and resilience of the power grid, optimized the power grid regulation strategy, and ensured the stability and security of power supply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121036079B_ABST
    Figure CN121036079B_ABST
Patent Text Reader

Abstract

The present disclosure provides a comprehensive energy intelligent agent collaborative fault-tolerant group control method for improving the resilience of an energy system. According to the differentiated goals of each intelligent agent in the comprehensive energy intelligent agent, the comprehensive energy intelligent agent is divided into multiple intelligent agent subsystems; for each intelligent agent subsystem, a leader-follower model including a leader, a follower, and a collaborative fault-tolerant control module is constructed to achieve group autonomy; wherein the leader in any intelligent agent subsystem is a control unit, the follower is a controlled device, the followers are neighbor intelligent agents, and the neighbor intelligent agents output consistency under any follower failure through the collaborative fault-tolerant control module; for the group autonomous intelligent agent group, a group control architecture is adjusted, the interest balance of each intelligent agent subsystem is taken as a target, a Q learning algorithm is used to solve the Nash equilibrium point of the game between multiple intelligent agent subsystems by using evolutionary game theory, and the comprehensive energy intelligent agent is scheduled and planned according to the solution result to achieve inter-group coordination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent agent control technology, and in particular to a comprehensive energy intelligent agent collaborative fault-tolerant group control method for improving the resilience of energy systems. Background Technology

[0002] With the continuous growth of energy demand and the frequent occurrence of extreme weather events, the resilience of power grids has become crucial to ensuring the security of power supply. Power grid resilience refers to the grid's ability to withstand disruptions and recover rapidly in the face of sudden events such as natural disasters, equipment failures, and external attacks. However, existing power grid control methods have shown certain limitations in dealing with complex and ever-changing external environments.

[0003] Integrated energy systems, as a new energy utilization model, combine multiple energy forms and achieve efficient energy utilization through intelligent management. However, in current technologies, the regulatory capabilities of integrated energy intelligent systems in supporting grid resilience are still insufficient, especially in terms of group dispatch and control, where effective regulation strategies and methods are lacking. Summary of the Invention

[0004] The purpose of this disclosure is to provide a comprehensive energy intelligent agent collaborative fault-tolerant group control method to enhance the resilience of energy systems, solve the problem of insufficient grid resilience under equipment failure, and provide strong protection for the safe and stable operation of the power grid.

[0005] According to the first aspect of this disclosure, a collaborative fault-tolerant group control method for integrated energy intelligent agents is provided to enhance the resilience of energy systems. The method includes: dividing the integrated energy intelligent agent into an intelligent agent group control architecture that includes multiple intelligent agent subsystems based on the differentiated objectives of each intelligent agent in the integrated energy intelligent agent.

[0006] For each intelligent agent subsystem, a leader-follower model including a leader, followers, and a collaborative fault-tolerant control module is constructed to achieve group autonomy. In any intelligent agent subsystem, the leader is the control unit, the followers are the controlled devices, and the followers are neighboring intelligent agents. The neighboring intelligent agents achieve output consistency under the failure of any follower through the collaborative fault-tolerant control module.

[0007] For the autonomous intelligent agent group scheduling and control architecture, with the goal of balancing the interests of each intelligent agent subsystem, the Q-learning algorithm and evolutionary game theory are used to solve the Nash equilibrium point of the game between the multiple intelligent agent subsystems. Based on the solution results, the integrated energy intelligent agent is scheduled and planned to achieve inter-group coordination.

[0008] According to a second aspect of this application, a comprehensive energy intelligent agent collaborative fault-tolerant group control device for enhancing the resilience of energy systems is disclosed, comprising:

[0009] The agent segmentation module is used to divide the integrated energy intelligent agent into an agent group control architecture that includes multiple agent subsystems based on the differentiated goals of each agent in the integrated energy intelligent agent.

[0010] The group autonomy module is used to construct a leader-follower model for each intelligent agent subsystem, including a leader, followers, and a collaborative fault-tolerant control module, to achieve group autonomy. In any intelligent agent subsystem, the leader is the control unit, the followers are the controlled devices, and the followers are neighboring intelligent agents. The neighboring intelligent agents achieve output consistency under the failure of any follower through the collaborative fault-tolerant control module.

[0011] The inter-group coordination module is used for the group coordination and control architecture of intelligent agents in an autonomous group. With the goal of balancing the interests of each intelligent agent subsystem, it uses the Q-learning algorithm and joint evolutionary game theory to solve the Nash equilibrium point of the game between the multiple intelligent agent subsystems. Based on the solution results, it performs scheduling planning for the integrated energy intelligent agent to achieve inter-group coordination.

[0012] According to a third aspect of this disclosure, an electronic device is provided, including a processor and a memory, the memory storing a computer program, the processor executing the computer program stored in the memory to perform the steps of the method described in the first aspect above.

[0013] According to a fourth aspect of this disclosure, a computer program is provided, including computer instructions that, when executed by a processor, implement the steps of the method described in the first aspect above.

[0014] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the steps of the method described in the first aspect above. Attached Figure Description

[0015] Figure 1 A flowchart illustrating a comprehensive energy intelligent agent collaborative fault-tolerant group control method for enhancing the resilience of energy systems, provided as an embodiment of this disclosure;

[0016] Figure 2 This is a schematic diagram of the structure of an intelligent agent group control architecture provided in one embodiment of the present disclosure;

[0017] Figure 3 A schematic diagram of a leader-follower model including a fault-tolerant control module is provided as an embodiment of this disclosure;

[0018] Figure 4 A timing diagram illustrating the control based on the Q-learning algorithm is provided as an embodiment of this disclosure;

[0019] Figure 5 A schematic diagram of the structure of a comprehensive energy intelligent agent collaborative fault-tolerant group control device for enhancing the resilience of energy systems provided in one embodiment of this disclosure;

[0020] Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present disclosure. Detailed Implementation

[0021] Before introducing the embodiments of this disclosure, it should be noted that:

[0022] Some embodiments of this disclosure are described as processing flows. Although the various operational steps of the flow may be numbered sequentially, the operational steps may be performed in parallel, concurrently, or simultaneously.

[0023] The embodiments disclosed herein may use terms such as "first," "second," etc., to describe various features, but these features should not be limited by these terms. These terms are used merely to distinguish one feature from another.

[0024] The term “and / or” may be used in embodiments of this disclosure, and “and / or” includes any and all combinations of one or more of the associated features listed.

[0025] It should be understood that when describing the connection or communication relationship between two components, unless it is explicitly stated that the two components are directly connected or communicate directly, the connection or communication between the two components can be understood as a direct connection or communication, or it can be understood as an indirect connection or communication through an intermediate component.

[0026] To make the technical solutions and advantages of the embodiments of this disclosure clearer, the exemplary embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this disclosure can be combined with each other.

[0027] Integrated energy systems, as a new energy utilization model, integrate multiple energy forms and achieve efficient energy utilization through intelligent management. However, in current technologies, the control capabilities of integrated energy intelligent agents in supporting grid resilience are insufficient, especially in terms of group control and dispatch, where effective control strategies and methods are lacking. Current integrated energy intelligent agent control methods have the following limitations when addressing grid resilience issues: ① Limited control strategies: lacking diversified control strategies for different levels and scenarios; ② Poor information exchange: imperfect information transmission and sharing mechanisms between different levels lead to poor control effects; ③ Insufficient adaptive capabilities: existing control methods struggle to achieve rapid response and adaptive adjustment when facing complex and ever-changing environments.

[0028] Therefore, in order to improve the resilience of the power grid and ensure the stability of power supply, this invention proposes a comprehensive energy intelligent agent collaborative fault-tolerant group control method to enhance the resilience of the energy system. It mainly solves the problems of insufficient grid resilience under equipment failure and low energy system regulation efficiency, and provides a strong guarantee for the safe and stable operation of the power grid.

[0029] like Figure 1 As shown, the integrated energy intelligent agent collaborative fault-tolerant group control method for enhancing the resilience of energy systems includes:

[0030] S101, based on the differentiated objectives of each intelligent agent in the integrated energy intelligent agent, the integrated energy intelligent agent is divided into an intelligent agent group control architecture that includes multiple intelligent agent subsystems;

[0031] S102, For each intelligent agent subsystem, a leader-follower model including a leader, followers, and a collaborative fault-tolerant control module is constructed to achieve group autonomy; In any intelligent agent subsystem, the leader is the control unit, the followers are the controlled devices, the followers are neighboring intelligent agents, and the neighboring intelligent agents achieve output consistency under the failure of any follower through the collaborative fault-tolerant control module.

[0032] S103, for the group-based autonomous intelligent agent group scheduling and control architecture, aims at balancing the interests of each intelligent agent subsystem. It uses the Q-learning algorithm and evolutionary game theory to solve the Nash equilibrium point of the game between the multiple intelligent agent subsystems. Based on the solution results, it performs scheduling planning for the integrated energy intelligent agent to achieve inter-group coordination.

[0033] Using the above approach, based on the differentiated objectives of each agent in the integrated energy intelligent agent, the integrated energy intelligent agent is divided into an intelligent agent group control architecture comprising multiple intelligent agent subsystems. This architecture enables classified management of equipment intelligent agents with different interests. Based on the proposed group control architecture, a collaborative fault-tolerant control module is designed. By adjusting the controllers of normally functioning agents and those that have failed, intra-group autonomy is achieved under equipment failure conditions, improving the adaptive capability and resilience of the power grid and realizing intelligent scheduling and control. At the same time, by utilizing the Q-learning algorithm and evolutionary game theory, inter-group coordination under the equilibrium of interests of multiple intelligent agent subsystems is achieved, optimizing the power grid's regulation strategy.

[0034] In a power grid integrated energy system, different devices and entities play different roles in energy production, conversion, storage, and consumption, and their interests (such as economic benefits, operating costs, and supply-demand balance) vary significantly. This invention's intelligent agent segmentation is based on the principle of profit-driven action. By identifying key participants with differentiated objectives within the system, it abstracts them into independent intelligent agent systems, thereby achieving targeted collaborative optimization.

[0035] In S101 above, specifically, it can involve identifying the differentiated interests of each agent within the integrated energy agent of the power grid, dividing the integrated energy agent into four agent subsystems: a Renewable Energy Provider (REP) responsible for the scheduling and control of distributed power sources such as photovoltaics (PV) and energy storage devices (ES), and its core equipment; an Integrated Energy Provider (IEP) responsible for the scheduling and control of energy conversion equipment such as gas-fired combined cooling, heating and power systems (CCHP), gas-fired heat pumps (GHP), and electric chillers (CAC), as well as thermal storage devices, and its core equipment; an Electric Vehicle Owner (EVO) responsible for the charging and discharging control of electric vehicles (EVs), and its core equipment; and Terminal Energy Users (TEUs) responsible for transmitting demand information of the system's terminal electrical / heating / cooling loads and reducing load when necessary, and their core equipment. The constructed group dispatch and control architecture comprising these four agent subsystems is as follows: Figure 2 As shown.

[0036] The following are the equipment composition, functional descriptions, and benefit control objectives of four types of intelligent subsystems:

[0037] Renewable Energy Suppliers (REPs)

[0038] Core equipment:

[0039] Photovoltaic (PV) power generation system: includes photovoltaic array, inverter, grid-connected controller, etc.

[0040] Wind turbine generator set (WT): wind turbine, converter, energy storage interface.

[0041] Energy storage devices (ES): lithium-ion battery packs, supercapacitors, energy storage converters (PCS), and energy management systems (EMS).

[0042] Control objectives: Maximize revenue from renewable energy generation (electricity sales revenue); optimize energy storage charging and discharging strategies to match grid price fluctuations (low storage, high discharge).

[0043] Managed variables:

[0044] Photovoltaic / wind power output; energy storage charging and discharging power, energy storage capacity status.

[0045] Integrated Energy Service Provider (IEP)

[0046] Core equipment:

[0047] Gas-fired combined cooling, heating and power (CCHP) system: gas turbine, waste heat boiler, absorption chiller.

[0048] Gas-fired heat pump (GHP): A gas-driven heat pump unit.

[0049] Electric Refrigeration Unit (CAC): Electric compression refrigeration equipment.

[0050] Thermal storage equipment (HS): phase change thermal storage tanks, heat exchangers.

[0051] External power grid interaction interface: bidirectional electricity meter, power regulation device.

[0052] Control objective:

[0053] Ensure real-time supply and demand balance of electricity, heat, and cooling loads to minimize operating costs.

[0054] Managed variables:

[0055] CCHP (electric / heating / cooling output), GHP (heating power), CAC (cooling power), and the amount of heat charged and released by the thermal storage equipment.

[0056] Electric vehicle users (EVO)

[0057] Core equipment:

[0058] Electric vehicle cluster (EV): pure electric vehicles and plug-in hybrid electric vehicles.

[0059] Charging station network: fast charging stations, slow charging stations, V2G bidirectional charging and discharging machines.

[0060] Aggregated control platform: charge / discharge scheduling server, user response model algorithm.

[0061] Control objective:

[0062] Optimize user charging and discharging costs (minimize charging costs and maximize discharging benefits) and respond to grid dispatching needs.

[0063] Managed variables:

[0064] Charge / discharge power, battery state of charge (SOC), and user response probability.

[0065] Terminal load users (TEU)

[0066] Core equipment:

[0067] Electrical load terminals: industrial / commercial / residential electrical equipment (such as motors, lighting, and IT equipment).

[0068] Heat load terminals: radiators, hot water circulation system.

[0069] Cooling load terminals: air conditioning terminals, refrigeration equipment.

[0070] Flexible load management devices: interruptible load controllers and demand response terminals.

[0071] Control objective:

[0072] Ensure the supply of basic load (non-reducible portion) and implement load reduction in case of system emergencies.

[0073] Managed variables:

[0074] Electricity / heating / cooling load demand, and load reduction ratio.

[0075] The content of S102 above will be described below.

[0076] like Figure 3The diagram illustrates a leader-follower model for each intelligent agent subsystem, comprising a leader, followers, and a collaborative fault-tolerant control module. The collaborative fault-tolerant control module includes a fault-tolerant controller, a distributed state and fault estimator, and a distributed external system observer. In any intelligent agent subsystem, the leader is the control unit, and the followers are the controlled devices. Specifically, the leader for any intelligent agent subsystem is one of the following: a renewable energy supplier (REP), an integrated energy service provider (IEP), an electric vehicle user (EVO), or a terminal load (TEU). The followers of the intelligent agent subsystem are the core devices corresponding to that leader. As shown in the diagram, the leaders are IEP, REP, EVO, and TEU, and the core devices are the followers. Each follower is a neighboring intelligent agent, and the neighboring intelligent agents achieve output consistency under the failure of any follower through the fault-tolerant control module. When a device fails, the fault-tolerant controllers of both the normally functioning intelligent agents and the failed intelligent agents are adjusted simultaneously, enabling all subsystems to work collaboratively to achieve fault-tolerant control and intra-group autonomy.

[0077] The follower is described by the following continuous-time nonlinear dynamic system:

[0078] (1)

[0079] In the formula, i =1,···, N Represents the number of intelligent agents. x i ∈ n , u i ∈ m and y i ∈ p They represent the first i The status, control inputs and outputs of a comprehensive energy device; g i ( t , x i ( t )): n n This indicates that both the Lipschitz condition and the initial condition g are satisfied simultaneously. i (0, x i A sufficiently smooth nonlinear function (0))=0; f i ( t )∈ rIndicates existence in the first i A fault in a comprehensive energy equipment, and meets the requirements i ( t ) 0; A , B , C E and E are respectively of dimension 1 n × n, n × m , p × n and n × r The real matrix, where the matrix C E and E are in full rank;

[0080] The leader is described by the following continuous-time nonlinear autonomous dynamic system:

[0081] (2)

[0082] In the formula, v ( t )∈ h and y r ( t )∈ p Represent the state and output of the leading agent, respectively; φ( t , v ( t )): h h and q ( t , v ( t )): h p These are respectively satisfying φ(0, v (0))=0 and q (0, v A sufficiently smooth nonlinear function (0))=0; φ( t , v ( t ()) is used to describe the state-space dynamic characteristics of a system, representing the state-space dynamic characteristics in time. t and state v ( t Under this condition, the next state of the system v ( t +1); q( t , v ( t )) indicates time t and state v ( t Under these conditions, the system output y r ( t The leader (2) is used to generate the output reference trajectory and can be regarded as a nonlinear external system in the output regulation problem.

[0083] To verify the consistency tracking effect as time approaches infinity, the output tracking error between the i-th follower and the leader is defined as:

[0084] (3)

[0085] The purpose of this invention is to design a suitable fault-tolerant control module output. ui(t), i=1,···,N Make

[0086] (4)

[0087] Furthermore, the problem of fault-tolerant output consistency under equipment failure can be described as follows:

[0088] Consider a heterogeneous nonlinear multi-agent system consisting of followers (1) and a leader (2). When the integrated energy device, i.e., the follower, fails, a fault-tolerant control module is designed to output... u i ( t ), such that for any initial conditions x i (0) and v (0), there is

[0089] (5)

[0090] This means that the intelligent agent subsystem has achieved cooperative fault-tolerant output consistency.

[0091] The collaborative fault-tolerant control module includes a fault-tolerant controller, a distributed state and fault estimator, and a distributed external system observer. The distributed state and fault estimator is used to estimate the state and faults in the followers. The distributed external system observer is used to provide reference tracking signals for some followers who cannot obtain dynamic information about the leader. The fault-tolerant controller is used to determine the fault-tolerant control input for the followers based on the output of the distributed state and fault estimator.

[0092] Specifically, the design methods for distributed state and fault estimators and distributed external system observers are given using the separation principle. The two work together to achieve group autonomy in a leader-follower form.

[0093] The following are the distributed state and fault estimators for each follower:

[0094] (6)

[0095] In the formula, ∈ n and ∈ r These are the state estimate and the fault estimate, respectively. ∈ n The output value of the estimator, u i ( t For followers during the period t The fault-tolerant control input. The coupling gain κ>0 is an undetermined constant. H ∈ n ×r ,Γ∈ r×r and F ∈ r×p Let Γ be the estimator gain to be determined, where Γ is a positive definite symmetric matrix. ē i (t) ∈ p This represents the measurement output error that depends on local information interaction, specifically in the form of:

[0096] (7)

[0097] In the formula, a ij Represents a directed augmenting directed graph Corresponding adjacency matrix The element. Define the first i The error in estimating the state and fault of each follower is:

[0098] (8)

[0099] According to equation (8), equation (7) can be rewritten as

[0100] (9)

[0101] An adaptive distributed external system observer is given, in the following form:

[0102] (10)

[0103] In the formula, η i ( t )∈ p z i ( t )∈ h It is the observer's state vector. ∈ This represents the adaptive parameter, where α>0 indicates a constant.

[0104] Define the following vector:

[0105] (11)

[0106] Based on this, a fault-tolerant controller is proposed for the above-mentioned intelligent agent subsystem to solve the cooperative fault-tolerant output consistency problem:

[0107] (12)

[0108] In the formula, K 1∈ m×n , K 2∈ m×n and B * ∈ m×n Let B be the gain matrix to be determined, and ∥B∥ be the Euclidean norm of the real matrix B. Furthermore, this formula can be rewritten as a fault-tolerant controller of the following form:

[0109] (13)

[0110] The aforementioned S102 ensures system stability during faults, i.e., ensures intra-group autonomy, providing a reliable operating environment for inter-group coordination in S103. Furthermore, inter-group coordination in S103 can further optimize the control and scheduling of the integrated power grid energy system, effectively improving its resilience. The content of S103 is described below.

[0111] Considering that IEP, REP, and EVO act as active controllers, responsible for energy production, conversion, storage, and trading; while TEU, as a demand-side terminal, is defined as a passive responder, participating in system regulation through load demand signals. The TEU's function focuses on demand transmission and emergency load reduction; its behavior pattern is embedded in the overall system optimization framework through power balance constraints and load reduction rules, achieving its functional objectives without independent modeling. Based on this, in the aforementioned S103, the benefit objectives of the three controllable resource intelligent agent subsystems—REP, IEP, and EVO—can be integrated to establish a multi-agent subsystem benefit optimization model.

[0112] The decision-making models for the three types of controllable resource intelligent agent subsystems, namely REP, IEP, and EVO, are as follows.

[0113] 1) IEP Decision Model

[0114] The Integrated Power Controller (IEP) plays a crucial and leading role in the integrated energy system, undertaking the main supply of end-user loads. The IEP optimizes the output allocation of adjustable equipment across different time periods to reduce operating costs and the expense of purchasing electricity from the external grid. It also handles the purchase of electricity from the external grid during power shortages or when prices are favorable. The IEP's control objective is to minimize total operating costs, and its operation and scheduling benefit function takes the following form:

[0115] (14)

[0116] In the formula, N T Indicates the total number of time periods in the scheduling cycle; p B (t) for t The cost of electricity generated by the system purchasing power from the external power grid during certain time periods; P PG (t) correspond t The amount of electrical energy purchased by the system from the external power grid during a given time period; for t Fuel consumption cost of IEP (Intelligent In-Effect Control) unit during specific time periods;

[0117] Characterization t Economic expenditures incurred during the operation and maintenance of IEP intelligent agent devices during specific time periods; reflect t Additional costs incurred by the start-up and shutdown operations of the IEP intelligent control unit during specific time periods; I HS ( t )for t The cost of wear and tear on time-period thermal storage equipment.

[0118] 2) REP Decision Model

[0119] Based on the electricity sales price for each time period, REP (Regulatory Power Utilization) adjusts the charging / discharging mode of energy storage devices to rationally plan the grid-connected electricity volume of distributed power sources and energy storage systems. Its economic benefits are derived by subtracting the costs of various equipment from the electricity sales revenue. The specific form of the operation and scheduling benefit function is as follows:

[0120] (15)

[0121] In the formula, pS (t) express t The grid connection tariff set by the external power grid operator during the specified time period; correspond t The power output of the REP agent to the external power grid during the specified time period; for t Economic expenditures incurred during the operation and maintenance of REP intelligent agent equipment during specific time periods; I ES ( t )for t Costs corresponding to the operating losses of time-limited energy storage devices; Costs required for equipment procurement.

[0122] 3) EVO Decision Model

[0123] As an aggregator, EVO implements centralized scheduling and management of electric vehicle clusters within the integrated energy system. Under a pre-defined charging and discharging strategy framework, EVO formulates a two-way energy interaction strategy between the vehicle and the grid based on real-time grid electricity price signals. In the constructed price response model, the economic incentive for V2G scheduling becomes more significant when the price difference between the vehicle-to-grid interaction electricity price and the marginal cost of scheduling increases. Combined with an electric vehicle user price sensitivity response model constructed based on consumer psychology, the price sensitivity response characteristics of the EVO agent can be described as follows:

[0124] (16)

[0125] In the formula, α The probability parameter representing the EVO's participation in scheduling; Δ p The price difference between the grid-connected electricity price for electric vehicles set by the grid operator and the benchmark electricity price. Considering the stochastic nature of the response probability, the expression for the EVO benefit function is:

[0126] (17)

[0127] In the formula, P EVd (t) characterization t The discharge power value of the EVO interacting with the main grid during the time period; r The random number is uniformly distributed in the interval [0,1), if and only if r ≥ α (Δ p The user scheduling participation mechanism is triggered when ) is active; [•] represents the rounding up function.

[0128] In the energy scheduling and management of intelligent agents, IEP and REP agents primarily undertake the supply of different energy loads within the integrated energy system, while simultaneously striving to optimize their respective economic costs / benefits. REP and EVO agents possess the capability to engage in bidirectional power trading with the external power grid to generate revenue; however, in pursuing revenue, they must also consider the supply of terminal loads within the system and the constraints of power transmission through the grid connection lines. The different agents, in their pursuit of their optimal benefits, constitute a completely informational static non-cooperative game problem. Based on ensuring the normal operation of the equipment and the overall system, each agent formulates its own operating strategy to achieve its optimal cost / benefit, thereby achieving coordination among the groups. Therefore, the problems involved in the aforementioned game relationship belong to multiple distributed optimization problems where each agent independently optimizes its own objectives, and the Nash equilibrium point reached is the result of this optimization.

[0129] Based on this, in S103 above, the Nash equilibrium point of the game between multiple agent subsystems can be solved by using the Q-learning algorithm in conjunction with evolutionary game theory. This can be achieved by first constructing a multi-agent cooperative game architecture based on the following objective function:

[0130] (18)

[0131] In the formula, G is the equilibrium point reached in the multi-agent game. f (·) is the game function for multiple agents; N g The number of intelligent agent subsystems participating in the game is 3, of which the participating intelligent agent subsystems are IEP, REP, and EVO; A g These are the operational actions of the intelligent agent subsystems participating in the game; i g It is the benefit function I of each intelligent agent subsystem;

[0132] Then, the Q-learning iterative algorithm is used to iterate on the multi-agent collaborative game architecture. For the intelligent subsystems participating in the game, when the selected action constitutes the optimal response strategy under the strategy combination of other intelligent subsystems, it is determined that the Q-learning has converged to the Nash equilibrium point.

[0133] Furthermore, when solving for the Nash equilibrium point of the game among the multiple intelligent agent subsystems, it is necessary to perform a feasibility check on the state-action combinations required for iteration based on power balance constraints, energy supply equipment operation constraints, and energy storage equipment operation constraints. Only valid strategy combinations that meet the equipment constraints and satisfy the game equilibrium conditions are retained for iteration. The three constraints are introduced below.

[0134] 1) Power balance constraints. While making full use of energy storage devices, ensure that the supply and demand of electricity / heat / cooling energy within the integrated energy system can achieve real-time balance:

[0135] (19)

[0136] (20)

[0137] (twenty one)

[0138] In the formula: P ( t )for t Power output values ​​of multi-source units during the time period η PG represents the energy efficiency conversion coefficient of the heterogeneous unit group, load represents the load, and elec, heat, and cold represent electrical energy, thermal energy, and cold energy, respectively. P load ( t )reflect t Energy demand of multiple load nodes during different time periods; the superscript [±] of energy storage equipment defines the charging and discharging working states, where "+" represents the charging state and "-" represents the discharging state.

[0139] 2) Operating constraints of energy supply equipment. When operating, the energy supply equipment configured in the integrated energy system plan must meet the rated power and ramp-up constraints of the equipment:

[0140] (twenty two)

[0141] (twenty three)

[0142] In the formula: Characterizing the first m Energy supply equipment in t The real-time power output value for a given period, its operating boundary is determined by... and limited; Δ and ∆ Defined as the first m The dynamic characteristics of power regulation of similar energy supply equipment correspond to the ramp rate of power reduction and power increase processes, respectively.

[0143] 3) Energy storage device operation constraints. The energy storage devices configured in the integrated energy system plan must meet the dynamic model of power and energy during operation, and satisfy the constraints on charging and discharging power and capacity:

[0144] (twenty four)

[0145] (25)

[0146] Where: M m ( t ) characterize them Group energy storage equipment in t The real-time capacity status parameters for a given time period, and their operating modes are determined by... and Define capacity configuration threshold boundaries; and The first m The charging and discharging power and maximum charging and discharging power of the energy storage equipment.

[0147] The implementation of the Q-learning algorithm described above will be explained below.

[0148] The value function and iterative process of the Q-learning algorithm are expressed as follows:

[0149] (26)

[0150] (27)

[0151] In the formula, s and s Define the current scheduling state and the next timing state respectively; R ( s , s ′, a ) Representing state s and performing actions a Migrate to s The instant reward value obtained; γ ∈ (0,1) is set as the time series discount factor; p ( s ′| s Describe the state s Through action a The probability distribution of triggering state transitions; Q k For the optimal action value function Q *The k The approximation value in the next iteration; α Defined as a policy learning factor, it represents the confidence weight of the experience replay mechanism; Q ( s , a ) is in state s Next action a The Q value.

[0152] This invention integrates evolutionary game theory with Q-learning methods. In evolutionary game theory, the continuous evolution of the system state drives the adaptive updating of the game agent's strategy. The evolutionary game process employs a cumulative influence mechanism of historical action sequences, and the state transition process of each game agent essentially maps to the dynamic optimization trajectory of its policy space.

[0153] Based on the objective function (18), a multi-agent collaborative game architecture is constructed: the game subjects include three types of intelligent agent subsystems (IEP, REP, EVO); the decision space consists of the choices made by each agent regarding the operation mode of the energy supply / storage equipment; the benefit functions correspond to the mathematical models of equations (14), (15), and (17), respectively. In the Q-learning framework of game theory, for the participating intelligent agents... N g,i When the action is selected The optimal response strategy under other agent policy combinations, i.e., satisfying a gi ∈A gi All achieve optimal operational economy, that is:

[0154] (28)

[0155] Then it can be determined that Q-learning converges to the Nash equilibrium point, where ( , This constitutes a Nash equilibrium strategy portfolio. Indicates except for intelligent agents i The equilibrium strategy set of other players.

[0156] In a game-theoretic environment, intelligent agents N g,i The expected total benefit and the updated value can be expressed by the Q-value function as follows:

[0157] (29)

[0158] (30)

[0159] In the formula, Define intelligent agents i The iterative evolution of the action value function, representing The value of the (k+1)th iteration; Characterizing intelligent agents i The time-domain observations of the instant reward mechanism correspond to the function. The k The value of the next iteration; N g The number of agents participating in the game; σ 1 ( s ′),…, σ Ng ( s ′)] represents a Nash equilibrium solution for a game.

[0160] During the strategy iteration process, optimal control decisions need to be made based on the current state space S of the system. The action selection probability calculation mechanism designed in this invention satisfies: the decision set of each agent { a 1 ,…, a Ng} and its associated reward vector { R 1 ,…, R Ng The state space evolution can be obtained through the Q-value function, and it is governed by the joint action policy of multiple agents. Therefore, it is necessary to construct a state transition probability model. p (s′ |a 1 ,…, a Ng Its mathematical representation satisfies:

[0161] (31)

[0162] This invention constructs a state transition probability model based on the Boltzmann distribution to describe the strategy selection process in an evolutionary game framework. This mechanism achieves optimal action selection through agent decision probability mapping y, in the state transition... s Lower intelligent agent i Choose Action a i The probability is:

[0163] (32)

[0164] In the formula, λ For the evolutionary game iteration period k The exponential function of (the number of iterations in a repeated game) is calculated as follows:

[0165] (33)

[0166] λ Parameter quantification of the randomness of agent decision-making. λ As the value increases, the randomness of the agent's decision-making increases; λ As the value decreases, the randomness of the agent's decision-making decreases. It is evident that the fusion optimization of the Boltzmann probability distribution method and the Q-learning algorithm significantly improves the algorithm's policy self-optimization efficiency.

[0167] In one specific implementation, the iterative process of the Q-learning algorithm described above is as follows:

[0168] Step 1) Initialize and configure the Q function parameters. In the offline pre-learning phase, establish a state-action value matrix Q(s,a)=0 with zero initialization; in the online learning phase, a transfer learning mechanism is adopted, and the feasible policy benchmark library generated by pre-learning is loaded as the initial value.

[0169] Step 2) Discretize the continuous state and action variables to form a <state, action> pair-valued function. Simulate and generate samples using the Boltzmann probability distribution method. Based on the Nash equilibrium objective of multi-agent interests, select the current running state and corresponding action strategy, which can be specifically expressed as:

[0170] (34)

[0171] In constructing the state space, this invention uses the following parameters as inputs: actual operating data of photovoltaic output / load demand in the integrated energy system at various time periods, total output power of controllable devices such as combined heat and power systems, heat pump units, and wind turbines, charging and discharging capabilities of energy storage / thermal storage equipment, and electric vehicle discharge power parameters. Given the continuous nature of these parameters, to meet the discretization requirements of the Q-learning framework, the state variables are discretized using an interval partitioning method. The length of each interval is represented as follows:

[0172] (35)

[0173] Based on the device output characteristics, the first m The operating power of the units is divided into categories. M m For discrete intervals, algorithm adaptation for continuous variables is achieved through piecewise quantization, while strictly adhering to equipment output boundary constraints. The system state is composed of multi-dimensional parameters such as photovoltaic power generation, diverse load demands, and energy storage equipment operating conditions, forming a composite state vector with time-series characteristics. S k =< k , S PV , S Load , S CCHP , S GHP , S CAC , S ES , S HS , S EV >, where load status S Load It is a combination of demand characteristics for the three energy types: electricity, heat, and cooling.

[0174] The action decision space encompasses the start-up and shutdown decisions of equipment such as combined heat and power (CHP) units and heat pump systems, as well as the mode selection for the charging and discharging behavior of energy storage units. To address the discretization requirements of the Q-learning algorithm, continuous control variables are transformed into step-by-step decision options, constructing an action vector that includes equipment operating modes and energy storage states. a k =< k , a CCHP , a GHP , a CAC , a ES , a HS , a EV The feasibility of state-action combinations is verified through a constraint screening mechanism. Only valid strategy combinations that meet the equipment operation constraints and game equilibrium conditions are retained as the input parameter set for subsequent Q-value calculation.

[0175] Step 3) Calculate the instantaneous reward value of each agent based on equations (14), (15) and (17); at the same time, predict the subsequent state S′ and record the prediction results in the Q table.

[0176] Step 4) After obtaining the state S′ for the next time period, the Q-value table is updated using the Q-learning iterative algorithm under Nash equalization to achieve iterative state update S←S′. This process needs to simultaneously incorporate the real-time energy parameters of the energy storage device and perform calculations in conjunction with the dynamic model and state-action space parameters.

[0177] Step 5) Determine if the learning process has converged. The learning process terminates when the Nash equilibrium condition is met and the volatility of each agent's Q-function is below a preset threshold, or when the maximum number of iterations is reached. If convergence fails, set k = k + 1 and return to step 2).

[0178] The timing control flow of the above algorithm is as follows: Figure 4 As shown.

[0179] like Figure 5 As shown, based on the same inventive concept, this invention also proposes a comprehensive energy intelligent agent collaborative fault-tolerant group control device to enhance the resilience of energy systems. This device includes:

[0180] The agent segmentation module 110 is used to divide the integrated energy intelligent agent into an agent group control architecture that includes multiple agent subsystems based on the differentiated goals of each agent in the integrated energy intelligent agent.

[0181] The group autonomy module 220 is used to construct a leader-follower model for each intelligent agent subsystem, including a leader, followers, and a collaborative fault-tolerant control module, to achieve group autonomy. In any intelligent agent subsystem, the leader is a control unit, the followers are controlled devices, and the followers are neighboring intelligent agents. The neighboring intelligent agents achieve output consistency under the failure of any follower through the collaborative fault-tolerant control module.

[0182] The inter-group coordination module 230 is used for the group coordination and control architecture of intelligent agents in the group autonomy. With the goal of balancing the interests of each intelligent agent subsystem, it uses the Q-learning algorithm and joint evolutionary game theory to solve the Nash equilibrium point of the game between the multiple intelligent agent subsystems. Based on the solution result, it performs scheduling planning for the integrated energy intelligent agent to achieve inter-group coordination.

[0183] The solutions in this application embodiment can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0184] In addition, embodiments of the present invention also provide an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are connected via a bus. When the computer program is executed by the processor, it implements the various processes of the above-mentioned integrated energy intelligent agent collaborative fault-tolerant group control method for improving the resilience of the energy system, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0185] For details, see Figure 6 As shown, the electronic device includes a bus 1110, a processor 1120, a transceiver 1130, a bus interface 1140, a memory 1150, and a user interface 1160.

[0186] In this embodiment of the invention, the electronic device further includes: a computer program stored in the memory 1150 and executable on the processor 1120, wherein when the computer program is executed by the processor 1120, it implements the above-mentioned integrated energy intelligent agent collaborative fault-tolerant group control method for improving the resilience of the energy system.

[0187] Transceiver 1130 is used to receive and send data under the control of processor 1120.

[0188] In this embodiment of the invention, a bus architecture (represented by bus 1110) is used. Bus 1110 may include any number of interconnected buses and bridges. Bus 1110 connects various circuits, including one or more processors represented by processor 1120 and memory represented by memory 1150.

[0189] Bus 1110 represents one or more of several types of bus architectures, including memory buses and memory controllers, peripheral buses, Accelerated Graphics Port (AGP), processors, or local buses using any bus architecture from various bus architectures. As an example and not a limitation, such architectures include: Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) buses, and Peripheral Component Interconnect (PCI) buses.

[0190] The processor 1120 can be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processors mentioned above include: general-purpose processors, central processing units (CPUs), network processors (NPs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microcontroller units (MCUs) or other programmable logic devices, discrete gates, transistor logic devices, and discrete hardware components. They can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. For example, the processor can be a single-core processor or a multi-core processor, and the processor can be integrated on a single chip or located on multiple different chips.

[0191] Processor 1120 can be a microprocessor or any conventional processor. The method steps disclosed in the embodiments of the present invention can be directly executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in readable storage media known in the art, such as Random Access Memory (RAM), Flash Memory, Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), registers, etc. The readable storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0192] Bus 1110 can also connect various other circuits, such as peripheral devices, voltage regulators, or power management circuits. Bus interface 1140 provides an interface between bus 1110 and transceiver 1130, all of which are well known in the art. Therefore, embodiments of the present invention will not be described further.

[0193] Transceiver 1130 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. For example, transceiver 1130 receives external data from other devices, and transceiver 1130 is used to send data processed by processor 1120 to other devices. Depending on the nature of the computer system, a user interface 1160 may also be provided, such as a touchscreen, physical keyboard, monitor, mouse, speaker, microphone, trackball, joystick, or stylus.

[0194] It should be understood that, in embodiments of the present invention, memory 1150 may further include memory remotely configured relative to processor 1120, and such remotely configured memory may be connected to a server via a network.

[0195] It should be understood that the memory 1150 in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile memory and non-volatile memory.

[0196] In this embodiment of the invention, the memory 1150 stores the following elements of the operating system 1151 and the application 1152: executable modules, data structures, or subsets thereof, or extended sets thereof.

[0197] Specifically, the operating system 1151 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 1152 includes various applications, such as a media player and a browser, used to implement various application functions. Programs implementing the methods of this embodiment of the invention can be included in the application program 1152. The application program 1152 includes applets, objects, components, logic, data structures, and other computer system executable instructions that perform specific tasks or implement specific abstract data types.

[0198] Furthermore, this invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the various embodiments of the comprehensive energy intelligent agent collaborative fault-tolerant group control method for improving the resilience of energy systems, and achieves the same technical effects. To avoid repetition, it will not be described again here.

[0199] This invention also provides a computer program product, including one or more computer instructions. When executed by a processor, the computer instructions implement the steps in the above-described integrated energy intelligent agent collaborative fault-tolerant group control method for improving the resilience of energy systems, and generate all or part of the processes or functions described in the embodiments of this application.

[0200] When the computer instructions are loaded and executed on the processor, all or part of the processes or functions described in the embodiments of this application are generated.

[0201] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0202] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A comprehensive energy intelligent agent collaborative fault-tolerant group control method for enhancing the resilience of energy systems, characterized in that, The method includes: Based on the differentiated objectives of each intelligent agent in the integrated energy intelligent agent, the integrated energy intelligent agent is divided into an intelligent agent group coordination and control architecture that includes multiple intelligent agent subsystems; For each intelligent agent subsystem, a leader-follower model is constructed, including a leader, followers, and a collaborative fault-tolerant control module, to achieve group autonomy. In any intelligent agent subsystem, the leader is the control unit, the followers are the controlled devices, and the followers are neighboring intelligent agents. The neighboring intelligent agents achieve output consistency under the failure of any follower through the collaborative fault-tolerant control module. For a group-based autonomous agent group control architecture, with the goal of balancing the interests of each agent subsystem, a multi-agent cooperative game architecture is constructed based on the following objective function: In the formula, G is the equilibrium point reached in the multi-agent game, and f(·) is the game function of the multi-agent game; N g The number of intelligent agent subsystems participating in the game is 3, including the integrated energy service provider IEP, the renewable energy supplier REP, and the electric vehicle user EVO; A g These are the operational actions of the intelligent agent subsystems participating in the game; i g It is the benefit function of each intelligent agent subsystem; the Q-learning iterative algorithm is used to iterate the multi-agent collaborative game architecture. For the intelligent subsystems participating in the game, when the selected action constitutes the optimal response strategy under the strategy combination of other intelligent agent subsystems, it is determined that the Q-learning converges to the Nash equilibrium point; according to the solution results, the integrated energy intelligent agent is scheduled and planned to achieve inter-group coordination.

2. The method according to claim 1, characterized in that, The integrated energy intelligent agent is divided into an intelligent agent group control architecture comprising multiple intelligent agent subsystems, based on the differentiated objectives of each intelligent agent in the integrated energy intelligent agent, including: The differentiated interests of each agent in the integrated energy intelligent agent are identified, and the integrated energy intelligent agent is divided into four intelligent agent subsystems. The four intelligent agent subsystems include the renewable energy supplier REP and its core equipment, which is responsible for the scheduling and control of distributed power sources and energy storage equipment; the integrated energy service provider IEP and its core equipment, which is responsible for the scheduling and control of energy conversion equipment and thermal storage equipment; the electric vehicle user EVO and its core equipment, which is responsible for the charging and discharging control of electric vehicles; and the terminal load TEU and its core equipment, which is responsible for transmitting the demand information of the system terminal electricity / heat / cooling load and performing load reduction.

3. The method according to claim 2, characterized in that, The core equipment of the renewable energy supplier REP includes at least photovoltaic power generation systems, wind turbine generators, and energy storage devices. The core equipment of the integrated energy service provider IEP includes at least a gas-fired combined cooling, heating and power system, a gas-fired heat pump, an electric chiller, and thermal storage equipment. The core equipment for the electric vehicle user EVO includes at least an electric vehicle cluster, a charging pile network, and an aggregation control platform. The core equipment of the terminal load unit (TEU) includes at least an electrical load terminal, a thermal load terminal, a cold load terminal, and a flexible load management device.

4. The method according to claim 3, characterized in that, The leader of any intelligent agent subsystem is one of the renewable energy supplier REP, integrated energy service provider IEP, electric vehicle user EVO, and terminal load TEU, and the follower of the intelligent agent subsystem is the core device corresponding to the leader.

5. The method according to claim 4, characterized in that, The follower is described by the following continuous-time nonlinear dynamic system: In the formula, i=1,···,N represents the number of agents, x i ∈ n ,u i ∈ m and y i ∈ p Represent the state, control input, and output of the i-th integrated energy device, respectively; g i (t,x i (t)): n n This indicates that both the Lipschitz condition and the initial condition g are satisfied simultaneously. i (0,x i A sufficiently smooth nonlinear function f(0))=0; i (t)∈ r This represents a fault existing in the i-th integrated energy device, and satisfies... i (t) 0; A, B, C and E are real matrices of dimensions n×n, n×m, p×n and n×r, respectively; The leader is described by the following continuous-time nonlinear autonomous dynamic system: In the formula, v(t)∈ h and y r (t)∈ p Represent the state and output of the leading agent, respectively; φ(t,v(t)): h h and q(t,v(t)): h p These are sufficiently smooth nonlinear functions satisfying φ(0,v(0))=0 and q(0,v(0))=0, respectively; φ(t,v(t)) describes the state-space dynamics of the system, representing the next state v(t+1) of the system at time t and state v(t); q(t,v(t)) represents the output y of the system at time t and state v(t). r (t).

6. The method according to claim 5, characterized in that, The neighboring agents achieve output consistency under the fault of any follower through the cooperative fault-tolerant control module, including: In the event of a follower failure, the cooperative fault-tolerant control module outputs u. i (t), such that for any initial condition x i (0) and v(0), have in, .

7. The method according to claim 6, characterized in that, The collaborative fault-tolerant control module includes a fault-tolerant controller, a distributed state and fault estimator, and a distributed external system observer; wherein, the distributed state and fault estimator is used to estimate the state and faults in the followers; the distributed external system observer is used to provide reference tracking signals for some followers who cannot obtain the leader's dynamic information; and the fault-tolerant controller is used to determine the fault-tolerant control input for the followers based on the output of the distributed state and fault estimator.

8. The method according to claim 7, characterized in that, The distributed state and fault estimator is described as follows: In the formula, ∈ n and ∈ r These are the state estimate and the fault estimate, respectively. ∈ n u is the output value of the estimator. i (t) represents the fault-tolerant control input of the follower at time t; the coupling gain κ>0 is an undetermined constant; H∈ n×r ,Γ∈ r×r and F∈ r×p Let Γ be the estimator gain to be determined, where Γ is a positive definite symmetric matrix; ē i (t)∈ p This indicates the measurement output error that depends on local information interaction; The distributed external system observer is described as follows: In the formula, η i (t)∈ p z i (t)∈ h It is the observer's state vector. ∈ This represents the adaptive parameter, where α > 0 indicates a constant; a ij Represents a directed augmenting directed graph Corresponding adjacency matrix Element; The fault-tolerant controller is: In the formula, K1∈ m×n , K2∈ m×n and B * ∈ m×n Let B be the gain matrix to be determined, and let ∥B∥ be the Euclidean norm of the real matrix B.

9. The method according to claim 3, characterized in that, The operation and scheduling benefit function of the integrated energy service provider (IEP) is as follows: In the formula, N T p represents the total number of time periods in the scheduling cycle. B (t) represents the electricity cost incurred by the system in purchasing electricity from the external power grid during time period t; P PG (t) represents the electrical power purchased by the system from the external power grid during time period t; The fuel consumption cost of the IEP intelligent in-vivo control unit during time period t; Characterizes the economic expenditures incurred by the operation and maintenance of IEP intelligent agent devices during time period t; Reflects the additional costs incurred by the IEP intelligent agent in controlling the start-up and shutdown operations of the generating units during time period t; I HS (t) represents the loss cost of the thermal storage equipment during time period t; The scheduling efficiency function of the renewable energy supplier REP is: Where, p S (t) represents the grid connection tariff set by the external power grid operator during time period t; The power value delivered by the REP agent to the external power grid during time period t; The economic expenditure incurred by the operation and maintenance of the REP intelligent agent equipment during period t; I ES (t) represents the cost corresponding to the operating losses of the energy storage device during time period t; Costs required for equipment procurement; The scheduling efficiency function for the electric vehicle user EVO is: Where α represents the probability parameter of EVO participating in scheduling; Δp is the price difference between the grid-connected electric vehicle price set by the grid operator and the benchmark price; P EVd (t) represents the discharge power value of the EVO interacting with the main grid during time period t; r is a random number uniformly distributed in the interval [0,1), and the user scheduling participation mechanism is triggered if and only if r≥α(Δp); [•] represents the rounding function.

10. The method according to claim 1, characterized in that, Also includes: When solving for the Nash equilibrium point of the game among the multiple intelligent agent subsystems, based on power balance constraints, energy supply equipment operation constraints, and energy storage equipment operation constraints, the feasibility of the state-action combinations required for iteration is verified, and the effective strategy combinations that meet the equipment constraints and satisfy the game equilibrium conditions are retained for iteration.

11. A comprehensive energy intelligent agent collaborative fault-tolerant group control device for enhancing the resilience of energy systems, characterized in that, The device includes: The agent segmentation module is used to divide the integrated energy intelligent agent into an agent group control architecture that includes multiple agent subsystems based on the differentiated goals of each agent in the integrated energy intelligent agent. The group autonomy module is used to construct a leader-follower model for each intelligent agent subsystem, including a leader, followers, and a collaborative fault-tolerant control module, to achieve group autonomy. In any intelligent agent subsystem, the leader is the control unit, the followers are the controlled devices, and the followers are neighboring intelligent agents. The neighboring intelligent agents achieve output consistency under the failure of any follower through the collaborative fault-tolerant control module. The inter-group coordination module is used for group coordination and control architecture of autonomous agents within a group. With the goal of balancing the interests of each agent subsystem, it constructs a multi-agent cooperative game architecture based on the following objective function: In the formula, G is the equilibrium point reached in the multi-agent game, and f(·) is the game function of the multi-agent game; N g The number of intelligent agent subsystems participating in the game is 3, including the integrated energy service provider IEP, the renewable energy supplier REP, and the electric vehicle user EVO; A g These are the operational actions of the intelligent agent subsystems participating in the game; i g It is the benefit function of each intelligent agent subsystem; the Q-learning iterative algorithm is used to iterate the multi-agent collaborative game architecture. For the intelligent subsystems participating in the game, when the selected action constitutes the optimal response strategy under the strategy combination of other intelligent agent subsystems, it is determined that the Q-learning converges to the Nash equilibrium point; according to the solution results, the integrated energy intelligent agent is scheduled and planned to achieve inter-group coordination.

12. An electronic device comprising a processor and a memory, the memory storing a computer program, characterized in that, The processor executes a computer program stored in the memory to implement the steps of the method as described in any one of claims 1 to 10.

13. A computer program product comprising computer instructions, characterized in that, When computer instructions are executed by a processor, they implement the steps of the method described in any one of claims 1 to 10.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it performs the steps of the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Regional integrated energy system cluster collaborative optimization method, system, equipment and medium

    CN115907232A

  • Comprehensive energy system microgrid group optimization method, system and related equipment

    CN120542650A