Distributed task allocation method and system oriented to multi-cluster game

By employing an online learning algorithm with a two-layer network topology and an event-triggered mechanism, the dynamic environment and network resource constraints of agents in multi-cluster games are addressed, enabling agents to adaptively adjust and achieve stable convergence in complex distributed systems.

CN120856789APending Publication Date: 2025-10-28JIAXING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510763304.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing distributed Nash equilibrium solutions for multi-cluster games cannot effectively address the competition and cooperation relationships between agents in dynamic environments, with limited network bandwidth, and time-varying network topologies, leading to network congestion and latency.

Method used

By employing a two-layer network topology and event-triggered mechanism, and combining the collaborative updating of policy estimation variables and gradient tracking variables, an online learning algorithm is designed to adapt to environments with time-varying cost functions and limited network resources.

Benefits of technology

It enables agents to adaptively adjust and learn online in dynamic environments, reduces network communication frequency, avoids network congestion, and ensures algorithm convergence and stability, making it suitable for complex distributed systems such as smart grids and energy networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856789A_ABST
    Figure CN120856789A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed task allocation method and system oriented to a multi-cluster game, and relates to the field of distributed computing and control. The method comprises the following steps: firstly, acquiring configuration information of a multi-cluster game system, including an action set, a time-varying cost function set and the like; based on this, a double-layer network topology structure is constructed; and initializing a strategy estimation variable and a gradient tracking variable of each agent. In the algorithm execution process, strategy estimation is updated according to an event triggering mechanism, gradient tracking variables are updated by fusing intra-cluster and cross-cluster gradient information, and strategy estimation variables are updated by adopting an online gradient descent strategy. The task allocation optimization is realized by calculating the dynamic regret of each cluster and verifying whether the dynamic Nash equilibrium condition is met or not. According to the method, the complex interaction relationship of coexistence of competition and cooperation in the multi-cluster game problem in the dynamic environment is solved, the communication overhead is effectively reduced through an event triggering mechanism, and the applicability and efficiency of the system are improved while the convergence is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed computing and control, and in particular to a distributed task allocation method and system for multi-cluster game playing. Background Technology

[0002] In distributed multi-agent systems, game theory provides an important theoretical foundation for solving decision-making problems among agents. Traditional non-cooperative game models perform well in handling purely competitive agent systems, and related distributed Nash equilibrium solving algorithms have been extensively studied. However, in complex practical applications such as smart grids, medical networks, and energy networks, agents often have both competitive and cooperative relationships. To address this, scholars have proposed a multi-cluster game theory framework, dividing agents into different clusters. Within each cluster, agents cooperate, while between clusters, they maintain a competitive relationship, effectively meeting the need for both competition and cooperation in complex systems.

[0003] However, existing methods for solving distributed Nash equilibrium in multi-cluster games suffer from the following technical problems: First, most existing research is based on the static environment assumption, meaning the environment in which the agents operate is time-invariant. This idealized assumption differs significantly from reality; in real-world scenarios, agents are often in dynamic environments, facing time-varying cost functions or constraints. Second, existing algorithms typically assume sufficient network communication resources, allowing agents to exchange information in real time. However, in real-world distributed systems, limited network bandwidth is a common problem; frequent information transmissions not only increase network burden but can also lead to network congestion and latency. Third, the communication network topology between agents in real-world distributed systems is often dynamically changing, and existing algorithms lack effective theoretical guarantees when dealing with time-varying network structures.

[0004] Therefore, there is an urgent need to study distributed online algorithms for multi-cluster games that adapt to dynamic environments, especially for the realistic constraints of limited network bandwidth and time-varying network topology, and to design online learning algorithms with event-triggered mechanisms. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a distributed task allocation method and system for multi-cluster games, which not only solves the multi-cluster game problem involving time-varying cost functions, but also alleviates the problem of limited network resources.

[0006] Firstly, this application proposes a distributed task allocation method for multi-cluster game.

[0007] The distributed task allocation method for multi-cluster game according to embodiments of this application includes:

[0008] Obtain system configuration information of a multi-cluster game system, wherein the multi-cluster game system includes multiple clusters, each cluster includes several agents, and the system configuration information includes the number of agents in each cluster, the action set of each agent, and the time-varying cost function of each agent;

[0009] A two-layer network topology is constructed based on the system configuration information. The two-layer network topology includes a global communication graph and an intra-cluster communication graph corresponding to each cluster.

[0010] Based on the two-layer network topology, initialize the policy estimation variables and gradient tracking variables for each agent;

[0011] Based on the policy estimation variables of the agent at the current moment and the preset event triggering mechanism, update the agent's transmission variables;

[0012] Based on the updated transmission variables and the two-layer network topology, update the policy estimate of each agent for other agents;

[0013] The gradient tracking variable is updated based on the updated policy estimate;

[0014] Based on the updated gradient tracking variables and the preset online gradient descent strategy, update the policy estimation variables;

[0015] Based on the time-varying cost function and the policy estimation variables, the dynamic regret of each cluster is calculated. If the dynamic regret meets the dynamic Nash equilibrium condition, the policy estimation variables are no longer updated.

[0016] The distributed task allocation method for multi-cluster games according to embodiments of this application has at least the following beneficial effects: By constructing a two-layer network topology, the distributed task allocation method for multi-cluster games effectively solves the complex game problem in traditional distributed systems where agents have both competitive and cooperative relationships. This method introduces an event-triggered mechanism, significantly reducing network communication frequency, lowering the system's demand for network bandwidth, and avoiding network congestion caused by frequent information exchange in traditional methods. Through a collaborative update mechanism of policy estimation variables and gradient tracking variables, online learning and adaptive adjustment in time-varying environments are achieved, overcoming the limitation of existing technologies being only applicable to static environments. The designed two-layer network structure supports collaborative cooperation among agents within a cluster and competitive relationships between clusters, better matching the needs of practical application scenarios. The Nash equilibrium determination mechanism based on a dynamic regret index ensures the convergence and stability of the algorithm in dynamic environments, providing a theoretically reliable and practically strong task allocation solution for complex distributed systems such as smart grids and energy networks. The entire method has good distributed characteristics and online learning capabilities, effectively addressing the challenges brought by network topology changes and environmental dynamics.

[0017] According to some embodiments of this application, updating the agent's transmission variables based on the agent's current policy estimation variables and a preset event triggering mechanism includes:

[0018] Calculate the error between the agent's policy estimation variable at the current moment and the transfer variable at the previous moment;

[0019] The event trigger threshold at the current moment is calculated based on the preset time-varying step size;

[0020] The transmission variable is updated based on the error amount and the event trigger threshold.

[0021] According to some embodiments of this application, the gradient tracking variables include intra-cluster fused data and global fused data; updating the gradient tracking variables according to the updated policy estimate includes:

[0022] Based on the policy estimate of the agent at the current moment, calculate the local gradient data corresponding to the agent;

[0023] Based on the intra-cluster communication graph and the local gradient data, the intra-cluster fused data is obtained;

[0024] The global fusion data is obtained based on the global communication graph and the fused data within the cluster.

[0025] According to some embodiments of this application, updating the policy estimation variable based on the updated gradient tracking variable and a preset online gradient descent policy includes:

[0026] Obtain the time-varying step size parameter at the current moment;

[0027] Based on the gradient tracking variables and the time-varying step size parameters, candidate policy data is obtained;

[0028] The candidate policy data is projected onto the action set to obtain the constraint policy data;

[0029] The policy estimation variables are updated based on the constraint policy data.

[0030] According to some embodiments of this application, the step of constructing a two-layer network topology based on the system configuration information includes:

[0031] Based on the physical connection relationships of the agents within each cluster, construct an intra-cluster communication graph;

[0032] Construct a global communication graph based on the communication needs of the agents in different clusters;

[0033] Generate an adjacency matrix that satisfies the double random property for the intra-cluster communication graph, and generate an adjacency matrix that satisfies the double random property for the global communication graph.

[0034] According to some embodiments of this application, the time-varying cost function satisfies the following condition:

[0035] The time-varying cost function of each agent is convex and continuously differentiable with respect to the policy estimation variable at each time step;

[0036] The gradient function of the time-varying cost function satisfies the L-Lipschitz continuity condition;

[0037] The pseudo-gradient mapping of the time-varying cost function satisfies the strong monotonicity condition in the action set.

[0038] According to some embodiments of this application, the step of calculating the dynamic regret of each cluster based on the event cost function and the policy estimation variable, and stopping the updating of the policy estimation variable when the dynamic regret meets the dynamic Nash equilibrium condition, includes:

[0039] Based on the time-varying cost function and the policy estimation variables at the current moment, calculate the actual cost of each cluster at the current moment;

[0040] Based on the time-varying cost function and the optimal policy estimation variable corresponding to the time-varying Nash equilibrium from the initial time to the current time, calculate the minimum possible cost of each cluster at the current time;

[0041] The dynamic regret of each cluster is obtained by calculating the cumulative difference between the actual cost and the corresponding minimum possible cost of each cluster at all times.

[0042] Determine whether the dynamic regret satisfies the sublinear growth condition. If the dynamic regret satisfies the sublinear growth condition, stop updating the policy estimation variable.

[0043] Secondly, this application also proposes a distributed task allocation system for multi-cluster game.

[0044] The distributed task allocation system for multi-cluster games according to the embodiments of this application, applied to the distributed task allocation method for multi-cluster games described in any embodiment of the first aspect, includes:

[0045] System configuration module: used to acquire and manage the configuration information of the multi-cluster game system;

[0046] Two-layer network construction module: used to construct and maintain the two-layer network topology based on the system configuration information;

[0047] Variable initialization module: used to initialize the policy estimation variables and gradient tracking variables for each agent;

[0048] Event-triggered control module: used to control the transmission variables between the intelligent agents according to a preset event-triggered mechanism;

[0049] Policy estimation update module: used to update the policy estimate based on the received transmission variables;

[0050] Gradient tracking module: used to update the gradient tracking variable using gradient tracking technology;

[0051] Online learning module: used to execute an online gradient descent strategy to update the policy estimation variables;

[0052] Dynamic Regret Assessment Module: Used to calculate and assess the dynamic regret of each cluster until Nash equilibrium is reached. Attached Figure Description

[0053] The present application will be further described below with reference to the accompanying drawings and embodiments, wherein:

[0054] Figure 1 Here is a flowchart of the main process of a distributed task allocation method for multi-cluster game in an embodiment;

[0055] Figure 2 This is a flowchart illustrating whether to update the agent's transmission variables in an embodiment.

[0056] Figure 3 This is a flowchart illustrating the updating of inter-agent policy estimation based on transmission variables in an embodiment.

[0057] Figure 4 This is a flowchart illustrating the updating of gradient tracking variables according to the gradient tracking technique in an embodiment.

[0058] Figure 5 This is a flowchart illustrating the updated strategy estimation variables in an example.

[0059] Figure 6 This is a flowchart illustrating how to determine whether a system has reached dynamic Nash equilibrium based on dynamic regret, as part of an embodiment. Detailed Implementation

[0060] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0061] In the description of this application, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0062] In the description of this application, "several" means more than one, "plurality" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0063] In the description of this application, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.

[0064] In the description of this application, reference to the terms "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.

[0065] Firstly, embodiments of this application provide a distributed task allocation method for multi-cluster game playing.

[0066] An example of a distributed task allocation method for multi-cluster game theory is shown below. Figure 1 As shown, including but not limited to steps S100 to S800:

[0067] S100. Obtain the system configuration information of the multi-cluster game system, wherein the multi-cluster game system includes several clusters, each cluster includes several agents, and the system configuration information includes the number of agents in each cluster, the action set of each agent, and the time-varying cost function of each agent.

[0068] S200. Construct a two-layer network topology based on system configuration information. The two-layer network topology includes an intra-cluster communication graph and a global communication graph.

[0069] S300. Based on the two-layer network topology, initialize the policy estimation variables and gradient tracking variables for each agent;

[0070] S400. Update the agent's transmission variables based on the agent's current policy estimation variables and the preset event triggering mechanism;

[0071] S500: Based on the updated transmission variables and the two-layer network topology, update the policy estimate of each agent for other agents;

[0072] S600. Update the gradient tracking variables based on the updated policy estimate;

[0073] S700. Update the policy estimation variables based on the updated gradient tracking variables and the preset online gradient descent policy;

[0074] S800: Based on the time-varying cost function and policy estimation variables, calculate the dynamic regret of each cluster. If the dynamic regret meets the dynamic Nash equilibrium condition, stop updating the policy estimation variables.

[0075] In step S100, for example, Γ([H],Ω,F) is used. tLet represent an online multi-cluster game task allocation problem involving H clusters of N agents, where each cluster h∈[H] contains N agents. h An intelligent agent. Where Ω represents the set of actions of all agents, and the action set Ω := Ω 1 ×…×Ω H ,in Let h be the private action set of agent i in cluster h∈[H].

[0076] Understandably, each cluster is treated as a virtual player participating in a non-cooperative game, but the actual decisions are made by agents within the cluster. The policy estimation variable for agent i in cluster h∈[H] is... The set of policy estimation variables for cluster h is denoted as The set of global policy estimation variables is represented as follows: In cluster h, the set of policy estimation variables excluding agent i is represented as: The set of global policy estimation variables excluding cluster h is denoted as . Agents within the same cluster aim to collaboratively minimize the cluster's time-varying cost function. Action set Ω := Ω 1 ×…×Ω H ,in Let h be the private action set of agent i in cluster h∈[H].

[0077] In a cluster h∈[H], the time-varying cost function of agent i is denoted as f. hi,t (z t This is a private cost function. Assume that the time-varying cost function of each agent is a time-varying convex function. The time-varying cost function of cluster h is denoted as... The global set of time-varying cost functions is F. t :=(F 1,t ,…,F H,t During the training and update process of the multi-cluster game system, the cluster h∈[H] aims to select a time-varying cost function F that satisfies its own limitations. h,t Minimum feasible strategy z h,t (i.e., the set of policy estimation variables for agents in cluster h).

[0078] Understandably, in step S200, a two-layer network topology is constructed based on the system configuration information. A two-layer communication framework and two different time-varying communication graphs are used to connect all agents. The first layer of communication structure is intra-cluster communication between agents within the same cluster, and all agents within the same cluster are connected based on the intra-cluster communication graph. Interactive information is used to collaboratively minimize the time-varying cost function of the cluster. The second-layer communication structure involves all agents communicating via a global communication graph. Perform global communication.

[0079] Understandably, in some embodiments, for any time t, the cluster intra-cluster communication graph connecting all agents in the cluster... It is strongly connected. The adjacency matrix is ​​denoted as In other embodiments, for any time t, All diagonal elements are positive; in other embodiments, for any time t, the adjacency matrix of all clusters is... It is non-negative and has the property of being doubly random, that is...

[0080] Understandably, in some embodiments, for any time t, a global communication graph connecting all agents... It is strongly connected. The adjacency matrix is ​​denoted as R. t In other embodiments, for any time t, R t The diagonal terms are positive; in other embodiments, for any time t, the adjacency matrix R... t It is non-negative and has the property of being doubly random, that is...

[0081] In step S300, based on the two-layer network topology, the policy estimation variables and gradient tracking variables for each agent are initialized. For any agent i ∈ [N] in the cluster h ∈ [H]... h ],initialization Choose any Gradient tracking variable in, This represents the estimate of the gradient of agent i on cluster h.

[0082] In step S400, the agent's transmission variables are updated based on the agent's current policy estimation variables and a preset event triggering mechanism. In this embodiment, the preset event triggering mechanism is used to determine whether the error quantity meets the conditions for data transmission. If the conditions are met, the current data (i.e., the policy estimation variables) is transmitted; otherwise, no transmission is performed. It is understood that... Figure 2 As shown, step S400 further includes, but is not limited to, steps S410 to S430:

[0083] S410. Calculate the error between the agent's policy estimate variable at the current moment and the transfer variable at the previous moment;

[0084] S420. Calculate the event trigger threshold at the current moment based on the preset time-varying step size;

[0085] S430. Update the transmission variable at the current moment based on the error amount and the time trigger threshold.

[0086] In steps S410 to S430, in order to relieve network pressure and reduce network transmission volume, an event-triggering mechanism is introduced into the algorithm. Taking the current time as the t-th moment and the previous time as the (t-1)-th moment as an example, for any agent Define as the transmission variable updated by agent j for agent i at the (t-1)-th moment, and z (hi)hj,t as the actual policy estimation variable of agent j for agent i at the t-th moment. Calculate as the error amount between the policy estimation variable at the current moment and the transmission variable at the previous moment. This error amount reflects the deviation degree between the actual policy and the last transmitted policy. The designed event-triggering function is as follows: where, C0α t is the event-triggering threshold, where C0>0 is a preset basic threshold constant, and α t is the time-varying step size at the t-th moment, 0 < α t < 1. For example, the specific form of α t is where 0 < a < 1, and at the same time α t satisfies the following conditions: the infinite series sum of the time-varying step size parameter sequence is infinite; the sum of the squares of the sequence of the time-varying step size parameters is finite. When it means that the error amount exceeds the event-triggering threshold, update the transmission variable at the t-th moment and send the updated transmission variable to agent i; otherwise, do not transmit data to agent i, and make agent i keep the previous transmission variable unchanged. This selective transmission mechanism effectively reduces the network communication load.

[0087] The design of the event-triggering mechanism fully considers the convergence requirements of the algorithm. Through theoretical analysis, it can be proved that under the event-triggering condition, the algorithm can still ensure that the policy estimation variable converges to the true Nash equilibrium point. Specifically, the decreasing property of the event-triggering threshold ensures that the cumulative transmission error is bounded and satisfies conditions, thus not affecting the overall convergence of the algorithm. By implementing the event-triggering mechanism, the communication overhead of the distributed system is significantly reduced. On the premise of ensuring the convergence of the algorithm, the communication times can be effectively reduced; it has strong adaptability. The time-varying characteristic of the event-triggering threshold enables the mechanism to automatically adjust the communication frequency according to the progress of the algorithm; it is simple to implement. Only need to compare the error amount with the event-triggering threshold, and the computational complexity is low.

[0088] It can be understood that in step S500, according to the transmission variable and the two-layer network topology structure, the policy estimation of each agent for other agents is updated, as Figure 3 shown, further including but not limited to:

[0089] S510. Update the first policy estimate of each agent based on the communication graph within each cluster and the corresponding transmission variables of each agent within the cluster.

[0090] S520. For each cluster, update the second policy estimate for each agent based on the global communication graph and the transmission variables of agents in other clusters.

[0091] In steps S510 to S520, two different time-varying communication graphs in a two-layer network topology are used to connect all agents. The first layer of the two-layer network topology is an intra-cluster communication graph between agents in the same cluster, and all agents in the same cluster are connected based on this intra-cluster communication graph. Interactive information is used to collaboratively minimize the cluster's cost function. The second layer of the two-layer network topology's communication structure involves all agents communicating via a global communication graph. Perform global communication.

[0092] The first policy estimate represents the policy estimate of the corresponding agent towards the other agents in its cluster; the second policy estimate represents the policy estimate of the corresponding agent towards the agents in other clusters in a multi-cluster game system.

[0093] For example, in S510, for any agent i∈[N] in the cluster... h ], h∈[H], which is relative to other agents j∈[N] in the same cluster h The first policy estimate for h∈[H] is updated as follows: in, Communication diagram within the cluster The corresponding elements in the adjacency matrix represent the weights of agent i and neighbor j. Let be the transmission variable of agent j's estimate of agent i at time t. This update process is essentially a weighted average of the transmission variables among agents in the cluster based on their adjacency matrix weights, thereby updating agent i ∈ [N]. h The first policy estimate for other agents in the cluster is given by [H], h ∈ [N]. In step S520, for any agent i ∈ [N] in the cluster... h ], h∈[H], which is relevant to other clusters The agent's second policy estimate is updated as follows: in, Global communication graph The corresponding element in the adjacency matrix corresponds to cluster h and its neighboring clusters. The weight, Neighbor cluster The transmission variables for agents within cluster h. This update process allows each agent to obtain policy information from other agents in the cluster, thereby updating the global first policy estimate. Combining the first and second policy estimates, the updated complete policy estimate for agent i is z. hi,t+1 =[z (h)hi,t+1 ,z (-h)hi,t+1 ] T It can be seen that at each time t, the system synchronously performs dual policy estimation updates according to the following steps: all agents simultaneously perform the first policy estimation update, based on the transmission variables of neighbors within the cluster; all agents simultaneously perform the second policy estimation update, based on the transmission variables of the global neighbor cluster. This dual update mechanism ensures that information propagates rapidly within the cluster while also effectively achieving cross-cluster information exchange, thus enabling the entire multi-cluster game system to converge to a global dynamic Nash equilibrium. Through this hierarchical policy estimation update method, the system maintains both close coordination within the cluster and effective cooperation between clusters. It is worth noting that the adjacency matrices of both the intra-cluster communication graph and the global communication graph have a doubly random property, i.e. and This property ensures convergence during the information fusion process.

[0094] Understandably, in step S600, the gradient tracking variables are updated based on the updated policy estimate. First, local gradient data is calculated, then gradient information from agents within the cluster is fused based on the cluster communication graph to obtain cluster-fused data. Next, gradient information across the cluster is fused based on the global communication graph to obtain global fused data, and finally, the gradient tracking variables are updated. Specifically, as follows... Figure 4 As shown, step S600 may further include, but is not limited to, steps S610 to S630:

[0095] S610. Calculate the local gradient data corresponding to the agent based on the policy estimate of the agent at the current moment;

[0096] S620. Based on the intra-cluster communication graph and local gradient data, intra-cluster fused data is obtained;

[0097] S630: Based on the global communication graph and the fused data within the cluster, global fused data is obtained;

[0098] In steps S610 to S630, the local gradient data is first calculated based on the agent's policy estimate at the current moment. For any i∈[N] h The local gradient data of agent i are: h∈[H], where h∈[H]. in, f represents the local gradient data of agent i at time t. hi,t (z hi,tLet z be the time-varying cost function of agent i at time t. hi,t This represents the policy estimate for agent i at time t. The local gradient data reflects the local optimization direction of agent i under the current policy. Due to the convexity and continuous differentiability of the cost function, this gradient calculation is uniquely determined. Subsequently, based on the intra-cluster communication graph... By fusing local gradient data with gradient information from agents within the cluster, we obtain fused data within the cluster. Among them, y hi,t+1 For agent i, the data is fused within the cluster at time t+1. The elements of the adjacency matrix of the communication graph within the cluster represent the weights of agent i and neighbor j, y. hj,t To fuse data within the cluster for neighboring intelligent agent j at time t, This represents the local gradient data of agent i at time t-1. Data fusion within the cluster enables the fusion and propagation of gradient information among agents within the same cluster. This fusion process ensures that all agents within the cluster can collaboratively optimize the overall cost function of the cluster.

[0099] Subsequently, based on the global communication graph By fusing data within the cluster and further fusing gradient information across the cluster, we obtain globally fused data. Among them, Y h,t+1 For the globally fused data of cluster h at time t+1, R hg Y represents the elements of the adjacency matrix of the global communication graph, corresponding to the weights of cluster h and its neighboring cluster g. g,t The global fusion data of neighbor cluster g at time t. This is a normalization factor for the number of agents in cluster h. For any agent i ∈ [N] h ], h∈[H], its complete gradient tracking variables consist of cluster-fused data and global fused data: v hi,t+1 =y hi,t+1 +Y h,t+1 The gradient tracking variable update method in this embodiment employs a two-layer gradient tracking mechanism to ensure that each agent can obtain tracking variables containing gradient information of the entire system, providing an accurate optimization direction for subsequent online gradient descent. Through this hierarchical gradient information fusion approach, the system maintains rapid propagation of gradient information within the cluster while achieving effective sharing of gradient information across clusters, thereby supporting the convergence of the entire multi-cluster game system to dynamic Nash equilibrium.

[0100] Understandably, in step S700, the policy estimation variable is updated based on the updated gradient tracking variable and the preset online gradient descent policy. For example... Figure 5 As shown, step S700 further includes, but is not limited to, steps S710 to S740:

[0101] S710, Obtain the time-varying step size parameter at the current moment;

[0102] S720. Based on the gradient tracking variables and time-varying step size parameters, candidate policy data is obtained;

[0103] S730. Project the candidate policy data onto the action set to obtain the constraint policy data;

[0104] S740. Update the strategy estimate variables based on the constraint data.

[0105] In step S710, for each time t, the system obtains the current time-varying step size parameter: α t Let be the time-varying step size parameter at time t, where α t Always greater than 0, this ensures that the system's learning step size is always positive, guaranteeing the algorithm's progressiveness. α t It is always less than 1, which prevents instability caused by excessively large step sizes. Secondly, α t It also needs to satisfy the following condition: the infinite series sum of the time-varying step size parameter sequence must be infinite, that is... To ensure the algorithm has sufficient learning ability to overcome arbitrarily large initial errors and converge to the optimal solution under finite perturbations, the sum of squares of the time-varying step size parameter sequence must be a finite value. This is used to control the decay rate of the time-varying step size sequence and prevent oscillations during convergence. For example, the time-varying step size parameter can take the following specific form: Among them, parameter 0

[0106] In step S720, candidate policy data is obtained based on the gradient tracking variables and the time-varying step size parameter. For any agent i ∈ [N] in the cluster h ∈ [H]... h Based on the current gradient tracking variables, candidate policy data is calculated: in, Let z be the candidate policy variable for agent i at time t+1, representing the direction of gradient descent. hi,t Let be the policy estimation variable for agent i at time t. Let v be the time-varying step size parameter at time t. hi,t+1 Let be the gradient tracking variable for agent i at time t+1. This candidate policy data represents the update result along the negative gradient direction in one step and is the core computational step of the gradient descent algorithm.

[0107] ​In step S730, the candidate policy data is projected onto the action set to obtain constrained policy data. The constraint domain of the candidate policy data projected onto the agent's action set is as follows: Where, x hi,t+1 The constraint policy data for agent i at time t+1 is the result of projecting the policy estimate variables onto the agent's action set. Let i be the set of actions of agent i. To the action set The projection operation. For example, the specific implementation of the projection operation is as follows: in, and These are the lower and upper bounds of the action set for agent i, respectively.

[0108] In step S740, the policy estimation variables are updated using the constrained policy data: Among them, z hi,t+1 Let be the policy estimate variable for agent i at time t+1. As weighting coefficients for the estimated variables that maintain the current strategy, As the weighting coefficients for the data using the new constraint policy, it can be seen that the new policy estimate variable is essentially a weighted average of the policy estimate variable at the previous time step and the new constraint policy data.

[0109] By employing this gradient-following-based online gradient descent strategy, multi-cluster game systems can continuously optimize the strategies of each agent in a time-varying environment, ultimately converging to a dynamic Nash equilibrium. This method combines the efficiency of distributed optimization with the adaptability of online learning, providing an effective solution approach for complex multi-cluster game problems.

[0110] Understandable, such as Figure 6 As shown, step S800 further includes, but is not limited to, steps S810 to S840:

[0111] S810. Calculate the actual cost of each cluster at the current time based on the time-varying cost function and the policy estimation variables at the current time.

[0112] S820. Based on the time-varying cost function and the optimal policy estimation variable corresponding to the time-varying Nash equilibrium from the initial time to the current time, calculate the minimum possible cost of each cluster at the current time.

[0113] S830. Calculate the cumulative difference between the actual cost and the corresponding minimum possible cost of each cluster at all times to obtain the dynamic regret of each cluster.

[0114] S840. Determine whether the dynamic regret satisfies the sublinear growth condition. If the dynamic regret satisfies the sublinear growth condition, stop updating the policy estimation variables.

[0115] In steps S810 to S840, the calculation of dynamic regret reflects the real-time performance of the algorithm. For each cluster h∈[H], its dynamic regret metric is: Where z h,t These are the policy estimation variables of the agent at the current time t. is the optimal policy estimate variable corresponding to the time-varying Nash equilibrium point from the initial time to time t, at which cluster h cannot reduce its cost by unilaterally changing its own policy. Dynamic regret measures the cumulative deviation between the actual cost and the ideal optimal cost of cluster h throughout the learning process. The smaller the dynamic regret, the higher the learning efficiency of the cluster and the closer its policy selection is to the optimal one. When the ratio of dynamic regret to learning time for all clusters approaches zero, it is confirmed that the dynamic regret satisfies the sublinear growth condition, and the multi-cluster game system is determined to have reached a dynamic Nash equilibrium.

[0116] Understandably, to avoid drastic fluctuations in the problem, in some embodiments, a path accumulation can also be calculated to assess the dynamic nature of the problem, wherein the strategy path accumulation is: Gradient path accumulation: These two indicators reflect the degree to which the Nash equilibrium changes over time.

[0117] The evaluation of the sublinear growth condition is crucial for determining the algorithm's performance. and This indicates that the dynamic changes in the problem are mild. Under this condition, if the dynamic regret of all clusters satisfies Reg h If (T) = o(T), then the algorithm has successfully tracked the time-varying Nash equilibrium. The convergence determination process includes: monitoring the average regret Reg h The trend of (T) / T; when this value is less than the preset threshold and remains stable, the system is considered to have reached dynamic Nash equilibrium.

[0118] Secondly, this application proposes a distributed task allocation system for multi-cluster games, applied to the distributed task allocation method for multi-cluster games in the first aspect embodiment, comprising:

[0119] System configuration module: used to acquire and manage configuration information for multi-cluster gaming systems;

[0120] Two-layer network construction module: used to build and maintain a two-layer network topology based on system configuration information;

[0121] Variable initialization module: Used to initialize the policy estimation variables and gradient tracking variables for each agent;

[0122] Event-triggered control module: Used to control the transmission variables between intelligent agents according to a preset event-triggered mechanism;

[0123] Policy estimation update module: used to update policy estimates based on received transmission variables;

[0124] Gradient tracking module: Used to update gradient tracking variables using gradient tracking techniques;

[0125] Online learning module: used to execute online gradient descent strategies to update policy estimation variables;

[0126] Dynamic Regret Assessment Module: Used to calculate and assess the dynamic regret of each cluster until Nash equilibrium is reached.

[0127] The distributed task allocation system for multi-cluster games in this embodiment is based on the distributed task allocation method for multi-cluster games described above. Therefore, the distributed task allocation system for multi-cluster games in this embodiment has the same beneficial effects as the distributed task allocation method for multi-cluster games described above. To save space, it will not be described again here.

[0128] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application. Furthermore, unless otherwise specified, the embodiments and features described in the embodiments of this application can be combined with each other.

Claims

1. A distributed task allocation method for multi-cluster game theory, characterized in that, include: Obtain system configuration information of a multi-cluster game system, wherein the multi-cluster game system includes multiple clusters, each cluster includes several agents, and the system configuration information includes the number of agents in each cluster, the action set of each agent, and the time-varying cost function of each agent; A two-layer network topology is constructed based on the system configuration information. The two-layer network topology includes a global communication graph and an intra-cluster communication graph corresponding to each cluster. Based on the two-layer network topology, initialize the policy estimation variables and gradient tracking variables for each agent; Based on the policy estimation variables of the agent at the current moment and the preset event triggering mechanism, update the agent's transmission variables; Based on the updated transmission variables and the two-layer network topology, update the policy estimate of each agent for other agents; The gradient tracking variable is updated based on the updated policy estimate; Based on the updated gradient tracking variables and the preset online gradient descent strategy, update the policy estimation variables; Based on the time-varying cost function and the policy estimation variables, the dynamic regret of each cluster is calculated. If the dynamic regret meets the dynamic Nash equilibrium condition, the policy estimation variables are no longer updated.

2. The distributed task allocation method for multi-cluster game as described in claim 1, characterized in that, The step of updating the agent's transmission variables based on the policy estimation variables at the current moment and a preset event triggering mechanism includes: Calculate the error between the agent's policy estimation variable at the current moment and the transfer variable at the previous moment; The event trigger threshold at the current moment is calculated based on the preset time-varying step size; The transmission variable is updated based on the error amount and the event trigger threshold.

3. The distributed task allocation method for multi-cluster game as described in claim 1, characterized in that, The gradient tracking variables include cluster-in-cluster fused data and global fused data; The step of updating the gradient tracking variable based on the updated policy estimate includes: Based on the agent's current policy estimate, calculate the local gradient data corresponding to the agent; Based on the intra-cluster communication graph and the local gradient data, the intra-cluster fused data is obtained; The global fusion data is obtained based on the global communication graph and the fused data within the cluster.

4. The distributed task allocation method for multi-cluster game as described in claim 1, characterized in that, The step of updating the policy estimation variable based on the updated gradient tracking variable and the preset online gradient descent policy includes: Obtain the time-varying step size parameter at the current moment; Based on the gradient tracking variables and the time-varying step size parameters, candidate policy data is obtained; The candidate policy data is projected onto the action set to obtain the constraint policy data; The policy estimation variables are updated based on the constraint policy data.

5. The distributed task allocation method for multi-cluster game as described in claim 1, characterized in that, The step of constructing a two-layer network topology based on the system configuration information includes: Based on the physical connection relationships of the agents within each cluster, construct an intra-cluster communication graph; Construct a global communication graph based on the communication needs of the agents in different clusters; Generate an adjacency matrix that satisfies the double random property for the intra-cluster communication graph, and generate an adjacency matrix that satisfies the double random property for the global communication graph.

6. The distributed task allocation method for multi-cluster game as described in claim 1, characterized in that, The time-varying cost function satisfies the following condition: The time-varying cost function of each agent is convex and continuously differentiable with respect to the policy estimation variable at each time step; The gradient function of the time-varying cost function satisfies the L-Lipschitz continuity condition; The pseudo-gradient mapping of the time-varying cost function satisfies the strong monotonicity condition in the action set.

7. The distributed task allocation method for multi-cluster game as described in claim 1, characterized in that, The step of calculating the dynamic regret of each cluster based on the event cost function and the policy estimation variables, and stopping the updating of the policy estimation variables when the dynamic regret meets the dynamic Nash equilibrium condition, includes: Based on the time-varying cost function and the policy estimation variables at the current moment, calculate the actual cost of each cluster at the current moment; Based on the time-varying cost function and the optimal policy estimation variable corresponding to the time-varying Nash equilibrium from the initial time to the current time, calculate the minimum possible cost of each cluster at the current time; The dynamic regret of each cluster is obtained by calculating the cumulative difference between the actual cost and the corresponding minimum possible cost of each cluster at all times. Determine whether the dynamic regret satisfies the sublinear growth condition. If the dynamic regret satisfies the sublinear growth condition, stop updating the policy estimation variable.

8. A distributed task allocation system for multi-cluster game playing, characterized in that... The distributed task allocation method for multi-cluster game as described in any one of claims 1-7 is characterized by comprising: System configuration module: used to acquire and manage the configuration information of the multi-cluster game system; Two-layer network construction module: used to construct and maintain the two-layer network topology based on the system configuration information; Variable initialization module: used to initialize the policy estimation variables and gradient tracking variables for each agent; Event-triggered control module: used to control the transmission variables between the intelligent agents according to a preset event-triggered mechanism; Policy estimation update module: used to update the policy estimate based on the received transmission variables; Gradient tracking module: used to update the gradient tracking variable using gradient tracking technology; Online learning module: used to execute an online gradient descent strategy to update the policy estimation variables; Dynamic Regret Assessment Module: Used to calculate and assess the dynamic regret of each cluster until Nash equilibrium is reached.