Heterogeneous agent distributed multi-alliance game control method, device and equipment and medium

By dividing a heterogeneous unmanned swarm system into multiple alliances and establishing a dynamic model and communication topology, and using a distributed Nash equilibrium search algorithm to optimize the decision state, the problem of rapid decision optimization in multi-alliance games of high-order heterogeneous intelligent agents is solved, achieving precise control and improved collaborative performance.

CN121284029APending Publication Date: 2026-01-06BEIHANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511362114.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing multi-agent system algorithms have poor convergence performance and limited applicability when dealing with high-order heterogeneous agent multi-alliance games, making it difficult to achieve efficient cooperative control with coexistence of cooperation and competition in complex scenarios.

Method used

The agents in the heterogeneous unmanned swarm system are divided into multiple alliances. A high-order heterogeneous linear state dynamics model and an inner and outer two-layer alliance communication topology are established. The distributed Nash equilibrium search algorithm is used to iteratively optimize the decision state through the search layer and the output adjustment layer until it converges to the Nash equilibrium state, thereby generating a multi-alliance game strategy.

Benefits of technology

It achieves precise control of heterogeneous unmanned swarm systems, improves task execution capabilities and collaborative performance, and solves the problem of rapid decision optimization in multi-alliance games of high-order heterogeneous intelligent agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121284029A_ABST
    Figure CN121284029A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous agent distributed multi-alliance game control method, device and equipment and a medium, and relates to the field of multi-agent system cooperative control and game theory cross application, and the method comprises the steps: dividing agents in a heterogeneous unmanned cluster system into a plurality of alliances; establishing a high-order heterogeneous linear state dynamical model for the intelligent agent in each alliance; constructing an internal and external double-layer alliance communication topological graph according to an actual communication condition; and on the basis of the model and the topological graph, through a distributed Nash equilibrium search algorithm containing a search layer and an output adjustment layer, performing iterative optimization on an agent decision state to obtain a multi-alliance game strategy, and controlling the agent according to the strategy. According to the method, the cooperative control problem of the heterogeneous agents in a multi-alliance game scene can be effectively solved, the Nash equilibrium point is quickly searched, the decision state is adjusted to be optimal, and the cooperative control performance and task execution efficiency of a multi-agent system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of collaborative control of multi-agent systems and the cross-application of game theory, and in particular to a method, device, equipment and medium for distributed multi-alliance game control of heterogeneous agents. Background Technology

[0002] With the rapid development of communication networking technology and distributed algorithms, the research on multi-agent systems has attracted much attention and has been widely applied.

[0003] However, many current multi-agent system algorithms focus only on single cooperative or competitive relationships. In cooperative relationships, consensus is often reached based on consensus algorithms or group formation is achieved, while in game theory, the focus is mainly on searching for Nash equilibrium solutions. But in real-world scenarios, cooperative and competitive relationships often coexist, which can be modeled as a multi-alliance game problem. Regarding multi-alliance games, related technologies disclose a distributed multi-alliance game algorithm for first-order dynamical players in the case of topology switching, and related technologies disclose a pre-set time Nash equilibrium search algorithm for multi-alliance games of first-order or second-order dynamical agents under event triggering. However, these algorithms have limitations, such as only being able to achieve asymptotic or exponential convergence, being applicable to simple clusters, and being limited to isomorphic and low-order dynamical agents.

[0004] Therefore, there is an urgent need for a pre-set time distributed Nash equilibrium search algorithm for high-order heterogeneous agents in multi-alliance games to solve problems such as poor convergence and limited applicability when dealing with high-order heterogeneous agents in multi-alliance games, thereby improving the game performance and task execution efficiency of multi-agent systems in complex scenarios. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, equipment, and medium for controlling distributed multi-alliance games of heterogeneous intelligent agents, which can effectively solve the multi-alliance game problem composed of high-order heterogeneous intelligent agents in complex scenarios where cooperation and competition coexist.

[0006] To achieve the above objectives, this application provides the following solution:

[0007] Firstly, this application provides a method for controlling heterogeneous intelligent agents in a distributed multi-alliance game, including:

[0008] The agents in the heterogeneous unmanned swarm system are divided into multiple alliances; each alliance includes multiple agents.

[0009] Establish a high-order heterogeneous linear state dynamics model for each agent in each alliance;

[0010] Based on the actual communication between all agents in the heterogeneous unmanned swarm system, an internal and external two-layer alliance communication topology diagram of the heterogeneous unmanned swarm system is established.

[0011] Based on a high-order heterogeneous linear state dynamics model and an internal and external two-layer alliance communication topology, a distributed Nash equilibrium search algorithm is used to iteratively optimize the decision states of all agents in a heterogeneous unmanned swarm system until the decision states of all agents converge to a Nash equilibrium state within a preset convergence time, thus obtaining a multi-alliance game strategy. The distributed Nash equilibrium search algorithm includes a search layer and an output adjustment layer. The search layer is used to iteratively adjust the decision states of agents based on a leader-follower consistency observer to search for a Nash equilibrium point. The output adjustment layer is used to generate the control input corresponding to each agent based on the Nash equilibrium point searched by the search layer, thus obtaining a multi-alliance game strategy.

[0012] Controlling agents in a heterogeneous unmanned swarm system based on multi-alliance game strategy.

[0013] Secondly, this application provides a heterogeneous intelligent agent distributed multi-alliance game control device, comprising:

[0014] The alliance partitioning module is used to divide the agents in the heterogeneous unmanned swarm system into multiple alliances; each alliance includes multiple agents.

[0015] The dynamics model building module is used to build a high-order heterogeneous linear state dynamics model for each agent in each alliance.

[0016] The communication topology establishment module is used to establish an internal and external two-layer alliance communication topology diagram of the heterogeneous unmanned swarm system based on the actual communication situation between all intelligent agents in the heterogeneous unmanned swarm system.

[0017] The strategy generation module is used to iteratively optimize the decision states of all agents in a heterogeneous unmanned swarm system based on a high-order heterogeneous linear state dynamics model and an internal and external two-layer alliance communication topology. This optimization is achieved through a distributed Nash equilibrium search algorithm until the decision states of all agents converge to a Nash equilibrium state within a preset convergence time, thus obtaining a multi-alliance game strategy. The distributed Nash equilibrium search algorithm includes a search layer and an output adjustment layer. The search layer iteratively adjusts the agent decision states based on a leader-follower consistency observer to search for a Nash equilibrium point. The output adjustment layer generates control inputs for each agent based on the Nash equilibrium point searched by the search layer, thus obtaining the multi-alliance game strategy.

[0018] The control module is used to control the agents in a heterogeneous unmanned swarm system according to a multi-alliance game strategy.

[0019] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the heterogeneous intelligent agent distributed multi-alliance game control method described in any one of the above.

[0020] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the heterogeneous intelligent agent distributed multi-alliance game control method described above.

[0021] According to the specific embodiments provided in this application, this application has the following technical effects:

[0022] This application provides a method, apparatus, device, and medium for controlling heterogeneous intelligent agents in a distributed multi-alliance game. By dividing the agents in a heterogeneous unmanned swarm system into multiple alliances and establishing a high-order heterogeneous linear state dynamics model for each agent within an alliance, it solves the problem of how to reasonably model and group complex heterogeneous intelligent agent systems, achieving a structured approach to heterogeneous unmanned swarm systems. By establishing an internal and external two-layer alliance communication topology based on actual communication conditions, it solves the problem of unclear communication relationships between agents, achieving a clear description of the agent communication structure. Based on the model and topology, it uses a distributed Nash equilibrium search algorithm with a search layer and an output adjustment layer to iteratively optimize the decision state. The search layer searches for the Nash equilibrium point using a leader-follower consistency observer, and the output adjustment layer generates control inputs. This solves the problem of decision optimization and control of heterogeneous intelligent agents in multi-alliance games, achieving rapid search for Nash equilibrium and generation of effective control strategies for precise control of the agents. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating a heterogeneous intelligent agent distributed multi-alliance game control method provided in an embodiment of this application;

[0025] Figure 2 A communication topology diagram provided in one embodiment of this application;

[0026] Figure 3 A group formation trajectory diagram provided in an embodiment of this application based on a heterogeneous intelligent agent distributed multi-alliance game control method;

[0027] Figure 4 This application provides an embodiment of a heterogeneous agent-based distributed multi-alliance game control method for agent output decision graphs.

[0028] Figure 5 A schematic diagram of the functional modules of a heterogeneous intelligent agent distributed multi-alliance game control device provided in an embodiment of this application;

[0029] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0030] First, some technical terms involved in the embodiments of this application will be introduced.

[0031] With the development and application of communication networking technology and distributed algorithms, research on multi-agent systems has received widespread attention and application. However, many algorithms only focus on a single cooperative or competitive relationship among multiple agents. In cooperative relationships, the focus is usually on reaching consensus or forming a formation based on consensus algorithms and interactions; in game theory, the focus is mostly on searching for Nash equilibrium solutions.

[0032] Nash equilibrium is an important concept in game theory. It describes an optimal set of strategies in a non-cooperative game, in which no agent can gain a greater benefit by unilaterally changing its decisions.

[0033] However, in many real-world scenarios, cooperation and competition coexist. Intelligent agents, driven by their own interests, form game alliances. Within an alliance, agents cooperate to maximize the alliance's benefits—this is a cooperative relationship. Between different alliances, however, they compete against each other. This cooperative-competitive coexistence model can be modeled as a multi-alliance game problem. In this type of problem, the alliance is represented by a virtual entity directly participating in the game, but its decision-making is still actually made by the individual agents that make up the alliance.

[0034] Regarding multi-alliance game theory, researchers have pointed out that the core issue lies in the fact that agents can only obtain their own cost functions. Optimizing the cost function of the entire alliance, i.e., optimizing the sum of the cost functions of all agents in the alliance, requires estimating the gradient of the alliance cost function by introducing dynamic average consistency. Subsequently, scholars studied distributed multi-alliance game algorithms for first-order dynamical agents in topology-switching scenarios. However, these algorithms only achieved asymptotic or exponential convergence and were designed for simple clusters of first-order dynamical agents. To address the need for faster convergence, some researchers have studied pre-defined timed Nash equilibrium search algorithms for multi-alliance games using first-order or second-order dynamical agents triggered by events. These algorithms allow for pre-specifying the convergence time; however, the agent clusters studied are homogeneous and limited to first- or second-order dynamics. In practice, heterogeneous clusters are more common and capable of handling more complex tasks.

[0035] In summary, this application proposes a pre-set time distributed Nash equilibrium search algorithm for multi-alliance games involving high-order heterogeneous agents. The algorithm consists of a search layer and an output adjustment layer. In the search layer, an observer design based on leader-follower consistency allows agents to search for Nash equilibrium points in multi-alliance games simply by interacting with their neighbors, and a time-varying gain function is introduced to achieve convergence within a pre-set time. Subsequently, in the output adjustment layer, based on the results of the preceding search layer, a time-varying Lyapunov equation is introduced to adapt to the high-order heterogeneous dynamic characteristics of the agents while adjusting the decision states of the participating heterogeneous multi-agent teams to the optimal Nash equilibrium solution within a pre-set time.

[0036] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] In one exemplary embodiment, such as Figure 1 As shown, a method for controlling a heterogeneous intelligent agent distributed multi-alliance game is provided, including the following steps 201 to 205. Wherein:

[0039] Step 201: Divide the agents in the heterogeneous unmanned swarm system into multiple alliances; each alliance includes multiple agents.

[0040] Step 202: Establish a high-order heterogeneous linear state dynamics model for each agent in each alliance.

[0041] Step 203: Based on the actual communication situation among all agents in the heterogeneous unmanned swarm system, establish an internal and external two-layer alliance communication topology diagram of the heterogeneous unmanned swarm system.

[0042] Step 204: Based on the high-order heterogeneous linear state dynamics model and the internal and external two-layer alliance communication topology, the decision states of all agents in the heterogeneous unmanned swarm system are iteratively optimized using a distributed Nash equilibrium search algorithm until the decision states of all agents converge to the Nash equilibrium state within a preset convergence time, thus obtaining a multi-alliance game strategy. The distributed Nash equilibrium search algorithm includes a search layer and an output adjustment layer. The search layer is used to iteratively adjust the decision states of agents based on a leader-follower consistency observer to search for a Nash equilibrium point. The output adjustment layer is used to generate the control input corresponding to each agent based on the Nash equilibrium point searched by the search layer, thus obtaining a multi-alliance game strategy.

[0043] Step 205: Control the agents in the heterogeneous unmanned swarm system according to the multi-alliance game strategy.

[0044] By implementing steps 201 to 205 above, this application can effectively address the collaborative control problem of heterogeneous unmanned swarm systems in multi-alliance game scenarios. It can rationally divide agent alliances, accurately model agent dynamics, clearly construct communication topologies, and quickly find the optimal decision state through an efficient distributed Nash equilibrium search algorithm, generating effective multi-alliance game strategies. This enables precise control of agents in heterogeneous unmanned swarm systems, improving the overall task execution capability and collaborative performance of the system.

[0045] In another exemplary embodiment of this application, the higher-order heterogeneous linear state dynamics model is as follows:

[0046]

[0047] in, Let represent the update rate of the internal state of the j-th agent in the i-th alliance. express 3D real vector; This represents the system dynamics matrix of the j-th agent in the i-th alliance; This represents the internal state of the j-th agent in the i-th alliance; This represents the system input matrix of the j-th agent in the i-th alliance; This represents the control input of the j-th agent in the i-th alliance. express 3D real vector; Let represent the decision variable of the j-th agent in the i-th alliance. Represents a q-dimensional real column vector; Let represent the decision matrix of the j-th agent in the i-th alliance.

[0048] In another exemplary embodiment of this application, based on the actual communication between each agent in the heterogeneous unmanned swarm system, an internal and external two-layer alliance communication topology diagram of the heterogeneous unmanned swarm system is established, specifically including:

[0049] As shown in the following equation, based on the actual communication between all agents in the heterogeneous unmanned swarm system, an internal and external two-layer alliance communication topology diagram of the heterogeneous unmanned swarm system is established:

[0050]

[0051]

[0052] in, This represents the outer consortium communication topology of a heterogeneous unmanned swarm system, depicting the overall communication situation within and between unmanned swarm consortia; G i This represents the communication topology of the inner alliance of the i-th alliance in a heterogeneous unmanned swarm system, depicting the communication situation within the i-th alliance of the unmanned swarm. E represents the set of all agents in the heterogeneous unmanned swarm system; E represents the set of communication edges between agents in the heterogeneous unmanned swarm system; N represents the total number of alliances. This represents the set of all agents participating in the i-th alliance; This represents the first agent within the i-th alliance; Indicates the m-th element within the i-th alliance i One intelligent agent; m i E represents the number of agents within the i-th alliance; i Represents the set of communication edges between agents within the i-th alliance; This represents a communication edge from the k-th agent in the o-th alliance to the j-th agent in the i-th alliance; This represents the j-th agent within the i-th alliance; This represents the k-th agent in the o-th alliance; This represents a communication edge from the k-th agent in the i-th alliance to the j-th agent in the i-th alliance; This represents the k-th agent within the i-th alliance. The outer alliance communication topology is shown in the diagram. if Then record the communication weight. otherwise, Similarly, for the inner alliance communication topology graph G i ,if Then the communication weight within the alliance is recorded as follows: otherwise,

[0053] In another exemplary embodiment of this application, the agent's decision state is iteratively adjusted based on a leader-follower consistency observer to search for a Nash equilibrium point, specifically including:

[0054] For each agent in each coalition, perform the following operation:

[0055] The agent's auxiliary variables and average gradient estimation variables are initialized to zero, and the observation vector and Nash equilibrium search vector are initialized to zero vectors.

[0056] Based on the leader-follower consistency observer, the update rate of the Nash equilibrium search vector, the update rate of the auxiliary variable, the update rate of the observed variable, and the average gradient estimate variable are calculated using the following formulas:

[0057]

[0058] in, φ represents the Nash equilibrium search vector update rate of the j-th agent in the i-th alliance; φ represents the preset time gain operator; κ1, κ2 and κ3 represent the first gain parameter, the second gain parameter and the third gain parameter, respectively; Let f represent the cost function f of the j-th agent in the i-th alliance for the i-th alliance. i ()about The estimated vector of the partial derivatives; Let represent the auxiliary variable constructed by the j-th agent in the i-th alliance when estimating the k-th agent in the i-th alliance; Representing auxiliary variables The update rate; l represents the agent number within the i-th consortium; m i This represents the number of agents within the i-th alliance; G represents the communication topology graph of the inner consortium of the i-th consortium. i The weight of intra-alliance communication between the l-th agent and the j-th agent; Let f represent the cost function f of the j-th agent in the i-th alliance for the i-th alliance. i ()about The estimated vector of the partial derivatives; Let f represent the cost function f of the l-th agent in the i-th alliance for the i-th alliance. i ()about The estimated vector of the partial derivatives; Let the decision variable of the k-th agent in the i-th alliance be represented. The observed variable update rate of the j-th agent in the i-th alliance for the Nash equilibrium search vector of the k-th agent in the o-th alliance; N represents the total number of alliances; o′ represents the o′-th alliance; m o The number of agents in the 0th alliance is represented by k; k′ represents the number of agents in the k'th alliance. This represents the communication weight of the communication edge between the j-th agent in the i-th alliance and the k'-th agent in the o'-th alliance; The observed variable value represents the value of the Nash equilibrium search vector of the j-th agent in the i-th alliance for the k-th agent in the o-th alliance; The observed variable value represents the value of the Nash equilibrium search vector of the k'th agent in the o'th alliance for the kth agent in the o'th alliance; Let represent the Nash equilibrium search vector of the k-th agent within the o-th alliance; f represents the cost function of the j-th agent in the i-th alliance; i () represents the cost function of the i-th alliance, which is the sum of the cost functions of all agents in the alliance, i.e. Let represent the observation vector of the j-th agent in the i-th alliance.

[0059] Based on the Nash equilibrium search vector update rate, auxiliary variable update rate, observation variable update rate, and average gradient estimation variable, it is determined whether the agent's decision state has reached the Nash equilibrium state.

[0060] If the judgment result is yes, then the Nash equilibrium search vector corresponding to the agent's decision state in the current iteration is the searched Nash equilibrium point.

[0061] If the result is negative, return to step "Based on the leader-follower consistency observer, calculate the Nash equilibrium search vector update rate, auxiliary variable update rate, observation variable update rate, and average gradient estimate variable using the following formulas".

[0062] In another exemplary embodiment of this application, the calculation formula for the preset time gain operator is:

[0063]

[0064] Where φ represents the preset time gain operator; t represents the algorithm running time; T p The second preset time is represented by μ(t); the first auxiliary dynamic function is represented by μ(t). Denotes the derivative of the first auxiliary dynamic function μ(t); ∈T p denoted as the first preset time; where 0 < ∈ < 1, ∈ represents the time ratio parameter; h > 1 represents the rate control parameter of the preset time gain operator.

[0065] In another exemplary embodiment of this application, the agent-based distributed multi-alliance game control method further includes:

[0066] Based on a high-order heterogeneous linear state dynamics model, the adjustment matrix for each agent in each coalition is constructed using the following formula:

[0067]

[0068] in, This represents the system dynamics matrix of the j-th agent within the i-th alliance; Let represent the first adjustment matrix of the j-th agent within the i-th alliance; This represents the system input matrix of the j-th agent in the i-th alliance; Let represent the second adjustment matrix of the j-th agent within the i-th alliance; I represents the decision matrix of the j-th agent in the i-th alliance; q This represents a q×q dimensional identity matrix.

[0069] In another exemplary embodiment of this application, the control input corresponding to each agent is generated based on the Nash equilibrium point searched by the search layer to obtain a multi-alliance game strategy, specifically including:

[0070] Obtain the trace of the system dynamics matrix for each agent.

[0071] If the trace of the system dynamics matrix of the current agent is 0, the time-varying parameters of the j-th agent in the i-th coalition are calculated using the following formula:

[0072]

[0073] in, This represents the time-varying parameters of the j-th agent within the i-th alliance; T represents the first positive constant of the j-th agent within the i-th consortium; t represents the algorithm's running time; p Indicates the second preset time; Let represent the Nash equilibrium search vector of the j-th agent in the i-th alliance.

[0074] If the trace of the system dynamics matrix of the current agent is not 0, the time-varying parameters of the j-th agent in the i-th coalition are calculated using the following formula:

[0075]

[0076] in, Let be the second normal number of the j-th agent within the i-th alliance. This represents the system dynamics matrix of the j-th agent within the i-th alliance; The trace of the system dynamics matrix of the j-th agent within the i-th alliance. The system dynamics matrix dimensionality Let be the third normal number of the j-th agent within the i-th alliance.

[0077] Based on the time-varying parameters of the j-th agent within the i-th consortium, the Lyapunov equation for the time-varying parameters is designed using the following formula to solve for the third output matrix of the j-th agent within the i-th consortium:

[0078]

[0079] in, This represents the transpose of the system dynamics matrix of the j-th agent within the i-th alliance; This represents the third output matrix of the j-th agent within the i-th alliance; This represents the system input matrix of the j-th agent in the i-th alliance; This represents the transpose of the system input matrix of the j-th agent in the i-th alliance.

[0080] The control input for each agent in each alliance is obtained using the following formula:

[0081]

[0082] The following example illustrates this application using a specific heterogeneous intelligent agent distributed multi-alliance game control process.

[0083] Step 1. Establish the dynamic model of decision state of multi-alliance game agents, the definition of Nash equilibrium in multi-alliance game, and the related concepts of topological communication within the alliance and among the agents participating in the game.

[0084] Step 1.1: Dynamic model of decision state of multi-alliance game agents.

[0085] In multi-alliance game problems, let m be the total number of heterogeneous agents participating in the game. all The alliances are formed based on the cooperation and competition relationships reflected in the cost functions of the agents. Let N be the total number of alliances after the division, and m be the number of agents in each alliance. i One intelligent agent, m i ≥2, the number of individuals within the alliance and the total number of agents satisfy the condition. After the partitioning is completed, the set of alliances is: The agents within the alliance are numbered sequentially, and the set of numbers for all agents in alliance i is denoted as .

[0086] The dynamics of each agent in the system are driven by a high-order heterogeneous general linear dynamics model of the following form:

[0087]

[0088] in, Let represent the update rate of the internal state of the j-th agent in the i-th alliance. express 3D real vector; This represents the system dynamics matrix of the j-th agent in the i-th alliance; This represents the internal state of the j-th agent in the i-th alliance; This represents the system input matrix of the j-th agent in the i-th alliance; This represents the control input of the j-th agent in the i-th alliance. express 3D real vector; Let represent the decision variable of the j-th agent in the i-th alliance. Represents a q-dimensional real column vector; Let represent the decision matrix of the j-th agent in the i-th alliance. and They all have matching dimensions.

[0089] Step 1.2: Definition of Nash equilibrium in multi-alliance games.

[0090] Each agent has a cost function, and the cost function of the j-th agent in the i-th alliance is denoted as... Where y i Let $\mathbf{i}$ be the sum vector of the decision variables of all agents participating in the game in the $i$-th alliance. Let represent the decision variables of the first agent participating in the game within the i-th alliance; This represents the transpose of the decision variables of the first agent participating in the game in the i-th alliance; Represents the m-th element in the i-th alliance participating in the game. i Decision variables for each agent; Represents the m-th element in the i-th alliance participating in the game. i Transpose of the decision variables of an agent; y -i Let $\mathbf{i}$ be the sum vector of decision variables for all coalitions except the $i$-th coalition. The cost function consists of two parts: the coalition cost function and the intra-coalition cost function. Cost function between alliances It only relates to the decision variables of all participating agents in the i-th alliance, and This relates to the decision variables of all agents participating in the game. The cost function of the agents satisfies...

[0091] The cost function f of the i-th alliance i (y i ,y -i ) is the sum of the cost functions of each agent, i.e. The goal of each agent participating in the game is to minimize the overall cost function of their alliance by changing their own decision state. However, each agent can only obtain its own cost function.

[0092] If there exists an optimal decision state vector y * =(y i* ,y -i* ),satisfy Then this state vector is the Nash equilibrium solution. i* y represents the sum vector of the optimal decision variables of all participating agents in the i-th alliance; -i* f represents the sum vector of the optimal decision variables for all alliances except the i-th alliance; i (y * ) represents the state vector y of all agents in the i-th alliance. * The cost function under; f i (y i* ,y -i* ) represents the state vector y of all agents in the i-th alliance. * Below, all alliances except the i-th alliance have the same state vector y. -i* The cost function; f i (y i ,y -i* ) represents the state vector of all agents in the i-th alliance. Below, all alliances except the i-th alliance have the same state vector y. -i* The cost function.

[0093] 1.3: Topological communication within the alliance and among all participating agents in the game.

[0094] Based on the actual communication between the agents in the unmanned swarm system, the communication topology of the unmanned swarm is established.

[0095] First, consider the communication topology within Alliance i. in m i ≥2 indicates participation in the set of all agents within the alliance. Let represent the set of communication edges between agents within the i-th alliance, where This refers to a communication edge that points from the j-th agent in the i-th alliance to the k-th agent in the i-th alliance; This represents the j-th agent within the i-th alliance; Let represent the k-th agent within the i-th alliance. If two agents communicate, there is an edge between their corresponding nodes. If there exists a pair of corresponding edges satisfying , then ... If a communication edge represents a link between the k-th agent in the i-th association and the j-th agent in the i-th association, then the graph is a directed graph. The weighted adjacency matrix of the communication topology graph. It is a square matrix, and the value of the element in the k-th row and j-th column is... satisfy If and only if Otherwise, the element's value is 0. Define the j-th node in the i-th alliance. in-degree is Out-degree If all nodes have the same out-degree and in-degree, then the graph is a weighted directed graph. Its Laplacian matrix is ​​defined as follows: Where D i It is a diagonal matrix constructed from the in-degree information of all nodes in the alliance.

[0096] Consider the communication topology among all participating agents, i.e., intra-alliance communication + inter-alliance communication. Let the alliance communication topology of the heterogeneous unmanned swarm system be denoted as follows: but in For a collection of intelligent agents, This represents the set of all intelligent agents in a heterogeneous unmanned swarm system. Let E represent the set of agent communication edges in the heterogeneous unmanned swarm system; and N represent the total number of alliance members. This refers to a communication edge that points from the j-th agent in the i-th alliance to the k-th agent in the o-th alliance. This represents the j-th agent within the i-th alliance; Let $k$ represent the $k$-th agent in the $o$-th alliance. Similarly, its weight adjacency matrix is ​​constructed. When the j-th agent in the i-th alliance receives information from the k-th agent in the o-th alliance, the corresponding element in the adjacency matrix... The value is 1 if the condition is met, and 0 otherwise. Similarly, the in-degree information of the agent in the overall communication topology graph can be calculated, and the Laplace matrix of the overall communication topology graph can be constructed based on this.

[0097] Step 2. Establish the search layer algorithm.

[0098] Step 2.1: Initialize the relevant parameters.

[0099] For the j-th agent in the i-th alliance, initialize all auxiliary variables and the average gradient estimation variable. and Set the value to 0 to initialize the observation vector. With Nash equilibrium search vector It is a zero vector.

[0100] Step 2.2: Establish a preset time gain operator φ.

[0101] Select the first preset time ∈ T p Where 0 < ∈ < 1, T p The second preset time, with the preset time gain operator φ as follows:

[0102]

[0103] Where h > 1 is the preset time gain operator rate control parameter.

[0104] Step 2.3 Design and execute the search layer algorithm, including the Nash equilibrium search vector. Update rate, auxiliary variables Update rate, observed variable Update rate, and average gradient estimate variables Calculation method.

[0105] This section utilizes the preset time gain operator φ from step 2.2. For the j-th agent in the i-th alliance, the search layer algorithm is designed as follows:

[0106]

[0107] in, φ represents the Nash equilibrium search vector update rate of the j-th agent in the i-th alliance; φ represents the preset time gain operator; κ1, κ2 and κ3 represent the first gain parameter, the second gain parameter and the third gain parameter, respectively; Let f represent the cost function f of the j-th agent in the i-th alliance for the i-th alliance. i ()about The estimated vector of the partial derivatives; Let represent the auxiliary variable constructed by the j-th agent in the i-th alliance when estimating the k-th agent in the i-th alliance; Representing auxiliary variables The update rate; l represents the agent number within the i-th consortium; m i This represents the number of agents within the i-th alliance; G represents the communication topology graph of the inner consortium of the i-th consortium. i The weight of intra-alliance communication between the l-th agent and the j-th agent; Let f represent the cost function f of the j-th agent in the i-th alliance for the i-th alliance. i()about The estimated vector of the partial derivatives; Let f represent the cost function f of the l-th agent in the i-th alliance for the i-th alliance. i ()about The estimated vector of the partial derivatives; Let the decision variable of the k-th agent in the i-th alliance be represented. The observed variable update rate of the j-th agent in the i-th alliance for the Nash equilibrium search vector of the k-th agent in the o-th alliance; N represents the total number of alliances; o′ represents the o′-th alliance; m o The number of agents in the 0th alliance is represented by k; k′ represents the number of agents in the k'th alliance. This represents the communication weight of the communication edge between the j-th agent in the i-th alliance and the k'-th agent in the o'-th alliance; The observed variable value represents the value of the Nash equilibrium search vector of the j-th agent in the i-th alliance for the k-th agent in the o-th alliance; The observed variable value represents the value of the Nash equilibrium search vector of the k'th agent in the o'th alliance for the kth agent in the o'th alliance; Let represent the Nash equilibrium search vector of the k-th agent within the o-th alliance; f represents the cost function of the j-th agent in the i-th alliance; i () represents the cost function of the i-th alliance, which is the sum of the cost functions of all agents in the alliance, i.e. Let represent the observation vector of the j-th agent in the i-th alliance.

[0108] Define the observation vector Average gradient estimated variables In the calculation method, It is based on the established observation vector. The gradient term is calculated. represents the observed variable value of the Nash equilibrium search vector of the j-th agent in the i-th alliance for all agents in the o-th alliance; col() represents a new column vector formed by stacking the column vectors in it. The observed variable value represents the value of the Nash equilibrium search vector of the j-th agent in the i-th alliance for the 1-th agent in the o-th alliance; This indicates that the j-th agent in the i-th alliance has a relationship with the m-th agent in the o-th alliance. o The observed variable values ​​of the Nash equilibrium search vector for each agent; Let represent the observed variable value of the j-th agent in the i-th alliance for the Nash equilibrium search vector of all alliances; This represents the observed variable value of the j-th agent in the i-th alliance for the Nash equilibrium search vector of the 1-th alliance; This represents the observed variable value of the j-th agent in the i-th alliance for the Nash equilibrium search vector of the N-th alliance; m all q-dimensional real vector.

[0109] By executing this algorithm, the j-th agent of the i-th alliance can control the Nash equilibrium search vector within a first preset time ∈ T. The Nash equilibrium solution was found.

[0110] Step 3. Output adjustment layer algorithm.

[0111] Step 3.1: Design the first and second adjustment matrices based on the dynamic model.

[0112] For the j-th agent in the i-th alliance, design the first adjustment matrix based on its dynamic model. Second adjustment matrix as follows.

[0113]

[0114] in, This represents the system dynamics matrix of the j-th agent within the i-th alliance; Let represent the first adjustment matrix of the j-th agent within the i-th alliance; This represents the system input matrix of the j-th agent in the i-th alliance; Let represent the second adjustment matrix of the j-th agent within the i-th alliance; I represents the decision matrix of the j-th agent in the i-th alliance; q This represents a q×q dimensional identity matrix.

[0115] Step 3.2: Establish the time-varying parameters of the j-th agent within the i-th coalition of the Lyapunov equations.

[0116] Based on the characteristics of the intelligent agent dynamics model, calculate its system dynamics matrix. traces (The trace is calculated by summing the diagonal elements).

[0117] If the trace of the system dynamics matrix of the current agent is 0, the time-varying parameters of the j-th agent in the i-th coalition are calculated using the following formula:

[0118]

[0119] in, This represents the time-varying parameters of the j-th agent within the i-th alliance; Let i represent the first positive constant of the j-th agent within the i-th alliance. For a relatively large positive number, control parameter The rate of change; t represents the algorithm's running time; T p Indicates the second preset time; Let represent the Nash equilibrium search vector of the j-th agent in the i-th alliance.

[0120] If the trace of the system dynamics matrix of the current agent is not 0, the time-varying parameters of the j-th agent in the i-th coalition are calculated using the following formula:

[0121]

[0122] in, Let be the second normal number of the j-th agent within the i-th alliance. This represents the system dynamics matrix of the j-th agent within the i-th alliance; The trace of the system dynamics matrix of the j-th agent within the i-th alliance. The system dynamics matrix dimensionality Let be the third normal number of the j-th agent within the i-th alliance.

[0123] Step 3.3: Design and calculate the third output matrix

[0124] Solve the Lyapunov equations for time-varying parameters to obtain the third output matrix. This matrix is ​​based on the parameters designed in section 3.2. And satisfy:

[0125]

[0126] in, This represents the transpose of the system dynamics matrix of the j-th agent within the i-th alliance; This represents the third output matrix of the j-th agent within the i-th alliance; This represents the system input matrix of the j-th agent in the i-th alliance; This represents the transpose of the system input matrix of the j-th agent in the i-th alliance.

[0127] Step 3.4: Design and implement agent control inputs.

[0128] Based on the Nash equilibrium search vector designed in the search layer algorithm in step 2 Step 3.1 involves designing the first adjustment matrix. Second adjustment matrix and the third output matrix designed in step 3.3 The control input for the j-th agent within the i-th alliance is constructed as follows:

[0129]

[0130] The algorithm is executed to drive heterogeneous agents to a preset time T. p It converges to the Nash equilibrium solution of the multi-alliance game.

[0131] In another exemplary embodiment of this application, consider a group formation game involving two alliances, whose communication topology is as follows: Figure 2 As shown. The circle represents a second-order dynamic model, specifically, its dynamic model is:

[0132]

[0133] Among them, 0 2×2 I represents a 2×2 zero matrix, and I2 represents a 2×2 identity matrix.

[0134] Figure 2 The agent represented by the square is a first-order dynamical model. Specifically, its dynamical model is as follows:

[0135]

[0136] In the game, each individual's decision variable is their positional information, i.e.

[0137] Consider its participation in a group formation scenario, the cost function f of each agent i j All are determined by the cost function within the alliance. Cost function between alliances Combining the above, the expression is as follows:

[0138]

[0139] The cost function expression within the alliance is as follows:

[0140]

[0141] The inter-alliance cost function expression is:

[0142]

[0143] Using the algorithm of this application, parameter T is selected. p =12s, κ1=3.7, κ2=20, κ3=15, The group formation trajectory diagrams and agent output decision diagrams of each agent based on the multi-alliance game algorithm are obtained as follows: Figure 3 and Figure 4 As shown.

[0144] This application also provides an application scenario in which the above-mentioned heterogeneous intelligent agent distributed multi-alliance game control method is applied. Specifically, the heterogeneous intelligent agent distributed multi-alliance game control method provided in this embodiment can be applied to multi-UAV collaborative task execution scenarios in complex environments. This scenario includes a task planning stage, a target search and localization stage, and a collaborative operation stage. The UAV swarm enters the target search and localization stage from the task planning stage. After using distributed sensing and communication technologies to conduct a comprehensive search and precise localization of the target area, the target location information is obtained, and then the swarm enters the collaborative operation stage. The heterogeneous intelligent agent distributed multi-alliance game control method provided in this embodiment belongs to the intelligent agent collaborative control stage in the collaborative operation stage. Specifically, during collaborative operation, UAVs of different types and with different operational capabilities are divided into multiple alliances. This method establishes a dynamic model, constructs a communication topology graph, uses a distributed Nash equilibrium search algorithm to optimize the decision state, and generates a multi-alliance game strategy, thereby accurately controlling each UAV to carry out collaborative operations on the target and improving task execution efficiency.

[0145] Based on the same inventive concept, this application also provides a heterogeneous agent distributed multi-alliance game control device for implementing the heterogeneous agent distributed multi-alliance game control method described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the heterogeneous agent distributed multi-alliance game control device provided below can be found in the limitations of the heterogeneous agent distributed multi-alliance game control method described above, and will not be repeated here.

[0146] In one exemplary embodiment, such as Figure 5 As shown, a heterogeneous intelligent agent distributed multi-alliance game control device is provided, comprising:

[0147] The alliance partitioning module 301 is used to divide the agents in the heterogeneous unmanned swarm system into multiple alliances; each alliance includes multiple agents.

[0148] The dynamics model building module 302 is used to build a high-order heterogeneous linear state dynamics model for each agent in each alliance.

[0149] The communication topology establishment module 303 is used to establish an internal and external two-layer alliance communication topology diagram of the heterogeneous unmanned swarm system based on the actual communication situation between all intelligent agents in the heterogeneous unmanned swarm system.

[0150] The strategy generation module 304 is used to iteratively optimize the decision states of all agents in a heterogeneous unmanned swarm system based on a high-order heterogeneous linear state dynamics model and an internal and external two-layer alliance communication topology, using a distributed Nash equilibrium search algorithm, until the decision states of all agents converge to a Nash equilibrium state within a preset convergence time, thus obtaining a multi-alliance game strategy. The distributed Nash equilibrium search algorithm includes a search layer and an output adjustment layer. The search layer is used to iteratively adjust the decision states of agents based on a leader-follower consistency observer to search for a Nash equilibrium point. The output adjustment layer is used to generate the control input corresponding to each agent based on the Nash equilibrium point searched by the search layer, thus obtaining a multi-alliance game strategy.

[0151] The control module 305 is used to control the intelligent agents in the heterogeneous unmanned swarm system according to the multi-alliance game strategy.

[0152] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores control data for heterogeneous intelligent agent distributed multi-alliance game theory. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a heterogeneous intelligent agent distributed multi-alliance game theory control method.

[0153] In summary, this application proposes a distributed, pre-set time-based Nash equilibrium search algorithm for multi-alliance games where cooperation and competition coexist. The algorithm consists of a search layer and an output adjustment layer. In the search layer, a leader-follower consistency-based observer design allows agents to search for the Nash equilibrium point of the multi-alliance game simply by interacting with their neighbors, and a time-varying gain function is introduced to achieve convergence within a pre-set time. Subsequently, in the output adjustment layer, based on the search layer results, an output adjustment matrix is ​​constructed, and a time-varying Lyapunov equation is introduced. This adapts to the high-order heterogeneous dynamic characteristics of the agents while adjusting the decision states of the participating heterogeneous multi-agent teams to the optimal Nash equilibrium solution within a pre-set time.

[0154] This algorithm enables high-order heterogeneous general linear agents to search for Nash equilibrium solutions in multi-alliance games by interacting with neighboring agents, even when only knowing their own cost function. The algorithm can be applied to various scenarios, such as maintaining connectivity in multi-cluster unmanned systems and multi-cluster grouping and formation, to control agents to find Nash equilibrium solutions within a preset time.

[0155] Those skilled in the art will understand that Figure 6 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0156] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0157] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0158] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0159] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0160] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0161] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0162] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A heterogeneous agent distributed multi-alliance game control method, characterized in that, The heterogeneous intelligent agent distributed multi-alliance game control method comprises the following steps: Intelligent agents in the heterogeneous unmanned cluster system are divided into multiple alliances; each alliance comprises multiple intelligent agents; A high-order heterogeneous linear state dynamics model is established for each intelligent agent in each alliance; Based on actual communication conditions between all intelligent agents in the heterogeneous unmanned cluster system, an inner-outer double-layer alliance communication topology graph of the heterogeneous unmanned cluster system is established; Based on the high-order heterogeneous linear state dynamics model and the inner-outer double-layer alliance communication topology graph, the decision states of all intelligent agents in the heterogeneous unmanned cluster system are iteratively optimized through a distributed Nash equilibrium search algorithm until the decision states of all intelligent agents converge to a Nash equilibrium state within a preset convergence time, thereby obtaining a multi-alliance game strategy; the distributed Nash equilibrium search algorithm comprises a search layer and an output adjustment layer; the search layer is configured to iteratively adjust the decision state of an intelligent agent based on a leader-follower consensus observer to search for a Nash equilibrium point; and the output adjustment layer is configured to generate a control input corresponding to each intelligent agent according to the Nash equilibrium point searched by the search layer to obtain the multi-alliance game strategy; The intelligent agents in the heterogeneous unmanned cluster system are controlled according to the multi-alliance game strategy. 2.The method of claim 1, wherein, The high-order heterogeneous linear state dynamics model is as follows: wherein, denotes the update rate of the internal state of the jth agent in the ith coalition, denotes a q-dimensional real vector; denotes the system dynamics matrix of the jth agent in the ith coalition; denotes the internal state of the jth agent in the ith coalition; denotes the system input matrix of the jth agent in the ith coalition; denotes the control input of the jth agent in the ith coalition, denotes a q-dimensional real vector; denotes the decision variable of the jth agent in the ith coalition, denotes a q-dimensional real column vector; denotes the decision matrix of the jth agent in the ith coalition. 3.The method of claim 1, wherein, Based on actual communication conditions between all intelligent agents in the heterogeneous unmanned cluster system, an inner-outer double-layer alliance communication topology graph of the heterogeneous unmanned cluster system is established, and specifically comprises the following steps: Based on actual communication conditions between all intelligent agents in the heterogeneous unmanned cluster system, an inner-outer double-layer alliance communication topology graph of the heterogeneous unmanned cluster system is established, and specifically comprises the following steps: in, This represents the outer consortium communication topology of a heterogeneous unmanned swarm system, depicting the overall communication situation within and between unmanned swarm consortia; G i This represents the communication topology of the inner alliance of the i-th alliance in a heterogeneous unmanned swarm system. E represents the set of all agents in the heterogeneous unmanned swarm system; E represents the set of communication edges between agents in the heterogeneous unmanned swarm system; N represents the total number of alliances. This represents the set of all agents participating in the i-th alliance; This represents the first agent within the i-th alliance; Indicates the m-th element within the i-th alliance i One intelligent agent; m i E represents the number of agents within the i-th alliance; i Represents the set of communication edges between agents within the i-th alliance; This represents a communication edge from the k-th agent in the o-th alliance to the j-th agent in the i-th alliance; This represents the j-th agent within the i-th alliance; This represents the k-th agent in the o-th alliance; This represents a communication edge from the k-th agent in the i-th alliance to the j-th agent in the i-th alliance; Let i represent the k-th agent within the i-th alliance. 4.The method of claim 1, wherein, The leader-follower consensus observer is configured to iteratively adjust the decision state of an intelligent agent to search for a Nash equilibrium point, and specifically comprises the following steps: For each intelligent agent in each alliance, the following operations are performed: The auxiliary variable and the average gradient estimation variable of the intelligent agent are initialized to zero, and the observation vector and the Nash equilibrium search vector are initialized to a zero vector; Based on the leader-follower consensus observer, the Nash equilibrium search vector update rate, the auxiliary variable update rate, the observation variable update rate, and the average gradient estimation variable are calculated through the following formula: wherein, denotes the Nash equilibrium search vector update rate of the jth agent of the ith coalition; φ denotes a preset time gain operator; κ1, κ2, and κ3 denote a first gain parameter, a second gain parameter, and a third gain parameter, respectively; denotes the jth agent of the ith coalition for the cost function f i of the ith coalition with respect to the estimation vector of the partial derivative of denotes the auxiliary variable constructed by the jth agent of the ith coalition for the kth agent of the ith coalition at time denotes the update rate of the auxiliary variable ; l denotes the agent number within the ith coalition; m i denotes the number of agents within the ith coalition; denotes the intra-coalition communication weight between the lth agent and the jth agent in the coalition communication topology graph G i of the ith coalition; denotes the jth agent of the ith coalition for the cost function f i of the ith coalition with respect to the estimation vector of the partial derivative of denotes the lth agent of the ith coalition for the cost function f i of the ith coalition with respect to the estimation vector of the partial derivative of denotes the decision variable of the kth agent of the ith coalition; denotes the observation variable update rate of the Nash equilibrium search vector of the jth agent of the ith coalition for the kth agent within the oth coalition; N denotes the total number of coalitions; o' denotes the o'th coalition; m o denotes the number of agents within the oth coalition; k' denotes the k'th agent; denotes the communication weight of the communication edge of the jth agent of the ith coalition pointing to the k'th agent of the o'th coalition; denotes the observation variable value of the Nash equilibrium search vector of the jth agent of the ith coalition for the kth agent within the oth coalition; denotes the observation variable value of the Nash equilibrium search vector of the k'th agent of the o'th coalition for the kth agent within the oth coalition; denotes the Nash equilibrium search vector of the kth agent within the oth coalition; denotes the cost function of the jth agent of the ith coalition; f i denotes the cost function of the ith coalition; denotes the observation value vector of the jth agent of the ith coalition; Based on the Nash equilibrium search vector update rate, the auxiliary variable update rate, the observation variable update rate, and the average gradient estimation variable, it is determined whether the decision state of the intelligent agent reaches a Nash equilibrium state; If the determination result is yes, the Nash equilibrium search vector corresponding to the decision state of the intelligent agent in the current iteration is the searched Nash equilibrium point; If the determination result is no, the step of calculating the Nash equilibrium search vector update rate, the auxiliary variable update rate, the observation variable update rate, and the average gradient estimation variable based on the leader-follower consensus observer is returned to.

5. The distributed multi-alliance game control method of heterogeneous intelligent agents according to claim 4, characterized in that, The calculation formula of the preset time gain operator is as follows: wherein φ represents a preset time gain operator; t represents an algorithm running time; T p represents a second preset time; μ(t) represents a first auxiliary dynamic function; represents a derivative of the first auxiliary dynamic function μ(t); ∈T p represents a first preset time; wherein 0<∈<1, ∈ represents a time proportion parameter; h>1, represents a rate control parameter of the preset time gain operator. 6.The method of claim 1, wherein, The heterogeneous intelligent agent distributed multi-alliance game control method further comprises the following steps: Based on the high-order heterogeneous linear state dynamics model, an adjustment matrix of each intelligent agent in each alliance is constructed through the following formula: wherein, denotes the system dynamics matrix of the jth agent within the ith coalition; denotes the first conditioning matrix of the jth agent within the ith coalition; denotes the system input matrix of the jth agent within the ith coalition; denotes the second conditioning matrix of the jth agent within the ith coalition; denotes the decision matrix of the jth agent within the ith coalition;I q denotes the q x q identity matrix.

7. The method of claim 1, wherein, According to the Nash equilibrium point searched by the search layer, a control input corresponding to each intelligent agent is generated to obtain the multi-alliance game strategy, and specifically comprises the following steps: The trace of the system dynamics matrix of each intelligent agent is obtained; If the trace of the system dynamics matrix of the current agent is 0, the time-varying parameter of the jth agent in the ith alliance is calculated by the following formula: wherein, represents a time-varying parameter of the jth agent in the ith alliance; represents a first constant of the jth agent in the ith alliance; t represents the running time of the algorithm; T p represents a second preset time; represents a Nash equilibrium search vector of the jth agent of the ith alliance; If the trace of the system dynamics matrix of the current agent is not 0, the time-varying parameter of the jth agent in the ith alliance is calculated by the following formula: wherein, is a second normal number of the jth agent within the ith coalition, denotes the system dynamics matrix of the jth agent within the ith coalition; denotes the trace of the system dynamics matrix of the jth agent within the ith coalition, is the dimension of the system dynamics matrix is a third normal number of the jth agent within the ith coalition; ​ Based on the time-varying parameter of the jth agent in the ith alliance, the third output matrix of the jth agent in the ith alliance is solved by the following formula of the time-varying parameter Lyapunov equation: wherein, denotes the transpose of the system dynamics matrix of the jth agent within the ith coalition; denotes the third output matrix of the jth agent within the ith coalition; denotes the system input matrix of the jth agent within the ith coalition; denotes the transpose of the system input matrix of the jth agent within the ith coalition; The control input of each agent in each alliance is obtained by the following formula:

8. A heterogeneous agent distributed multi-alliance game control device, characterized in that, The heterogeneous agent distributed multi-alliance game control device applies the heterogeneous agent distributed multi-alliance game control method in any one of claims 1-7, and the heterogeneous agent distributed multi-alliance game control device comprises: An alliance division module is configured to divide the agents in the heterogeneous unmanned swarm system into multiple alliances; each alliance comprises multiple agents; A dynamics model construction module is configured to respectively establish a high-order heterogeneous linear state dynamics model for each agent in each alliance; A communication topology establishment module is configured to establish an inner-outer double-layer alliance communication topology graph of the heterogeneous unmanned swarm system based on actual communication conditions between all agents in the heterogeneous unmanned swarm system; A strategy generation module is configured to perform iterative optimization on decision states of all agents in the heterogeneous unmanned swarm system based on the high-order heterogeneous linear state dynamics model and the inner-outer double-layer alliance communication topology graph by using a distributed Nash equilibrium search algorithm until the decision states of all agents converge to a Nash equilibrium state within a preset convergence time, so as to obtain a multi-alliance game strategy; the distributed Nash equilibrium search algorithm comprises a search layer and an output adjustment layer; the search layer is configured to iteratively adjust the decision states of the agents based on a leader-follower consensus observer to search for a Nash equilibrium point; and the output adjustment layer is configured to generate a control input corresponding to each agent according to the Nash equilibrium point searched by the search layer, so as to obtain the multi-alliance game strategy. A control module is configured to control the agents in the heterogeneous unmanned swarm system according to the multi-alliance game strategy.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the heterogeneous agent distributed multi-alliance game control method in any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the heterogeneous agent distributed multi-alliance game control method in any one of claims 1-7.

Citation Information

Cited By

  • Heterogeneous multi-agent distributed Nash equilibrium optimization adaptive fault-tolerant control method

    CN122085707A

  • Adaptive fault-tolerant control method for heterogeneous multi-agent distributed nash equilibrium optimization

    CN122085707B