An agent control method, device and equipment for multi-alliance game and a medium
By dividing alliances in multi-alliance games, establishing a directed unbalanced communication topology graph, and designing a preset time gain operator, the problem of low agent convergence efficiency caused by the directed unbalanced communication topology graph is solved by utilizing the left zero eigenvector of the Laplace matrix and the Nash equilibrium search term. This enables rapid control of unmanned swarms and rapid search for Nash equilibrium solutions.
Patent Information
- Application Number
- CN202510037892.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-01-10
AI Technical Summary
In multi-alliance games, existing technologies suffer from low efficiency in converging agents to Nash equilibrium solutions due to the directed unbalanced communication topology, making it difficult to quickly iterate to the optimal decision state.
By dividing the agents in the unmanned swarm system into multiple alliances, establishing a directed unbalanced communication topology graph, determining the target alliance and the estimated vectors of the agents, designing a preset time gain operator, and using the left zero eigenvector of the Laplace matrix and the average gradient estimation term, combined with the Nash equilibrium search term, the agents can converge to the Nash equilibrium solution within a preset time.
It improves the iterative evolution efficiency of intelligent agents in games, realizes rapid control of unmanned swarms, can search for Nash equilibrium solutions within a preset time, and improves the convergence speed of multi-alliance games.
Smart Images

Figure CN119882439B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of agent control, in particular to an agent control method and device for multi-alliance game, equipment and medium. BACKGROUND
[0002] With the development and application of communication networking technology and unmanned system technology, the research on multi-agent system has attracted extensive attention and application. In recent years, the research on multi-alliance game algorithm has become a focus problem. The purpose of game is to obtain greater benefits. In the game, an important concept and strategy set is Nash equilibrium, which describes an optimal strategy set for different agents in the game. In this state, any agent cannot obtain greater benefits by changing its own decision. Therefore, each agent changes its decision state in the game to make its alliance obtain greater benefits.
[0003] A coalition is composed of agents with common interests; multiple agents with different interests form multiple alliances. Unlike the basic inter-agent game problem, cooperation and game relationship coexist in the multi-alliance game problem. In multi-alliance game, each agent forms a game alliance due to its own interest relationship. Within the alliance, each agent cooperates with each other to optimize the alliance cost function, and competes with other alliances. The alliance can be regarded as a virtual individual participating in the game of other alliances, but the actual decision maker is each agent participating in the game.
[0004] In practice, the speed of algorithm convergence is an important evaluation index of the performance of the algorithm. The preset time algorithm can preset a desired algorithm convergence time in advance. By designing time-varying control gain, the algorithm completes convergence in the preset time. In addition, due to the differences in signal transmission power and communication channel establishment, directed unbalanced communication topology is more common in practice. However, most related methods are for relatively simple communication topologies.
[0005] Therefore, eliminating the interference caused by the directed unbalanced communication topology of the unmanned cluster and improving the efficiency of the agent converging to the Nash equilibrium solution can help the agent to evolve to the optimal decision state faster in the game, so as to realize the rapid control of the unmanned cluster. SUMMARY
[0006] The purpose of the present application is to provide an agent control method and device for multi-alliance game, equipment and medium, which can eliminate the interference caused by the directed unbalanced communication topology of the unmanned cluster system and improve the efficiency of the agent converging to the Nash equilibrium solution, thereby providing a fast solution method for the unmanned cluster control problem from the perspective of game, so as to realize the rapid control of the unmanned cluster.
[0007] To achieve the above object, the application provides the following scheme.
[0008] In a first aspect, the application provides a multi-alliance game agent control method, comprising:
[0009] Dividing agents in an unmanned cluster system into multiple alliances; each of the alliances comprises multiple agents;
[0010] Based on actual communication conditions between the agents in the unmanned cluster system, establishing a directed unbalanced communication topology graph of the unmanned cluster system;
[0011] According to the agents in all the alliances, establishing a multi-alliance game agent decision state dynamics model;
[0012] According to the multi-alliance game agent decision state dynamics model, determining a target alliance; the target alliance is an alliance in which there is an agent satisfying a trigger condition; the agent satisfying the trigger condition in the target alliance is a target agent;
[0013] According to the directed unbalanced communication topology graph, determining a Laplace matrix of the target alliance and an estimated vector of the target agent in the target alliance; the estimated vector converges to a left zero eigenvector of the Laplace matrix of the target alliance to which the target agent belongs at a first preset time;
[0014] According to the estimated vector of the target agent, determining an average gradient estimation item of the target agent; the average gradient estimation item converges to an average estimation value of an overall cost function of the target alliance to which the target agent belongs at a second preset time;
[0015] According to the average gradient estimation item of the target agent, determining a Nash equilibrium search item of the target agent; the Nash equilibrium search item converges to a Nash equilibrium solution of the target agent at the second preset time; the Nash equilibrium solution is a decision state vector that minimizes the overall cost function of the target alliance to which the target agent belongs;
[0016] According to the Nash equilibrium search items of all the target agents in all the target alliances, controlling the agents in the unmanned cluster system.
[0017] Optionally, the expression of the multi-alliance game agent decision state dynamics model is:
[0018]
[0019] y ij =C ij x ij ;
[0020] wherein, yij denotes the decision state vector of the jth agent in the ith coalition, y ij ∈R q , R q denotes a q-dimensional real vector, u ij denotes the control input of the jth agent in the ith coalition, denotes a p ij -dimensional real vector, x ij denotes the internal state of the jth agent in the ith coalition, denotes an n ij -dimensional real vector, denotes the update rate of the internal state of the jth agent in the ith coalition, A ij denotes the system dynamics matrix of the jth agent in the ith coalition, B ij denotes the system input matrix of the jth agent in the ith coalition, C ij denotes the decision matrix of the jth agent in the ith coalition.
[0021] Optionally, based on the actual communication between each agent in the unmanned swarm system, a directed unbalanced communication topology graph of the unmanned swarm system is established, specifically comprising:
[0022] Based on the actual communication between each agent in the unmanned swarm system, a communication topology within the ith coalition is established wherein, denotes the set of all agents participating in the ith coalition, m i denotes the number of agents in the ith coalition, denotes the 1st agent in the ith coalition, denotes the m i th agent in the ith coalition; E i denotes the set of communication edges of agents in the ith coalition, denotes a communication edge from the jth agent in the ith coalition to the kth agent in the ith coalition, denotes the jth agent in the ith coalition, denotes the kth agent in the ith coalition;
[0023] A weight adjacency matrix of the ith coalition is established according to the communication topology within the ith coalition; the weight adjacency matrix is a square matrix, when the element value in the jth column of the kth row of the weight adjacency matrix of the ith coalition is otherwise
[0024] determine the in-degree of the jth agent in the ith alliance and the out-degree of the jth agent in the ith alliance according to the weighted adjacency matrix of the ith alliance; the in-degree of the jth agent in the ith alliance denotes an element value in the jth row and the kth column of the weighted adjacency matrix of the ith alliance, and the out-degree of the jth agent in the ith alliance
[0025] construct a directed unbalanced communication topology graph of the unmanned cluster according to the in-degree of all agents in the ith alliance and the out-degree of all agents in the ith alliance.
[0026] Optionally, according to the directed unbalanced communication topology graph, determine the Laplacian matrix of the target alliance and the estimated vector of the target agent in the target alliance, specifically comprising:
[0027] determine the Laplacian matrix of the target alliance; the expression of the Laplacian matrix is:
[0028]
[0029] wherein, L i is the Laplacian matrix of the ith target alliance, D i is a diagonal matrix constructed by the out-degree of all agents in the ith target alliance, is the weighted adjacency matrix of the ith alliance;
[0030] establish a first preset time gain operator; the expression of the first preset time gain operator is:
[0031]
[0032] wherein, h>1, h is a first preset time gain operator gain rate control parameter, T is a second preset time, and φ(t;T) is a first preset time gain operator with a preset time of T, t is time;
[0033] determine the estimated vector of the target agent in the target alliance according to the first preset time gain operator and the directed unbalanced communication topology graph; the expression of the estimated vector is:
[0034]
[0035] wherein, z ij is the estimated vector of the jth target agent in the ith target alliance, is the update rate of the estimated vector of the jth target agent in the ith target alliance, φ(t;∈T) is a first preset time gain operator with a preset time of ∈T, ∈T is a first preset time, and 0<∈<1.
[0036] Optionally, the average gradient estimation term of the target agent is determined according to the estimation vector of the target agent, and specifically includes:
[0037] A second preset time gain operator is established; an expression of the second preset time gain operator is:
[0038]
[0039] wherein, is the second preset time gain operator, and φ(t; T) is the first preset time gain operator with a preset time T;
[0040] An auxiliary variable of a jth target agent in an ith target coalition is initialized A kth element of the average gradient estimation term g ij of the jth target agent is 0, i ∈ {1, 2, …, N}, N is a number of target coalitions participating in the game, j, k ∈ {1, 2, …, m i};
[0041] The average gradient estimation term of the target agent is determined according to the estimation vector of the target agent; an expression of the average gradient estimation term of the target agent is:
[0042]
[0043] wherein, f ij () is a cost function of a jth target agent in an ith target coalition, y represents a union of decision state vectors of all target agents in all target coalitions, η represents a union of Nash equilibrium search terms of all target agents in all target coalitions, η i represents a union of Nash equilibrium search terms of all target agents in an ith target coalition, N is a number of target coalitions participating in the game, represents a Nash equilibrium search term of an m i th target agent in an ith target coalition, represents an m i th target agent in an ith target coalition, is an update rate of , z ijj is a jth element of an estimation vector z ij of a jth target agent in an ith target coalition, is a second preset time gain operator, g ijk is a kth element of an average gradient estimation term g ij of a jth target agent in an ith target coalition.
[0044] Optionally, the Nash equilibrium search term of the target agent is determined according to the average gradient estimation term of the target agent, and specifically includes:
[0045] The Nash equilibrium search term of the target agent is initialized as a zero vector;
[0046] The Nash equilibrium search term of the target agent is determined according to the average gradient estimation term of the target agent; and an expression of the Nash equilibrium search term of the target agent is:
[0047]
[0048] wherein η ij is the Nash equilibrium search term of the jth target agent in the ith target coalition, is an update rate of η ij , κ is a Nash equilibrium search rate control parameter, κ>0, g ijj is the jth element of the average gradient estimation term g ij of the jth target agent in the ith target coalition.
[0049] Optionally, the agents in the unmanned swarm system are controlled according to the Nash equilibrium search terms of all target agents in all target coalitions, and specifically includes:
[0050] The first gain matrix and the second gain matrix are constructed according to the multi-coalition game agent decision state dynamics model; and a calculation formula of the first gain matrix and the second gain matrix is:
[0051]
[0052] wherein K ij1 is the first gain matrix of the jth target agent in the ith target coalition, K ij2 is the second gain matrix of the jth target agent in the ith target coalition, I q is a q×q dimensional unit matrix, diagonal elements are 1, and other elements are 0;
[0053] The control input of all target agents in all target coalitions is obtained according to the Nash equilibrium search terms of all target agents in all target coalitions, the first gain matrix and the second gain matrix; and a formula of the control input of the target agent is:
[0054]
[0055] wherein u' ij is the control input of the jth target agent in the ith target coalition;
[0056] control the agents in the unmanned swarm system according to the control inputs of all target agents in all target coalitions.
[0057] In a second aspect, the present application provides an agent control device for multi-coalition game, comprising:
[0058] a coalition division module, configured to divide the agents in the unmanned swarm system into a plurality of coalitions; each of the coalitions comprises a plurality of agents;
[0059] a directed unbalanced communication topology graph establishment module, configured to establish a directed unbalanced communication topology graph of the unmanned swarm system based on actual communication conditions between the agents in the unmanned swarm system;
[0060] a model establishment module, configured to establish a multi-coalition game agent decision state dynamics model according to the agents in all coalitions;
[0061] a target coalition determination module, configured to determine a target coalition according to the multi-coalition game agent decision state dynamics model; the target coalition is a coalition in which there is an agent satisfying a trigger condition; the agent satisfying the trigger condition in the target coalition is a target agent;
[0062] an estimation vector determination module, configured to determine a Laplacian matrix of the target coalition and an estimation vector of the target agent in the target coalition according to the directed unbalanced communication topology graph; the estimation vector converges to a left zero eigenvector of the Laplacian matrix of the target coalition to which the target agent belongs at a first preset time;
[0063] an average gradient estimation term determination module, configured to determine an average gradient estimation term of the target agent according to the estimation vector of the target agent; the average gradient estimation term converges to an average estimation value of an overall cost function of the target coalition to which the target agent belongs at a second preset time;
[0064] a Nash equilibrium search term determination module, configured to determine a Nash equilibrium search term of the target agent according to the average gradient estimation term of the target agent; the Nash equilibrium search term converges to a Nash equilibrium solution of the target agent at the second preset time; the Nash equilibrium solution is a decision state vector that minimizes the overall cost function of the target coalition to which the target agent belongs;
[0065] a control module, configured to control the agents in the unmanned swarm system according to the Nash equilibrium search terms of all target agents in all target coalitions.
[0066] In a third aspect, the present application provides a computer device, comprising: a memory, a processor to store a computer program on the memory and run the computer program on the processor, and the processor executes the computer program to implement the steps of the agent control method for multi-alliance game according to any one of the above.
[0067] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the agent control method for multi-alliance game according to any one of the above.
[0068] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0069] The present application provides an agent control method, device, equipment and medium for multi-alliance game. The estimation vector is determined through the directed unbalanced communication topology graph of the unmanned cluster. Through the determination of the estimation vector, the problem that the weight value is not matched due to the direct communication between agents in the directed unbalanced communication topology graph when the average gradient estimation term is calculated subsequently, so that the average gradient estimation term cannot be obtained, is solved. Then, based on the average gradient estimation term, the problem of the limitation of the agent which can only obtain the own cost function but needs to optimize the whole alliance cost function is solved, and the Nash equilibrium search term is designed in combination with the estimation vector. Based on the Nash equilibrium search term, the target agent is driven to converge to the Nash equilibrium solution at the second preset time, the efficiency of the agent converging to the Nash equilibrium solution is improved, which helps the agent to evolve to the optimal decision state faster in the game, so as to realize the rapid control of the unmanned cluster. BRIEF DESCRIPTION OF DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0071] Figure 1 The flowchart of the agent control method for multi-alliance game in an embodiment of the present application is shown in the figure.
[0072] Figure 2 The communication topology relationship diagram provided by an embodiment of the present application is shown in the figure.
[0073] Figure 3 The agent decision trajectory diagram in the multi-alliance game is shown in the figure.
[0074] Figure 4 The estimation error curve of the left zero characteristic vector is shown in the figure.
[0075] Figure 5 A functional module schematic diagram of an agent control device for a multi-alliance game according to an embodiment of the present application is provided.
[0076] Figure 6 A structural schematic diagram of a computer device according to an embodiment of the present application is provided. DETAILED DESCRIPTION
[0077] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0078] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0079] Regarding the multi-alliance game algorithm, Ye et al. proposed a multi-alliance game Nash equilibrium search algorithm based on dynamic average consensus estimation and singular perturbation principle, which estimates the average consensus of the gradient. Liu et al. introduced the method of penalty function and studied the multi-alliance game algorithm of second-order agents under the condition of existing constraints. Deng et al. studied the multi-alliance game Nash equilibrium search algorithm of second-order agents under the more general communication condition of balanced directed graph. However, it is worth noting that these algorithms can only obtain asymptotic or exponential convergence results, and are all for relatively simple communication topologies. In practice, the speed of algorithm convergence is an important evaluation index of the performance of the algorithm, and the preset time algorithm can preset a desired algorithm convergence time in advance, and through the design of time-varying control gain, the algorithm is controlled to complete convergence in the preset time. In addition, due to the differences in signal transmission power and communication channel establishment, unbalanced directed communication topology graph is more common in practice. Therefore, the present application proposes an agent control method for multi-alliance game to solve the influence brought by the unbalanced directed communication topology graph and realize the rapid control of unmanned clusters.
[0080] In one exemplary embodiment, as shown in Figure 1 a multi-alliance game agent control method is provided, comprising:
[0081] S1: dividing the agents in the unmanned cluster system into multiple alliances; each alliance includes multiple agents. Specifically, the alliances are divided according to the cooperation and competition relationship reflected by the cost function.
[0082] S2: Based on the actual communication between each agent in the unmanned cluster system, a directed unbalanced communication topology graph of the unmanned cluster system is established. Specifically, it includes:
[0083] Based on the actual communication between each agent in the unmanned cluster system, the communication topology within the i-th alliance is established Wherein, denotes the set of all agents participating in the i-th alliance, m i denotes the number of agents in the i-th alliance, m i ≥2, denotes the first agent in the i-th alliance, denotes the m i th agent in the i-th alliance; E i denotes the set of agent communication edges in the i-th alliance, denotes a communication edge from the jth agent in the i-th alliance to the kth agent in the i-th alliance, denotes the jth agent in the i-th alliance, denotes the kth agent in the i-th alliance. If two agents communicate, there is an edge between the nodes corresponding to the two agents. If there is a pair of corresponding edges, satisfying then the graph is a directed graph.
[0084] According to the communication topology within the i-th alliance, the weight adjacency matrix of the i-th alliance is established; the weight adjacency matrix is a square matrix, when the element value of the jth row and kth column in the weight adjacency matrix of the i-th alliance is otherwise
[0085] According to the weight adjacency matrix of the i-th alliance, the in-degree of the jth agent in the i-th alliance and the out-degree of the jth agent in the i-th alliance are determined; the in-degree of the jth agent in the i-th alliance denotes the element value of the jth row and kth column in the weight adjacency matrix of the i-th alliance, the out-degree of the jth agent in the i-th alliance If there is a node whose out-degree and in-degree are not equal, then the graph is a directed unbalanced graph.
[0086] The directed unbalanced communication topology graph of the unmanned cluster is constructed by the in-degree of all agents in the i-th alliance and the out-degree of all agents.
[0087] Further consider the communication topology between alliances which depicts the communication relationship between alliances, similar to the implementation scheme described in the above method, which will not be repeated here. The communication topology of the whole multi-alliance game is Intra-alliance communication + inter-alliance communication.
[0088] S3: Establish a multi-alliance game agent decision state dynamics model according to the agents in all alliances.
[0089] The dynamics of each agent in the system is driven by a model having the following form, and the expression of the multi-alliance game agent decision state dynamics model is:
[0090]
[0091] wherein y ij represents the decision state vector of the jth agent in the ith alliance, y ij ∈R q , R q represents a q-dimensional real vector, u ij represents the control input of the jth agent in the ith alliance, represents a p ij -dimensional real vector, x ij represents the internal state of the jth agent in the ith alliance, represents an n ij -dimensional real vector, represents the update rate of the internal state of the jth agent in the ith alliance, A ij represents the system dynamics matrix of the jth agent in the ith alliance, B ij represents the system input matrix of the jth agent in the ith alliance, C ij represents the decision matrix of the jth agent in the ith alliance.
[0092] S4: Determine the target alliance according to the multi-alliance game agent decision state dynamics model; the target alliance is an alliance in which there exists an agent satisfying a trigger condition; the agent satisfying the trigger condition in the target alliance is a target agent. The trigger condition is that the row of C ij B ij is full rank.
[0093] S5: Determine the Laplacian matrix of the target alliance and the estimation vector of the target agent in the target alliance according to the directed unbalanced communication topology graph; the estimation vector converges to the left zero eigenvector of the Laplacian matrix of the target alliance to which the target agent belongs at a first preset time.
[0094] In this step, for the established communication topology graph , the Laplacian matrix L i of the communication topology graph is established, and the Laplacian matrix L ian estimated vector of the left zero eigenvector of the Laplacian matrix of the target alliance, so as to eliminate the influence that the direct communication among agents in the subsequent average gradient term estimation cannot obtain the average gradient estimation value due to the weight mismatch in the unbalanced directed communication topology graph. According to the directed unbalanced communication topology graph, the Laplacian matrix of the target alliance and the estimated vector of the target agent in the target alliance are determined, and the specific steps include:
[0095] determining the Laplacian matrix of the target alliance; the expression of the Laplacian matrix is:
[0096]
[0097] wherein, L i is the Laplacian matrix of the i th target alliance, D i is a diagonal matrix constructed by the out-degree of all agents in the i th target alliance, is the weight adjacency matrix of the i th alliance.
[0098] establishing a first preset time gain operator; the expression of the first preset time gain operator is:
[0099]
[0100] wherein, is the left derivative of μ(t), is an intermediate auxiliary variable for constructing φ(t;T), h>1, h is a first preset time gain operator gain rate control parameter, h is a constant, the larger the value is, the larger the gain is, φ(t;T) is the first preset time gain operator with a preset time T, T is a second preset time and can be set to any value, t is time (i.e. the current time of algorithm running, such as running for 3s, t is 3), the first preset time gain operator is a function related to time t, and time is the independent variable.
[0101] in the estimated vector of the j th target agent in the i th target alliance, the initial parameter value z ijj (0) = 1, z ijk (0) = 0, k≠j, wherein z ijj represents the j th element of the vector z ij , z ijk represents the k th element of the vector z ij , and (0) represents the value of the element at the initial time, i.e. t = 0.
[0102] determining the estimated vector of the target agent in the target alliance according to the first preset time gain operator and the directed unbalanced communication topology graph; the expression of the estimated vector is:
[0103]
[0104] wherein, zij is the estimation vector of the jth target agent in the ith target coalition, is the update rate of the estimation vector of the jth target agent in the ith target coalition, φ(t; ∈T) is a first preset time gain operator with preset time ∈T, ∈T is a first preset time, and 0 < ∈ < 1.
[0105] By applying the algorithm, z ij converges to the left zero eigenvector ξ i of the Laplacian matrix L i in the first preset time ∈T. i T L i = 0 ξ i T .
[0106] S6: determining an average gradient estimation item of the target agent according to the estimation vector of the target agent; the average gradient estimation item converges to an average estimation value of an overall cost function of a target coalition to which the target agent belongs in a second preset time
[0107] Specifically, S6 includes the following steps.
[0108] establishing a second preset time gain operator; an expression of the second preset time gain operator is as follows:
[0109]
[0110] wherein, is the second preset time gain operator, and φ(t; T) is the first preset time gain operator with preset time T.
[0111] initializing the kth element of an auxiliary variable z of the jth target agent in the ith target coalition and the kth element of the average gradient estimation item g ij , i ∈ {1, 2,..., N}, N is the number of coalitions participating in the game, j, k ∈ {1, 2,..., m i}.
[0112] determining the average gradient estimation item of the target agent according to the estimation vector of the target agent; an expression of the average gradient estimation item of the target agent is as follows:
[0113]
[0114] wherein, f ij () is a cost function of the jth target agent in the ith target coalition, and y represents a union of decision state vectors of all target agents in all target coalitions. η represents the set of Nash equilibrium search terms for all target agents in all target coalitions. i Let N represent the set of Nash equilibrium search terms for all target agents in the i-th target alliance, where N is the number of target alliances participating in the game. In the i-th target alliance, the m-th... i Nash equilibrium search terms for a target intelligent agent m i q-dimensional real vector for The update rate, z ijj Let z be the estimated vector of the j-th target agent in the i-th target coalition. ij The j-th element, It is the second preset time gain operator, g ijk The average gradient estimate term g for the j-th target agent in the i-th target coalition. ij The kth element.
[0115] S7: Determine the Nash equilibrium search term of the target agent based on the average gradient estimation term of the target agent; the Nash equilibrium search term converges to the Nash equilibrium solution of the target agent at the second preset time; the Nash equilibrium solution is the decision state vector that minimizes the overall cost function of the target alliance to which the target agent belongs.
[0116] Each agent has a corresponding cost function, and the cost function of the j-th agent in the i-th alliance is denoted as f. ij (y i ,y -i ), where y i Let $\mathbf{i}$ be the sum vector of the decision variables of all agents participating in the game in the $i$-th alliance. y -i This represents the resultant vector of decision variables from other alliances, i.e.
[0117] The cost function f of the alliance i (y i ,y -i ) represents the sum of the cost functions of each agent in the alliance, i.e.
[0118] The desired Nash equilibrium solution (optimal decision state vector): The goal of each agent participating in the game is to minimize the overall cost function of its alliance by changing its own decision state.
[0119] If there exists an optimal decision state vector satisfy Then the optimal decision state vector is the Nash equilibrium solution. It can be seen that the Nash equilibrium solution is the optimal decision state pursued by all alliances and intelligent agents in the alliance.
[0120] S7 specifically comprises:
[0121] Initialize the Nash equilibrium search term of the target intelligent agent to a zero vector.
[0122] According to the average gradient estimation term of the target intelligent agent, determine the Nash equilibrium search term of the target intelligent agent; the expression of the Nash equilibrium search term of the target intelligent agent is:
[0123]
[0124] wherein η ij is the Nash equilibrium search term of the jth target intelligent agent in the ith target alliance, is the update rate of η ij , κ is the Nash equilibrium search rate control parameter, κ > 0, g ijj is the jth element of the average gradient estimation term g ij of the jth target intelligent agent in the ith target alliance.
[0125] S8: controlling intelligent agents in the unmanned swarm system according to the Nash equilibrium search terms of all target intelligent agents in all target alliances, specifically comprising:
[0126] According to the multi-alliance game intelligent agent decision state dynamics model, a first gain matrix and a second gain matrix are constructed; the calculation formula of the first gain matrix and the second gain matrix is:
[0127]
[0128] wherein K ij1 is the first gain matrix of the jth target intelligent agent in the ith target alliance, K ij2 is the second gain matrix of the jth target intelligent agent in the ith target alliance, I q is a q x q dimensional unit matrix with diagonal elements being 1 and other elements being 0.
[0129] According to the Nash equilibrium search terms, the first gain matrix and the second gain matrix of all target intelligent agents in all target alliances, the control input of all target intelligent agents in all target alliances is obtained; the formula of the control input of the target intelligent agent is:
[0130]
[0131] wherein u' ij is the control input of the jth target intelligent agent in the ith target alliance.
[0132] Controlling the agents in the swarm system according to the control inputs of all target agents in all target coalitions.
[0133] The present application is directed to a multi-coalition game scenario where cooperation and game coexist, and proposes an agent control method for multi-coalition game. By introducing a preset time gain operator (a first gain matrix and a second gain matrix), the algorithm converges in a preset time. By designing and executing an estimated vector of left zero eigenvectors, the influence brought by the directed unbalanced communication topology graph is solved. Then, based on the average gradient estimation term, the limitation that the agent can only obtain its own cost function and needs to optimize the entire coalition cost function is solved, and a Nash equilibrium search term is designed in combination with the estimated vector. Finally, a Nash equilibrium search term output adjustment control algorithm is designed, which drives the agent (a high-order heterogeneous general linear agent) to converge to a Nash equilibrium solution in a preset time based on the Nash equilibrium search term and its integral designed in the previous step.
[0134] Consider a connectivity maintenance game involving 3 coalitions, i.e., coalition number i∈{1,2,3}. There are 3 individuals in coalition 1, 4 individuals in coalition 2, and 4 individuals in coalition 3, i.e., m1=3, m2=4, m3=4, and their communication topology relationship is as shown in Figure 2
[0135] During the game process, the decision variable of each individual is its position information, i.e., y ij =[y xij ,y yij ] T , where y xij represents the horizontal coordinate of the jth agent in the ith coalition, and y yij represents the vertical coordinate of the jth agent in the ith coalition.
[0136] Consider a multi-coalition game scenario where each individual participates in a sensor connectivity maintenance game. The cost function f ij of each individual is composed of the intra-coalition cost function and the inter-coalition cost function , and the expression is as follows:
[0137]
[0138] , where the expression of the intra-coalition cost function is:
[0139]
[0140] d1=[0,0] T m, d2=[5,5] T m, d3=[-5,-5] T m
[0141] ||·|| denotes the norm operator, where m is in meters, and d i It is the bias vector.
[0142] The inter-alliance cost function expression is:
[0143]
[0144] Using the method of this application, the trajectory diagrams of the decision variables of each agent and the estimation error curve of the left zero eigenvector of the pre-set time directed graph are obtained, as shown in the figure. Figures 3-4 As shown, Figure 3 This indicates that by applying the algorithm, the agent searches for its Nash equilibrium point from its initial position, demonstrating that the agent can find the Nash equilibrium solution within a preset time. Figure 4 The parameter in the upper right corner represents the magnitude (norm) of the error between the estimated vector and the true value of the left zero eigenvector. According to Figure 4 It can be seen that the estimation error of each agent for the left zero eigenvector is basically zero after 0.5S, that is, this application can effectively eliminate the influence of the directed unbalanced communication topology graph on the communication between agents.
[0145] Based on the same inventive concept, this application also provides an agent control device for multi-alliance games to implement the agent control method for multi-alliance games described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the agent control device for multi-alliance games provided below can be found in the limitations of the agent control method for multi-alliance games described above, and will not be repeated here.
[0146] In one exemplary embodiment, such as Figure 5 As shown, a smart agent control device for multi-alliance games is provided, comprising:
[0147] The alliance partitioning module is used to divide the agents in the unmanned swarm system into multiple alliances; each alliance includes multiple agents.
[0148] The directed unbalanced communication topology graph establishment module is used to establish a directed unbalanced communication topology graph of the unmanned swarm system based on the actual communication situation between various agents in the unmanned swarm system.
[0149] The model building module is used to build a dynamic model of the decision-making states of agents in a multi-alliance game based on the agents in all alliances.
[0150] The target alliance determination module is used to determine the target alliance based on the dynamic model of decision state of multi-alliance game agents; the target alliance is an alliance with agents that meet the triggering conditions; the agents in the target alliance that meet the triggering conditions are the target agents.
[0151] an estimation vector determination module configured to determine, according to the directed unbalanced communication topology graph, a Laplacian matrix of a target coalition and an estimation vector of a target agent in the target coalition; the estimation vector converging to a left null eigenvector of the Laplacian matrix of the target coalition to which the target agent belongs within a first preset time.
[0152] an average gradient estimation term determination module configured to determine, according to the estimation vector of the target agent, an average gradient estimation term of the target agent; the average gradient estimation term converging to an average estimation value of an overall cost function of the target coalition to which the target agent belongs within a second preset time.
[0153] a Nash equilibrium search term determination module configured to determine, according to the average gradient estimation term of the target agent, a Nash equilibrium search term of the target agent; the Nash equilibrium search term converging to a Nash equilibrium solution of the target agent within the second preset time; the Nash equilibrium solution being a decision state vector that minimizes the overall cost function of the target coalition to which the target agent belongs.
[0154] a control module configured to control the agents in the unmanned swarm system according to the Nash equilibrium search terms of all the target agents in all the target coalitions.
[0155] The main feature of the present application is to design and utilize the estimation of the left null eigenvector to eliminate the interference caused by the unbalanced directed graph, and to complete the preset time multi-coalition game Nash equilibrium search of high-order heterogeneous agents by designing the average gradient estimation term, the Nash equilibrium search term and the output adjustment term that converge within a preset time.
[0156] The present application focuses on the multi-coalition game scenario of intra-coalition cooperation and inter-coalition game under the unbalanced directed communication topology graph, and provides a scheme for searching to the Nash equilibrium solution within any preset time for the individual agents participating therein. By applying the present application, the efficiency of converging all agent states to the Nash equilibrium solution in multi-coalition game can be effectively improved, and the agents participating in the multi-coalition game can integrate the game-related information and converge to the Nash equilibrium solution point of connectivity preservation within a preset time in multi-coalition game scenarios such as multi-group connectivity preservation game, thereby improving the search rate of the agents to the Nash equilibrium solution point of connectivity preservation.
[0157] By using the method, high-order heterogeneous general linear agents can be controlled to search for the Nash equilibrium solution of multi-coalition game by interacting in the unbalanced directed graph only knowing their own cost function. The method can be applied to various scenarios such as multi-cluster unmanned system connectivity preservation and multi-coalition congestion game, to control the agents to search for the Nash equilibrium solution within a preset time.
[0158] In an example embodiment, a computer device is provided, which can be a server or a terminal, and an internal structure diagram thereof can be as shown in FIG. 1. Figure 6 The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store agent control data of a multi-alliance game. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement an agent control method for a multi-alliance game.
[0159] Those skilled in the art can understand that Figure 6 the structure shown in FIG. 1 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an example embodiment, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in each of the method embodiments described above.
[0160] In an example embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in each of the method embodiments described above.
[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0162] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0163] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0164] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.
[0165] The principles and implementation modes of the present application are described by applying specific examples herein. The above description of the embodiments is only used to help understand the method and its core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range can be changed. In conclusion, the content of the present application should not be understood as a limitation.
Claims
1. A method for controlling intelligent agents in multi-alliance game theory, characterized in that, The multi-alliance game intelligent agent control method comprises the following steps: Intelligent agents in the unmanned cluster system are divided into multiple alliances; each alliance comprises multiple intelligent agents; A directed unbalanced communication topology graph of the unmanned cluster system is established based on actual communication conditions between the intelligent agents in the unmanned cluster system; A multi-alliance game intelligent agent decision state dynamics model is established according to the intelligent agents in all alliances; A target alliance is determined according to the multi-alliance game intelligent agent decision state dynamics model; the target alliance is an alliance in which there is an intelligent agent satisfying a trigger condition; the intelligent agent satisfying the trigger condition in the target alliance is a target intelligent agent; A Laplace matrix of the target alliance and an estimation vector of the target intelligent agent in the target alliance are determined according to the directed unbalanced communication topology graph; the estimation vector converges to a left zero eigenvector of the Laplace matrix of the target alliance to which the target intelligent agent belongs at a first preset time; An average gradient estimation item of the target intelligent agent is determined according to the estimation vector of the target intelligent agent, comprising: a second preset time gain operator is established; an expression of the second preset time gain operator is: ; wherein, is a second preset time gain operator, is a first preset time gain operator for a preset time of Initialize the first element of the auxiliary variable of the i-th target agent in the i-th target coalition to 0, the i-th element of the first element of the auxiliary variable of the i-th target agent in the i-th target coalition to the average gradient estimate term, and the i-th element of the average gradient estimate term to 0, where N is the number of coalitions participating in the game, ; An average gradient estimation item of the target intelligent agent is determined according to the estimation vector of the target intelligent agent; an expression of the average gradient estimation item of the target intelligent agent is: ; in, , For the first The first of the target alliances The cost function of a target agent This represents the set of decision state vectors of all target agents in all target alliances. , , This represents the set of Nash equilibrium search terms for all target agents in all target coalitions. Indicates the first The set of Nash equilibrium search terms for all target agents in a target alliance. The number of target alliances participating in the game. Indicates the first The first of the target alliances Nash equilibrium search terms for a target intelligent agent express 3D real vector, for update rate, The estimated vector of the j-th target agent in the i-th target coalition. The One element, It is the second preset time gain operator. The average gradient estimate term for the j-th target agent in the i-th target coalition. The kth element; the average gradient estimation term converges to the average estimate of the overall cost function of the target alliance to which the target agent belongs at the second preset time; A Nash equilibrium search item of the target intelligent agent is determined according to the average gradient estimation item of the target intelligent agent; the Nash equilibrium search item converges to a Nash equilibrium solution of the target intelligent agent at the second preset time; the Nash equilibrium solution is a decision state vector that minimizes an overall cost function of the target alliance to which the target intelligent agent belongs; Intelligent agents in the unmanned cluster system are controlled according to the Nash equilibrium search items of all target intelligent agents in all target alliances. 2.The multi-alliance game agent control method of claim 1, wherein, An expression of the multi-alliance game intelligent agent decision state dynamics model is: in, Indicates the first The first in the alliance The decision state vector of each agent. , express 3D real vector, Indicates the first The first in the alliance Control input for an intelligent agent , express 3D real vector, Indicates the first The first in the alliance The internal state of an agent , express 3D real vector, Indicates the first The first in the alliance The update rate of the internal state of an agent. , Indicates the first The first in the alliance The system dynamics matrix of an agent, Indicates the first The first in the alliance The system input matrix of each agent, Indicates the first The first in the alliance The decision matrix of an agent. 3.The multi-alliance game agent control method of claim 2, wherein, A directed unbalanced communication topology graph of the unmanned cluster system is established based on actual communication conditions between the intelligent agents in the unmanned cluster system, specifically comprising: Based on the actual communication between various agents in the unmanned swarm system, establish the first Communication topology within each alliance ;in, Indicates participation in the The set of all intelligent agents within a coalition, , Indicates the first The number of agents within a consortium, Indicates the first The first intelligent agent within the alliance. Indicates the first The first in the league One intelligent agent; Indicates the first A set of communication edges between intelligent agents within a consortium. , Indicates from the first The first in the alliance The agent points to the first The first in the alliance A communication edge for an intelligent agent Indicates the first The first in the league An intelligent agent. Indicates the first The first in the league One intelligent agent; According to the first communication topology within the first weight adjacency matrix of the first weight adjacency matrix of the first element value in the i th row and the j th column of the weight adjacency matrix of the first According to the weight adjacency matrix of the first alliance, the in-degree of the first agent in the first alliance and the out-degree of the first agent in the first alliance are determined; the in-degree of the first agent in the first alliance , represents the element value of the first row and the first column in the weight adjacency matrix of the first alliance, and the out-degree of the first agent in the first alliance ; A directed unbalanced communication topology graph of the swarm of unmanned vehicles is constructed from the in-degrees of all agents in the coalition and the out-degrees of all agents. A directed unbalanced communication topology graph of the swarm of unmanned vehicles is constructed from the in-degrees of all agents in the coalition and the out-degrees of all agents.
4. The multi-alliance game agent control method according to claim 3, characterized in that, A Laplace matrix of the target alliance and an estimation vector of the target intelligent agent in the target alliance are determined according to the directed unbalanced communication topology graph, specifically comprising: The Laplace matrix of the target alliance is determined; an expression of the Laplace matrix is: ; wherein, is the Laplacian matrix of the ith target coalition, is a diagonal matrix constructed from the out-degrees of all agents within the ith target coalition, is the weight adjacency matrix of the coalition, the ith coalition. A first preset time gain operator is established; an expression of the first preset time gain operator is: ; wherein, , is a first preset time gain operator gain rate control parameter, is a second preset time, is a preset time is a first preset time gain operator, t is time; The estimation vector of the target intelligent agent in the target alliance is determined according to the first preset time gain operator and the directed unbalanced communication topology graph; an expression of the estimation vector is: ; wherein, is an estimated vector of the jth target agent in the ith target coalition, is an update rate of the estimated vector of the jth target agent in the ith target coalition, is a preset time as a first preset time gain operator, is a first preset time, .
5. The multi-alliance game agent control method according to claim 4, characterized in that, A Nash equilibrium search item of the target intelligent agent is determined according to the average gradient estimation item of the target intelligent agent, specifically comprising: The Nash equilibrium search item of the target intelligent agent is initialized as a zero vector; The Nash equilibrium search item of the target intelligent agent is determined according to the average gradient estimation item of the target intelligent agent; an expression of the Nash equilibrium search item of the target intelligent agent is: ; wherein, is the Nash equilibrium search term for the jth target agent in the ith target coalition, is the update rate of , is the Nash equilibrium search rate control parameter, , is the jth element of the average gradient estimate term for the jth target agent in the ith target coalition .
6. The multi-alliance game agent control method according to claim 5, characterized in that, Intelligent agents in the unmanned cluster system are controlled according to the Nash equilibrium search items of all target intelligent agents in all target alliances, specifically comprising: constructing a first gain matrix and a second gain matrix according to the multi-alliance game intelligent agent decision state dynamics model; the calculation formulas of the first gain matrix and the second gain matrix are respectively: ; wherein, is a first gain matrix of the jth target agent in the ith target coalition, is a second gain matrix of the jth target agent in the ith target coalition, is is an identity matrix of dimension n x n with diagonal elements equal to 1 and other elements equal to 0. obtaining control inputs of all target intelligent agents in all target alliances according to the Nash equilibrium search items of all target intelligent agents in all target alliances, the first gain matrix and the second gain matrix; the formula of the control inputs of the target intelligent agents is: ; wherein, is the control input of the jth target agent in the ith target coalition. controlling the intelligent agents in the unmanned swarm system according to the control inputs of all target intelligent agents in all target alliances.
7. An agent control apparatus for a multi-alliance game, characterized by comprising: The intelligent agent control device for the multi-alliance game is used to implement the intelligent agent control method for the multi-alliance game in any one of claims 1-6. The intelligent agent control device for the multi-alliance game comprises: a coalition division module, which is used to divide the intelligent agents in the unmanned swarm system into multiple alliances; each of the alliances comprises multiple intelligent agents; a directed unbalanced communication topology graph establishing module, which is used to establish a directed unbalanced communication topology graph of the unmanned swarm system based on actual communication conditions between the intelligent agents in the unmanned swarm system; a model establishing module, which is used to establish a multi-alliance game intelligent agent decision state dynamics model according to the intelligent agents in all alliances; a target alliance determining module, which is used to determine target alliances according to the multi-alliance game intelligent agent decision state dynamics model; the target alliances are alliances in which there are intelligent agents satisfying trigger conditions; the intelligent agents satisfying the trigger conditions in the target alliances are target intelligent agents; an estimation vector determining module, which is used to determine a Laplace matrix of the target alliances and estimation vectors of the target intelligent agents in the target alliances according to the directed unbalanced communication topology graph; the estimation vectors converge to left zero eigenvectors of the Laplace matrix of the target alliance to which the target intelligent agents belong at a first preset time; an average gradient estimation item determining module, which is used to determine average gradient estimation items of the target intelligent agents according to the estimation vectors of the target intelligent agents; the average gradient estimation items converge to average estimation values of the overall cost function of the target alliance to which the target intelligent agents belong at a second preset time; a Nash equilibrium search item determining module, which is used to determine Nash equilibrium search items of the target intelligent agents according to the average gradient estimation items of the target intelligent agents; the Nash equilibrium search items converge to Nash equilibrium solutions of the target intelligent agents at the second preset time; the Nash equilibrium solutions are decision state vectors that minimize the overall cost function of the target alliance to which the target intelligent agents belong; a control module, which is used to control the intelligent agents in the unmanned swarm system according to the Nash equilibrium search items of all target intelligent agents in all target alliances.
8. A computer device comprising: A memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the intelligent agent control method for the multi-alliance game in any one of claims 1-6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the intelligent agent control method for the multi-alliance game in any one of claims 1-6.
Citation Information
Patent Citations
Fault-tolerant control method for multi-agent system in non-cooperative game
CN116009395A
Unknown environment and unknown dynamics-oriented unmanned aerial vehicle cluster security upheaval control method
CN118884998A