A distributed nash equilibrium search method for parameter uncertainty heterogeneous multi-agent game
By constructing a strongly connected directed graph and a local position estimator, and combining it with an adaptive controller to estimate unknown parameters online, the problems of high-order state interaction and topological constraints in multi-agent systems are solved. This enables distributed Nash equilibrium search in complex environments, improving the robustness and adaptability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-19
AI Technical Summary
Existing multi-agent systems face challenges such as reliance on high-order state information interaction, heavy communication burden, topological constraints, and insufficient adaptability to disturbances when dealing with parameter uncertainties and heterogeneous high-order systems, making it difficult to achieve distributed Nash equilibrium search.
A strongly connected directed graph is constructed for information exchange, relying solely on location information. A local location estimator and an adaptive controller are designed, and unknown parameters are estimated online using the model reference adaptive control concept. Distributed Nash equilibrium search is performed using neighbor location information.
It reduces the requirements for sensor accuracy and communication bandwidth, improves the robustness and adaptability of the system, supports strongly connected directed graph topologies, is suitable for high-order heterogeneous systems in complex environments, and realizes lightweight and scalable distributed control.
Smart Images

Figure CN122239447A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control, specifically to a distributed Nash equilibrium search method for heterogeneous multi-agent games with uncertain parameters, which can be applied to fields such as drone formations, autonomous vehicle fleets, smart grids, and industrial production collaboration. Background Technology
[0002] With the development of intelligent control systems, multi-agent systems (MAS) have been widely applied in fields such as drone formation control, autonomous vehicle grouping, robot collaborative operations, and intelligent logistics scheduling. In these scenarios, each agent often possesses autonomous decision-making capabilities and needs to achieve global objective optimization or reach a certain strategic equilibrium under limited communication conditions. Therefore, the search for Nash equilibrium in non-cooperative games has become an important topic in the interdisciplinary research of control and game theory.
[0003] Several distributed Nash equilibrium search methods exist for heterogeneous high-order dynamics players; however, these methods suffer from the following significant shortcomings:
[0004] Interaction relying on higher-order state information: Most methods require frequent exchange of higher-order derivative information such as velocity and acceleration between agents, making the system highly sensitive to measurement errors, noise and sensor configuration, while greatly increasing the communication burden;
[0005] Graph structure limitations: Some algorithms require the communication graph to be undirected or balanced, making it difficult to directly apply to the more general unbalanced directed graphs;
[0006] Weak adaptability to disturbances and uncertainties: Most current methods focus on external disturbances, but lack effective mechanisms for handling parameter uncertainties (such as unknown system gain, unclear dynamic structure, etc.).
[0007] Limitations of applicable scenarios: Some algorithms in the literature are only applicable to low-order systems (such as second-order integrators only), and lack a unified modeling and control framework for high-order, heterogeneous, and complex structural systems.
[0008] Especially in real-world scenarios, such as robot control and automated production systems, the cost function of an agent often depends only on positional information, while higher-order information such as velocity does not directly participate in decision optimization. Therefore, how to design distributed control based solely on neighbor positional information, and achieve Nash equilibrium search for heterogeneous high-order systems with parameter uncertainties without relying on global information such as higher-order states and Laplace matrix eigenvalues, has become a core challenge that urgently needs to be overcome. Summary of the Invention
[0009] To address the aforementioned problems, this invention provides a distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty. This method can be applied to unmanned system swarms operating in strongly connected directed graphs (applicable to both balanced and unbalanced directed graphs), such as drone swarms and autonomous vehicle systems.
[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0011] The first aspect of this invention provides a distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty, comprising:
[0012] S1: Construct a strongly connected directed graph in a multi-agent system, and realize information communication and data exchange between nodes based on the strongly connected directed graph.
[0013] This step establishes a communication framework that reflects the real network topology by defining the direction of information transmission between nodes. Unlike traditional methods that rely on symmetric or undirected graphs, this scheme allows information flow in any direction (e.g., A can send data to B, but B cannot transmit in the opposite direction), thus more closely resembling the asymmetric communication scenarios in real-world unmanned systems caused by obstructions, sensor blind spots, or link failures. For example, during UAV formation flight, the lead aircraft may only be able to broadcast its position information to the aircraft behind it, but cannot receive feedback from them. This asymmetry is one of the core features supported by this invention. Therefore, this design significantly improves the system's robustness and adaptability under dynamic network changes.
[0014] Based on the established strongly connected directed graph communication structure, each agent can send its local location state to neighboring nodes through periodic broadcasts or event-triggered mechanisms, and receive location information from other neighboring nodes. This process does not require a centralized scheduler; all interactions are driven entirely by local rules. Because it relies only on location data (rather than higher-order quantities such as velocity and acceleration), it reduces the requirements for sensor accuracy and communication bandwidth, making it particularly suitable for lightweight perception needs in low-cost robotic platforms or vehicle systems. Furthermore, this mechanism ensures the scalability of the distributed control architecture—regardless of system size, each node only needs to maintain connections related to its neighborhood to complete collaborative tasks.
[0015] S2: Construct a heterogeneous high-order integrator system model for a multi-agent system considering N agents; in the heterogeneous high-order integrator system model, the cost function of each agent depends only on the position variables of all agents and has unknown constant parameters.
[0016] Specifically, the heterogeneous high-order integrator system model of the constructed multi-agent system is as follows:
[0017] Consider N agents, and the dynamic model of each agent i is as follows:
[0018]
[0019] in:
[0020]
[0021] here , , where r i ≥1 represents the system dimension, x i,1 The position of agent i is represented by u, and the rest are higher-order derivatives. i Indicates control input; Let θ represent a known Lipschitz continuous function. i It is an unknown constant parameter; A i and B i It is the system matrix;
[0022] The cost function of agent i is defined as:
[0023] J i =f i (x s,1 )=f i (x 1,1 x 2,1 , ..., x N,1 )
[0024] The cost function depends only on the position variables x of all agents. j,1 x s,1 =col(x 1,1 x 2,1 , ..., x N,1 The goal of the cost function is to find the Nash equilibrium point x of the system. s,1 * , making ▽ i f i (x i,1 * x -i,1 * )=0.
[0025] The core of step S2 lies in simplifying the originally complex multivariate optimization problem into a location-based local decision-making mechanism. Each agent i defines its individual cost function, which does not include velocity, acceleration, or time derivative terms, thus avoiding dependence on higher-order information. Simultaneously, each agent is modeled as a high-order integrator system (e.g., third order or higher) with unknown parameters. By using only location information for feedback design, the complexity of state observation is reduced, and the engineering practicality of the system is enhanced. This is particularly relevant in practical applications such as collaborative assembly of industrial robots, where only the final spatial positioning of the workpiece is typically a concern, rather than its transient behavior during motion.
[0026] S3: In a multi-agent system, each agent i independently runs a local position estimator i, a gradient calculation module, a virtual agent i, an error tracking module, and an adaptive controller i.
[0027] Local position estimator i is used to update agent i's estimate of the positions of all other agents using its own and its neighbors' position information. ;
[0028] The designed local location estimator is as follows:
[0029]
[0030] It is a vector representing the estimate of the position of virtual agent i for all other virtual agents; It is an adjustable parameter.
[0031] The gradient calculation module is used for... Calculate the position gradient .
[0032] Virtual agent i, used for position gradient based on input The estimates of the local location estimator i simulate the Nash equilibrium evolution process under ideal conditions;
[0033] The method for constructing virtual intelligent agent i is as follows:
[0034] Design a linear virtual reference system for each agent:
[0035]
[0036] in, It refers to the state of the virtual intelligent agent; δ is the estimate of the positions of agent i for all other agents; i >0 represents the gradient step size; a ij Let G be the adjacency weights of a strongly connected directed graph; initial conditions ;
[0037] The dynamic model of the virtual intelligent agent is expressed as follows:
[0038]
[0039] in, express The m-th derivative, in an r i In a linear system of order 1, the input is The output is m=0,…,r i -1.
[0040] The error tracking module is used to track the position x of agent i. i The position of the linear virtual agent i The error between;
[0041] Adaptive controller i, used for design-based parameter adaptive control law u i It outputs control over agent i and updates the estimates of unknown constant parameters. .
[0042] Design parameters and adaptive control law u i The method is as follows:
[0043] Define tracking error Construct sliding mode variables:
[0044]
[0045] definition , can be obtained
[0046]
[0047] The designed parameter adaptive control law u i for:
[0048]
[0049] Where, k i For gain; For θ i The online estimate satisfies ρ is the gain.
[0050] To address the failure of existing methods when faced with issues such as "interactions dependent on higher-order state information," "asymmetric topology," and "uncertainty in dynamic parameters," this invention designs an online parameter estimation algorithm based on the Model Reference Adaptive Control (MRAC) concept. Specifically, an auxiliary variable is introduced within each agent. This is used to estimate its own unknown parameters (such as system order, gain coefficient, etc.). Error signals are constructed... And update using gradient descent strategy This mechanism enables real-time identification of uncertain parameters. Even when initial parameters are inaccurate or drift over time, the system can automatically adjust the control law to maintain stability and convergence, significantly improving the algorithm's robustness and adaptability in real-world environments.
[0051] This invention designs a distributed Nash equilibrium search algorithm to simulate an ideal Nash equilibrium trajectory. It utilizes neighbor location information without requiring global information such as global state (e.g., velocity, acceleration) or Laplace matrix eigenvalues to achieve iterative updates of the optimal policy for each agent. This is the key innovation of the entire technical solution. Traditional methods often rely on Laplace matrix eigenvalues from graph analysis to determine system convergence or design gain coefficients, but such information is difficult to obtain in a distributed environment and is easily affected by topology changes. This invention abandons this dependence and instead adopts an iterative update rule based on local gradient projection. This algorithm only needs to interact with neighbor location information to complete the update, without involving higher-order state variables, truly realizing a "centralized, parameter-free" distributed control logic.
[0052] In each iteration, the algorithm of this invention modifies the policy using local location gradients or virtual agents, causing the policy of the entire population to converge towards Nash equilibrium while satisfying the stability constraints of higher-order systems. This ensures that the algorithm not only approximates the mathematically significant Nash equilibrium point but also maintains the internal dynamic stability of the system. During each update, local gradient information is used to guide the policy towards a better direction.
[0053] A second aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described above.
[0054] A third aspect of the present invention provides a computer-readable storage medium having a computer program / instruction stored thereon, characterized in that, when the computer program / instruction is executed by a processor, it implements the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described above.
[0055] A fourth aspect of the present invention provides a computer program product, including a computer program / instruction, characterized in that, when the computer program / instruction is executed by a processor, it implements the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described above.
[0056] This invention has outstanding substantive features and significant progress compared to the prior art, specifically:
[0057] 1) Based solely on location information interaction, it has low communication overhead and strong adaptability;
[0058] This solution breaks away from the existing technical approach of relying on high-order state (such as velocity and acceleration) interactions in high-order system game algorithms. The proposed local position estimator and adaptive controller rely only on the position information of neighbors, which effectively reduces the requirements for system measurement accuracy and network bandwidth, and significantly improves the actual deployment capability and robustness of the system. It is especially suitable for industrial field environments with limited resources or communication (such as robot systems, AGV systems, and drone swarms).
[0059] 2) Supports strongly connected directed graph topologies, providing stronger network adaptability;
[0060] This solution supports strongly connected directed graph structures, breaking through the limitations of many existing works that are limited to undirected or balanced directed graphs. It enables stable operation in distributed systems with heterogeneous communication directions and asymmetric information transmission, and has greater flexibility and universality.
[0061] 3) The controller has a lightweight structure and the algorithm can be implemented in a distributed manner;
[0062] The controller design only involves state feedback, gradient feedback, virtual agent projection updates, and standard adaptive update laws. All variables can be obtained locally or through neighbor interactions. The structure is simple, easy to implement in a distributed manner, does not rely on a central coordinator, and has good engineering feasibility.
[0063] 4) Applicable to heterogeneous high-order systems with parameter uncertainties
[0064] For scenarios where the dynamic models of intelligent agents differ and the parameters are not fully known in reality, this invention designs a tracking strategy based on the idea of model reference adaptive control. This strategy effectively compensates for the impact of parameter uncertainty on system stability and game convergence, and improves the algorithm's adaptability in complex environments. It is one of the few distributed Nash equilibrium search methods that can handle multiple characteristics of "high-order + heterogeneous + uncertain parameters".
[0065] 5) Applicable to multiple scenarios such as intelligent manufacturing, energy dispatching, and traffic coordination.
[0066] This invention has good versatility, and the proposed method is applicable to game problems in various multi-agent systems, especially in typical scenarios such as intelligent manufacturing, logistics collaboration, and unattended equipment group control, which have broad application prospects. Attached Figure Description
[0067] Figure 1This is a system block diagram of an intelligent agent with uncertain dynamics under this control scheme.
[0068] Figure 2 It is a directed graph describing the information interaction between six players.
[0069] Figure 3 It represents the positional error between a player and their virtual counterpart in a directed graph.
[0070] Figure 4 It is the trajectory of each player's position change in a directed graph. Detailed Implementation
[0071] The technical solution of the present invention will be further described in detail below through specific embodiments.
[0072] The terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion.
[0073] Example 1
[0074] like Figure 1 As shown, this embodiment provides a distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty, including:
[0075] S1: Construct the communication topology diagram;
[0076] In a multi-agent system, a strongly connected directed graph is constructed, and information communication and data exchange between nodes are realized based on this strongly connected directed graph.
[0077] In a non-cooperative game involving N agents, the set of participants is denoted as N = {1, 2, ..., N}; each agent i corresponds to a cost function f. i (x): N → The set x, consisting of an ordered array of all N real numbers, is defined as the set of action configurations; where, N Let x be an N-dimensional Euclidean space, where N is a positive integer, and the set x = [x1, x2, ..., xn]. N ] T x i ∈ N Represents the action strategy of the i-th participant; y = [y i1 ,y i2 ,…,y iN ] T ∈ Nm It is an estimate of the entire action set x, where y ijIt is an approximate estimate of the action of participant i to participant j; neighboring agents communicate with each other through a strongly connected directed graph, and agents i and j only exchange position information through the path between nodes.
[0078] If the constructed communication topology graph is a strongly connected directed graph G, then:
[0079] For any x -i ∈X -i $, cost function In x i The upper part is convex and continuously differentiable; and there exists a positive constant μ such that For all Established;
[0080] For each intelligent agent ,exist The above is globally Lipschitz continuous and has a constant. , making .
[0081] S2: System Modeling;
[0082] Construct a heterogeneous high-order integrator system model for a multi-agent system considering N agents; in the heterogeneous high-order integrator system model, the cost function of each agent depends only on the position variables of all agents and has unknown constant parameters.
[0083] Specifically, the heterogeneous high-order integrator system model of the constructed multi-agent system is as follows:
[0084] Consider N agents, and the dynamic model of each agent i is as follows:
[0085] (1)
[0086] in:
[0087]
[0088] here , , where r i ≥1 represents the system dimension, x i,1 The position of agent i is represented by u, and the rest are higher-order derivatives. i Indicates control input; Let θ represent a known Lipschitz continuous function. i It is an unknown constant parameter; A i and B i It is the system matrix;
[0089] The cost function of agent i is defined as:
[0090] J i =f i (x s,1 )=f i (x 1,1 x 2,1 , ..., x N,1 )
[0091] The cost function depends only on the position variables x of all agents. j,1 x s,1 =col(x 1,1 x 2,1 , ..., x N,1 The goal of the cost function is to find the Nash equilibrium point x of the system. s,1 * , making ▽ i f i (x i,1 * x -i,1 * )=0.
[0092] S3: System operation logic and interaction process;
[0093] In a multi-agent system, each agent i independently runs the following modules:
[0094] Local position estimator i is used to update agent i's estimate of the positions of all other agents using its own and its neighbors' position information. ;
[0095] The gradient calculation module is used for... Calculate the position gradient ;
[0096] Virtual agent i, used for position gradient based on input The estimates of the local location estimator i simulate the Nash equilibrium evolution process under ideal conditions;
[0097] The error tracking module is used to track the position x of agent i. i The position of the linear virtual agent i The error between;
[0098] Adaptive controller i, used for design-based parameter adaptive control law u i It outputs control over agent i and updates the estimates of unknown constant parameters. .
[0099] System operation logic and interaction process: Each agent i first updates its own location information and that of its neighbors. Then based on Calculate the position gradient Then drive the linear virtual reference system Finally, the linear virtual reference system is tracked, and the adaptive controller u is updated. i Simultaneously update the estimates of unknown parameters. All steps in the system's operational logic and interaction process involve local computation, and communication only involves position estimations, ensuring the algorithm's distributed and lightweight characteristics.
[0100] Local position estimator design:
[0101] Since the agent can only perceive neighbor information, a distributed state estimator is designed as follows to estimate the positions of all other players. It is a vector representing the estimate of the position of virtual agent i for all other virtual agents; The update rules are as follows
[0102]
[0103] It is an adjustable parameter.
[0104] Virtual intelligent agent construction:
[0105] A linear virtual reference system (i.e., virtual agent) is designed for each agent to simulate the ideal Nash equilibrium trajectory in a distributed structure.
[0106] To avoid directly dealing with uncertainty and high-order structures, a linear virtual reference system (linear agent) is first designed for each agent to simulate the Nash equilibrium evolution process under ideal conditions:
[0107] (2)
[0108] in, It refers to the state of the virtual intelligent agent; δ is the estimate of the positions of agent i for all other agents; i >0 represents the gradient step size; a ij Let G be the adjacency weights of a strongly connected directed graph; initial conditions ;
[0109] The dynamic model of the virtual intelligent agent is expressed as follows:
[0110]
[0111] in, express The m-th derivative, in an r i In a linear system of order 1, the input is The output is m=0,…,r i -1.
[0112] Adaptive controller design:
[0113] Define tracking error Construct sliding mode variables:
[0114]
[0115] definition , can be obtained
[0116] (3)
[0117] To make the actual state x i asymptotic tracking The designed parameter adaptive control law u i for:
[0118] (4)
[0119] Where, k i For gain; For θ i The online estimate satisfies ρ is the gain.
[0120] The entire control law only involves position information, sliding mode error, and estimated values, completely avoiding higher-order measurements such as velocity / acceleration.
[0121] The proof of the Nash equilibrium solution for players with uncertain parameters consists of two steps. The first step is to prove the solution in a distributed algorithm. Under the influence of this, the error between the uncertain player and the linear virtual player converges to zero, that is... The second step is to prove that the game among all virtual players achieves a Nash equilibrium.
[0122] first step:
[0123] control algorithm Substituting into the dynamic model (1), we can obtain
[0124]
[0125] in,
[0126] For agent i, consider Lyapunov candidate functions.
[0127]
[0128] right Taking the time derivative yields
[0129]
[0130] This indicates ,and .
[0131] Pair Solving for the given information yields the following results:
[0132] (5)
[0133] because It is a Hurwitz matrix, therefore α exists. i1 >0 and α i2 >0 makes .
[0134] From the formula get
[0135] (6)
[0136] Notice, According to the input-output stability lemma, we can obtain...
[0137] (7)
[0138] This means , .
[0139] because From equation (6), we can obtain From the formula get According to Barbalat's lemma, we can obtain .
[0140] Step 2: Prove the dynamics of each virtual player It converges to the Nash equilibrium solution.
[0141] First, for each player i, define a virtual variable.
[0142] (8)
[0143] The following dynamic model can be obtained.
[0144] (9)
[0145] If Considered as an independent node in the new graph, each node has the following dynamic model.
[0146] (10)
[0147] definition , , , l=1,…,r i . We can obtain
[0148] (11)
[0149] make ,in, Then, the following compact form is obtained.
[0150] (12)
[0151] Based on the above state transitions, the Nash equilibrium search problem in heterogeneous multiplayer games is reformulated as the stability problem of the corresponding system (11). Next, we will prove... and , l=1,…,r i .
[0152] Consider the following Lyapunov candidate functions
[0153] V=V z +V y
[0154] in, ,
[0155]
[0156]
[0157] For V z,i Taking the derivative, we get
[0158]
[0159] set up , and Δ=diag(δ1,…,δ N ). This yields
[0160] (13)
[0161] in, Notice .
[0162] because , ,have
[0163]
[0164] Since the cost function satisfies the conditions of smoothness and strong convexity, we can obtain
[0165]
[0166] in, .
[0167] Therefore, the above equation can be rewritten as
[0168] (14)
[0169] Substituting it into equation (13), we can obtain
[0170] (15)
[0171] For V y Taking the derivative, we get
[0172] (16)
[0173] Rearrange using Young's inequality and The intersection terms between them. Substituting equations (13) and (16) into V, we get
[0174]
[0175] Among them, η1, η2, η3, and η4 are positive constants.
[0176] To ensure Negative definite, choose γ i and δ i Make
[0177]
[0178] Specifically, let η1=1, η3=η4=1 / 2 and η2= γ i and δ i Need to meet
[0179]
[0180] Since G is a strongly connected graph, d i ≥1.
[0181] choose For player i, when You can get The function is negative definite for all t. Since V is a positive definite function, the system is globally asymptotically stable (11), i.e. .
[0182] Therefore, in this embodiment, if the graph structure is strongly connected and the cost function satisfies the conditions of smoothness and strong convexity, then the position trajectories of all agents asymptotically converge to the Nash equilibrium point: , where x s,1 * This represents the Nash equilibrium point of the system. That is, all agent positions converge to the Nash equilibrium point, and their higher-order derivative states converge to zero.
[0183] Effect verification
[0184] To verify the effectiveness of the proposed distributed Nash equilibrium search method, consider a game system with N=6 players, whose interaction topology is as follows: Figure 2 As shown. This topology is a directed strongly connected graph, which satisfies the assumptions about network structure in this invention.
[0185] The objective function f for each player i i (x) is defined as follows, depending only on its own and some of its neighbors' location state variables x. j,1 :
[0186] f1(x) = x 1,1 2 +4x 1,1 +2x 1,1 x 2,1 +1
[0187] f2(x)=x 2,1 2 +2x 2,1 +x 2,1 x 3,1 +2
[0188] f3(x)=x 3,1 2 +6x 3,1 +2x 2,1 x 3,1 +3
[0189] f4(x) = x 4,1 2 +2x 4,1 +2x 4,1 x 2,1 +4
[0190] f5(x) = x 5,1 2 +2x 5,1 +x 5,1 x 2,1 +5
[0191] f6(x) = x 6,1 2 -6x 6,1 +6
[0192] Analytical calculations show that the Nash equilibrium solution of this game system is:
[0193] x s,1 * ={-3,1,-4,-2,-1.5,3}
[0194] In system modeling, players are considered to have different dynamic models, reflecting heterogeneity, and are divided into three categories:
[0195] The first category is first-order systems, suitable for some simplified control platforms (such as ground robots):
[0196]
[0197] The second category is second-order systems, typically represented by common autonomous vehicle dynamics:
[0198]
[0199] The third category is third-order systems, mainly used for modeling acceleration control of unmanned aerial vehicle (UAV) platforms.
[0200]
[0201] Where, x i,1 Represents the position variable, x i,2 and x i,3 These represent velocity and acceleration, respectively; u i To control the input. Function Uncertain parameter θ i =i.
[0202] The initial system settings are as follows:
[0203] x i,1 (0)=i, i∈{1,4}
[0204] x i,2 (0)=0, i∈{2,5}
[0205] x i,3 (0)=0, i∈{3,6}
[0206] (0)=x i,1 (0), (0)=0, (0)=0
[0207] y i (0)=[0, 0.1i, 0.2i, -0.2, -0.1, 0]T
[0208] The controller parameters are set as follows:
[0209] =5, δ=1, k i =1,ρ i =10 i
[0210] Figure 3 This demonstrates that the positional error between the player and their virtual player converges to 0. Figure 4 The simulation results show the trajectory of each player's position changing over time. From the simulation results, it can be observed that after t=35s, the position variable x of all players... i,1 (t) all converge to their respective Nash equilibrium values x. s,1 * This verifies the consistency and effectiveness of the algorithm of the present invention in heterogeneous high-order systems and strongly connected non-equilibrium graph structures.
[0211] Example 2
[0212] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in Embodiment 1.
[0213] Example 3
[0214] This embodiment provides a computer-readable storage medium storing a computer program / instruction thereon, characterized in that, when the computer program / instruction is executed by a processor, it implements the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in Embodiment 1.
[0215] Example 4
[0216] This embodiment provides a computer program product, including a computer program / instruction, characterized in that, when the computer program / instruction is executed by a processor, it implements the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in Embodiment 1.
[0217] Those skilled in the art will understand that embodiments of the present invention can be provided as methods or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0218] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.
Claims
1. A distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty, characterized in that, include: S1: Construct a strongly connected directed graph G in a multi-agent system, and realize information communication and data exchange between nodes based on the strongly connected directed graph G; S2: Construct a heterogeneous high-order integrator system model for a multi-agent system considering N agents; in the heterogeneous high-order integrator system model, the cost function of each agent depends only on the position variables of all agents and has unknown constant parameters; S3: In a multi-agent system, each agent i independently runs the following modules: Local position estimator i is used to update agent i's estimate of the positions of all other agents using its own and its neighbors' position information. ; The gradient calculation module is used for... Calculate the position gradient ; Virtual agent i, used for position gradient based on input The estimates of the local location estimator i simulate the Nash equilibrium evolution process under ideal conditions; The error tracking module is used to track the position x of agent i. i The position of the linear virtual agent i The error between; Adaptive controller i, used for design-based parameter adaptive control law u i It outputs control over agent i and updates the estimates of unknown constant parameters. .
2. The distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in claim 1, characterized in that, The heterogeneous high-order integrator system model of the multi-agent system constructed in step S2 is as follows: Consider N agents, and the dynamic model of each agent i is as follows: in: here , , where r i ≥1 represents the system dimension, x i,1 The position of agent i is represented by u, and the rest are higher-order derivatives. i Indicates control input; Let θ represent a known Lipschitz continuous function. i It is an unknown constant parameter; A i and B i It is the system matrix; The cost function of agent i is defined as: J i =f i (x s,1 )=f i (x 1,1 ,x 2,1 ,…,x N,1 ) The cost function depends only on the position variables x of all agents. j,1 x s,1 =col(x 1,1 x 2,1 , ..., x N,1 The goal of the cost function is to find the Nash equilibrium point x of the system. s,1 * , making ▽ i f i (x i,1 * x -i,1 * )=0.
3. The distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in claim 1, characterized in that, The designed local location estimator is as follows: It is a vector representing the estimate of the position of virtual agent i for all other virtual agents; It is an adjustable parameter.
4. The distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in claim 1, characterized in that, The method for constructing virtual intelligent agent i is as follows: Design a linear virtual reference system for each agent: in, It refers to the state of the virtual intelligent agent; δ is the estimate of the positions of agent i for all other agents; i >0 represents the gradient step size; a ij Let G be the adjacency weights of a strongly connected directed graph; initial conditions ; The dynamic model of the virtual intelligent agent is expressed as follows: in, express The m-th derivative, in an r i In a linear system of order 1, the input is The output is m=0,…,r i -1.
5. The distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty according to claim 1, characterized in that, Design parameters and adaptive control law u i The method is as follows: Define tracking error Construct sliding mode variables: definition , can be obtained The designed parameter adaptive control law u i for: Where, k i For gain; For θ i The online estimate satisfies ρ is the gain.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in any one of claims 1 to 5.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the distributed Nash equilibrium search method for heterogeneous multi-agent games with parameter uncertainty as described in any one of claims 1 to 5.