Distributed optimal control method for nonlinear multi-robot cluster system under graph game framework
By constructing the pilot-following auxiliary variables and time-varying cost function under the graph game framework, the pilot-following distributed optimal controller is solved, and the distributed graph game control method for non-linear multi-robot cluster system in the existing technology is not applicable to the weakest connected network structure, achieving the optimal control effect of global Nash equilibrium.
Patent Information
- Application Number
- CN202510244442.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-10
AI Technical Summary
The existing distributed graph game control method for nonlinear multi-robot cluster systems cannot be applied to network communication structures with the weakest connectivity characteristics.
A distributed optimal control method for nonlinear multi-motion body cluster system under the graph game framework is proposed. By constructing a pilot-follow auxiliary variable, a new graph game cost function with time-varying characteristics is established, and a pilot-follow distributed optimal controller is designed to ensure that all follower robots can track the motion trajectory of the pilot robot, and analyzing the cost function can achieve global Nash equilibrium optimization.
This method can be applied to network communication structures with the weakest connectivity characteristics, realizes the optimal control effect of global Nash equilibrium and reduces the value of the cost function.
Smart Images

Figure CN120122583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a distributed control method for a non - linear multi - robot cluster system, in particular to a distributed optimal control method under a graph game framework, and relates to the technical field of cluster system control. Background Art
[0002] A multi - robot cluster is a network system composed of multiple intelligent individuals with autonomous motion capabilities. The robots complete operation tasks through relatively local information interaction. Compared with linear robot systems, non - linear robot systems are more practical and are usually used to describe the complex dynamic characteristics of systems such as underwater robots and unmanned ships. In some specific task scenarios, non - cooperative confrontation characteristics will be reflected among the robots, such as pursuit - evasion, intelligent transportation, etc. Graph game theory is an effective means to handle these non - cooperative tasks.
[0003] In the prior art regarding multi - robot clusters, some solutions have been provided. For example, the existing patent document with the document number CN117518833B (authorized) discloses an improved high - order multi - agent cluster distributed non - cooperative game method and system, which relates to the technical field of cluster system control. It includes: constructing a consistency auxiliary variable for a high - order multi - agent cluster system, establishing a cost function under a non - cooperative game framework, designing a distributed consistency controller to ensure that the states of all agents reach a pre - set consistent final value, and analyzing that the cost function can achieve Nash equilibrium optimality; finally, using the distributed consistency controller to perform motion control on the high - order multi - agent cluster system. This invention can be applied to a network communication structure with the weakest connectivity that only contains one directed tree and ensures that the cost function achieves Nash equilibrium. For example, the existing patent document with the document number CN117762151B (authorized) discloses a distributed shape control method and agent for an agent cluster without labels. In this method, the size of the grid unit of the binary target image in physical space is determined according to the binary target image, the number of agents, and the expected distance between agents; negotiate with neighbor agents within its own perception range to update its own understanding of the acceleration and angular acceleration of the target shape, and then update its own understanding of the position, speed, direction, and angular velocity of the target shape based on its own understanding of the acceleration and angular acceleration of the target shape; convert the binary target image into a grayscale image, where the grayscale value transitions smoothly in the grayscale image; update its own expected speed according to the resultant force of the speed alignment force, internal collision avoidance force, stability maintenance force, and formation control force. It is possible to achieve formation shape control without setting a unique label for the agent and without assigning an accurate position or trajectory to the agent. However, there are few graph game control methods proposed for non - linear multi - robot systems.
[0004] The literature "A novel distributed optimal adaptive control algorithm for nonlinear multi-agent differential graphical games, IEEE / CAA Journal of Automatica Sinica, 2018, 5(1): 331-341" proposed a distributed graph game control method for nonlinear multi-robot systems. This method uses neural networks to estimate the optimal solution and achieve global Nash equilibrium. However, the technical problem of the method described in this literature is that it has strong restrictions on the network communication structure of multi-robot clusters and cannot be applied to network communication structures with the weakest connectivity characteristics. Summary of the Invention
[0005] To overcome the deficiency that the existing distributed graph game control method for nonlinear multi-robot cluster systems cannot be applied to network communication structures with the weakest connectivity characteristics, the present invention proposes a distributed optimal control method for nonlinear multi-agent cluster systems under a new graph game framework. This method constructs leader-follower auxiliary variables for multi-robot cluster systems with Lipschitz nonlinear characteristics, establishes a new graph game cost function with time-varying characteristics, designs a leader-follower distributed optimal controller to ensure that all follower robots can track the motion trajectory of the leader robot, and analyzes that the cost function can achieve the global Nash equilibrium optimum. The method provided by the present invention can be applied to network communication structures with the weakest connectivity characteristics.
[0006] The technical solution adopted by the present invention to solve its technical problems: A distributed optimal control method for nonlinear multi-robot cluster systems under a graph game framework, which is characterized by including the following steps:
[0007] Step 1: Construct leader-follower auxiliary variables for multi-robot cluster systems with Lipschitz nonlinear characteristics, establish a new graph game cost function with time-varying characteristics, and design a leader-follower distributed optimal controller;
[0008] Step 2: For the controller designed in Step 1, analyze that all follower robots can track the motion trajectory of the leader robot, and analyze that the cost function can achieve the global Nash equilibrium optimum.
[0009] Step 1. Construct leader-follower auxiliary variables for multi-robot cluster systems with Lipschitz nonlinear characteristics, establish a new graph game cost function with time-varying characteristics, and design a leader-follower distributed optimal controller. The specific implementation process is as follows:
[0010] First, give the following multi-robot cluster system model with Lipschitz nonlinear characteristics:
[0011]
[0012] where \(i = 0\) represents the leader robot, and \(i = 1,\cdots,N\) represent the follower robots. represents the \(n\)-dimensional state variable of the robot, and \(u\) i represents the control input, and \(f(x\) i ) is a nonlinear term satisfying the Lipschitz condition, and \(a\) 1 ,\(\cdots,a\) n represent the system linear parameters; construct the leader-follower auxiliary variables as follows:
[0013]
[0014] where the parameters \(\theta\) 1 ,\(\theta\) 2 ,\(\cdots,\theta\) n-1 are selected such that all the roots of the following equation have negative real parts:
[0015] p n-1 +\(\theta\) n-1 p n-2 +\(\cdots+\theta\) 2 p+\(\theta\) 1 = 0. (3)
[0016] Meanwhile, assume that the Lipschitz condition \(\vert f(x\) i ) - f(x\) j )\vert\leq\gamma\) i \vert y\) i - y\) j \vert holds, where \(\gamma\) i > 0 is the Lipschitz parameter;
[0017] Design the following controller:
[0018]
[0019] where is used to compensate for the linear parameters in the system model, is used to compensate for the local parameters in the auxiliary variables, is the distributed control protocol used to achieve leader-follower consensus and global Nash equilibrium optimality;
[0020] Establish the cost function under the graph game framework as follows:
[0021]
[0022] where represents the distributed control protocol of the neighbors of robot \(i\), and \(e\) ijare the communication parameters between robot i and its neighbors, R i and R j are any positive numbers, δ ij and Q ij The expressions of are as follows:
[0023]
[0024] In the formula,
[0025]
[0026] In addition, c should also satisfy the following conditions:
[0027]
[0028] In the formula, represents the largest singular value of, represents the smallest singular value of, represents γ 1 P 1 ,..., γ N P N the maximum value in, max{g i +h i} i∈{1,...,N} represents g 1 +h 1 ,..., g N +h N the maximum value in, is the Laplacian matrix corresponding to the network communication structure, G = diag{g 1 ,..., g N} is a positive definite diagonal matrix;
[0029] Select the following Hamiltonian function according to the cost function (5):
[0030]
[0031] In the formula, V i represents the value function corresponding to formula (5), represents the partial derivative of the value function with respect to δ i ; Using the extreme value condition the optimal distributed control protocol can be obtained as:
[0032]
[0033] Let the value function be Therefore, it can be obtained:
[0034]
[0035] Step 2. For the controller designed in Step 1, analyze that all follower robots can track the motion trajectory of the leader robot, and analyze that the cost function can achieve the global Nash equilibrium optimum. The specific implementation process is as follows:
[0036] The robot tracking error system is as follows:
[0037]
[0038] Substituting Equation (4) and Equation (11) into Equation (12), we can get:
[0039]
[0040] Therefore, the closed-loop system expression is:
[0041]
[0042] In the formula,
[0043]
[0044] Select the following Lyapunov function:
[0045]
[0046] Taking the derivative of the Lyapunov function, we have:
[0047]
[0048] Therefore, δ i →0, that is, y i →y 0 ;
[0049] When y i = y 0 ,we have:
[0050]
[0051] In the formula, x 0,1 and its derivatives of all orders are known quantities, while x i,1 is the variable of the differential equation to be solved; let where is the particular solution of the differential equation, is the general solution of the following homogeneous differential equation:
[0052]
[0053] Since the parameters θ 1 , θ 2 ,..., θ n-1The selection principle is to ensure that all the roots of Equation (3) have negative real parts. Therefore, Furthermore, In addition, let the particular solution be:
[0054]
[0055] where m 0 , m 1 ,..., m n-1 are the parameters to be solved; according to Equations (18) and (20), m 0 = m 1 =... = m n-2 = 0 and Therefore, we have:
[0056]
[0057] In summary, it can be seen that x i → x 0 , that is, all follower robots can track the motion trajectory of the leader robot;
[0058] Substituting Equation (11) into Equation (9), we get:
[0059]
[0060] According to the expression of Q ij given in Equation (6), it can be known that:
[0061]
[0062] That is, the Hamilton-Jacobi equation holds.
[0063] After the Hamilton-Jacobi equation holds, based on the fact that the value function V i is twice continuously differentiable and its partial derivative is strictly monotonic, the necessary and sufficient condition for the non-cooperative game optimization problem of the swarm system to converge to a unique Nash equilibrium solution is According to it can be seen that when y i → y 0 , holds. Therefore, the global Nash equilibrium optimum can be achieved.
[0064] A distributed optimal control system for a nonlinear multi-robot swarm under a graph game framework, the system has program modules corresponding to the steps of the above technical solution, and executes the steps in the distributed optimal control method for a nonlinear multi-robot swarm system under the graph game framework when running.
[0065] A computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the distributed optimal control method for a nonlinear multi-robot cluster system under a graph game framework when called by a processor.
[0066] The beneficial effects of the present invention are as follows:
[0067] The method provided by the present invention has less restrictions on the network communication structure of the multi-robot cluster and is fully applicable to the network communication structure with the weakest connected characteristic. It overcomes the deficiency that the existing distributed graph game control method for a nonlinear multi-robot cluster system cannot be applied to the network communication structure with the weakest connected characteristic. The method of the present invention constructs leader-follower auxiliary variables for a multi-robot cluster system with Lipschitz nonlinear characteristics, establishes a new graph game cost function with time-varying characteristics, designs a leader-follower distributed optimal controller to ensure that all follower robots can track the motion trajectory of the leader robot, and analyzes that the cost function can achieve the global Nash equilibrium optimum. The method provided by the present invention can be applied to the network communication structure with the weakest connected characteristic.
[0068] In the existing distributed graph game methods for nonlinear multi-robot systems, it is required that the network communication structure is a strongly connected graph, while the method proposed by the present invention can be applied to the network communication structure with the weakest connected characteristic, that is, it only contains a directed tree, and ensures that the cost function achieves the Nash equilibrium. The non-Nash equilibrium control protocol can ensure that the cost function values are 325, 409, and 525 respectively, while the Nash equilibrium control protocol proposed by the present invention can ensure that the cost function value is only 294, which can clearly show the advantages of the technical solution of the present invention. Description of the Drawings
[0069] Figure 1 It is the network communication structure diagram of the multi-robot cluster in the embodiment of the present invention, where the leader robot is numbered 0, and the follower robots are numbered 1, 2, 3, 4, and 5;
[0070] Figure 2 It is the state norm curve of 6 robots in the embodiment of the present invention;
[0071] Figure 3 It is the cost function curve of the first robot in the embodiment of the present invention, where the solid line is the curve under the action of the Nash equilibrium control protocol proposed by the present invention, and the thin dotted line is the non-Nash equilibrium control protocol The curve under the action, and the bold dotted line is the non-Nash equilibrium control protocol The curve under the action, and the dash-dotted line is the non-Nash equilibrium control protocol The curve under the action. Detailed Implementation Manner
[0072] The following combines specific implementation manners and the appended Figures 1-3 to describe the present invention in detail.
[0073] Step 1: Construct a leader-follower auxiliary variable for a multi-robot swarm system with Lipschitz nonlinear characteristics, establish a new graph game cost function with time-varying characteristics, and design a leader-follower distributed optimal controller. First, give the following multi-robot swarm system model with Lipschitz nonlinear characteristics:
[0074]
[0075] In the formula, i = 0 represents the leader robot, and i = 1,..., N represent the follower robots. represents the n-dimensional state variable of the robot, and u i represents the control input quantity, and f(x i ) is a nonlinear term that satisfies the Lipschitz condition, and a 1 ,..., a n represent the system linear parameters. Construct the leader-follower auxiliary variable as follows:
[0076]
[0077] In the formula, the selection principle of the parameters θ 1 , θ 2 ,..., θ n-1 is to ensure that the roots of the following equation all have negative real parts:
[0078] p n-1 + θ n-1 p n-2 +... + θ 2 p + θ 1 = 0. (3)
[0079] At the same time, assume that the Lipschitz condition |f(x i ) - f(x j )| ≤ γ i |y i - y j | holds, and γ i > 0 is the Lipschitz parameter.
[0080] Design the following controller:
[0081]
[0082] In the formula, is used to compensate the linear parameters in the system model, is used to compensate the local parameters in the auxiliary variable, The part is a distributed control protocol for achieving leader-follower consensus and global Nash equilibrium optimality.
[0083] The cost function under the graph game framework is established as follows:
[0084]
[0085] In the formula, represents the distributed control protocol of the neighbors of robot i, and e ij is the communication parameter between robot i and its neighbors, and R i and R j are arbitrary positive numbers, and the expressions of δ ij and Q ij are as follows respectively:
[0086]
[0087] In the formula,
[0088]
[0089] In addition, c should also satisfy the following conditions:
[0090]
[0091] In the formula, represents the maximum singular value of, represents the minimum singular value of, represents γ 1 P 1 ,..., γ N P N the maximum value in, max{g i +h i}, i∈{1,...,N} represents the maximum value of g 1 +h 1 ,..., g N +h N in, is the Laplacian matrix corresponding to the network communication structure, and G = diag{g 1 ,..., g N} is a positive definite diagonal matrix.
[0092] Select the following Hamiltonian function according to the cost function (5):
[0093]
[0094] In the formula, V i represents the value function corresponding to formula (5), represents the value function with respect to δi The partial derivative. Using the extreme value condition The optimal distributed control protocol can be obtained as:
[0095]
[0096] Let the value function be Therefore, it can be obtained that:
[0097]
[0098] Step 2: For the controller designed in Step 1, analyze that all follower robots can track the motion trajectory of the leader robot, and analyze that the cost function can achieve the global Nash equilibrium optimum. The robot tracking error system is as follows:
[0099]
[0100] Substituting Equation (4) and Equation (11) into Equation (12), we can get:
[0101]
[0102] Therefore, the closed-loop system expression is:
[0103]
[0104] In the formula,
[0105]
[0106] Select the following Lyapunov function:
[0107]
[0108] Taking the derivative of the Lyapunov function, we have:
[0109]
[0110] Therefore, δ i → 0, that is, y i → y 0 .
[0111] When y i = y 0 , we have:
[0112]
[0113] In the formula, x 0,1 and its derivatives of all orders are known quantities, while x i,1 is the variable of the differential equation to be solved. Let where is the particular solution of the differential equation, is the general solution of the following homogeneous differential equation:
[0114]
[0115] Since the selection principle of the parameters θ 1 , θ 2 ,..., θ n-1 is to ensure that all the roots of Equation (3) have negative real parts, so Furthermore, there is In addition, let the particular solution be:
[0116]
[0117] where m 0 , m 1 ,..., m n-1 are the parameters to be solved. According to Equations (18) and (20), m 0 = m 1 =... = m n-2 = 0 and Therefore, there is:
[0118]
[0119] In summary, it can be seen that x i → x 0 , that is, all follower robots can track the motion trajectory of the leader robot.
[0120] Substituting Equation (11) into Equation (9) gives:
[0121]
[0122] According to the expression of Q ij given by Equation (6), it can be known that:
[0123]
[0124] That is, the Hamilton-Jacobi equation holds. In addition, since the value function V i is twice continuously differentiable and its partial derivative is strictly monotonic, so according to the optimization theory, the necessary and sufficient condition for the non-cooperative game optimization problem of the cluster system to converge to a unique Nash equilibrium solution is And according to it can be known that when y i → y 0 , holds, so the global Nash equilibrium optimum can be achieved.
[0125] The beneficial effects of the present invention are verified by the following embodiments (such asFigures 1-3 )
[0126] Suppose there are 6 robots in the cluster system, including 1 leader robot and 5 follower robots. The order of each robot is 4, and a 1 = 1, a 2 = 2, a 3 = 3, a 4 = 4 and f(x i ) = 0.8sin(8x i,1 + 12x i,2 + 6x i,3 + x i,4 ). Let r = 10 and q 1 = q 2 = q 3 = q 4 = 0.1, and take θ 1 = 8, θ 2 = 12, θ 3 = 6. In the existing distributed graph game control methods for nonlinear multi-robot clusters, it is required that the network communication structure is a strongly connected graph. However, the network communication structure diagram adopted in the embodiments of the present invention only contains a directed tree and has the weakest connectivity characteristic. Therefore, the method proposed by the present invention has stronger adaptability.
[0127] Select the initial 4th-order state values of the 6 robots as x i,1 (0) = 0.8(i - 10), x i,2 (0) = 0.6(i - 15), x i,3 (0) = 0.5(i + 15), x i,4 (0) = 0.3(i + 20), i ∈ {0, 1,..., 5}. In addition, select g 1 = 5, g 2 = 3.6, g 3 = 2.5, g 4 = 1.5, g 5 = 0.7. According to Equation (8), it can be calculated that c ≥ 61.26. Therefore, c = 62 is selected. The state norm curves of each robot can be obtained. From the simulation curves, it can be seen that all follower robots can track the leader robot. In addition, to further verify the advantages of the method proposed by the present invention, the cost function curves under multiple groups of distributed control protocols are also given. By comparing the curves, it can be seen that the 3 non-Nash equilibrium control protocols can ensure that the cost function values are 325, 409, and 525 respectively, while the Nash equilibrium control protocol proposed by the present invention can ensure that the cost function value is only 294. Therefore, the Nash equilibrium optimal control protocol designed by the present invention can minimize the cost function.
[0128] The content not detailed in the present invention (such as algebraic graph theory and matrix theory) belongs to the common knowledge in the field.
[0129] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention. Modifications, partial substitutions, and application expansions made to the description and drawings of the present invention, or the direct or indirect application of the present invention in other related technical fields, should all be included within the protection scope of the patent of the present invention.
Claims
1. A distributed optimal control method for a nonlinear multi-robot cluster system under a graph game framework, characterized in that: The method comprises the following steps: Step 1: Construct a pilot-follower auxiliary variable for a multi-robot swarm system with Lipschitz nonlinear characteristics, establish a new graph game cost function with time-varying characteristics, and design a pilot-follower distributed optimal controller; Step 2: For the controller designed in step 1, analyze whether all follower robots can track the motion trajectory of the leader robot, and analyze whether the cost function can achieve the optimal global Nash equilibrium.
2. The distributed optimal control method for a nonlinear multi-robot cluster system under a graph game framework according to claim 1 is characterized in that: Step 1: Construct a pilot-follower auxiliary variable for a multi-robot cluster system with Lipschitz nonlinear characteristics, establish a new graph game cost function with time-varying characteristics, and design a pilot-follower distributed optimal controller. The specific implementation process is as follows: First, the following multi-robot cluster system model with Lipschitz nonlinear characteristics is given: Where i=0 represents the leader robot, i=1,...,N represents the follower robot, Represents the robot's n-dimensional state variable, u i represents the control input, f(x i ) is a nonlinear term that satisfies the Lipschitz condition, a1,...,a n Represents the system linear parameters; constructs the pilot-follower auxiliary variables as follows: In the formula, the parameters θ1, θ2, ..., θ n-1 The selection principle is to ensure that the roots of the following equations all have negative real parts: p n-1 +θ n-1 p n-2 +...+θ2p+θ1=0. (3) At the same time, assuming the Lipschitz condition |f(x i )-f(x j )|≤γ i |y i -y j |Founded, γ i >0 is the Lipschitz parameter; Design the following controller: In the formula, Part of it is used to compensate the linear parameters in the system model. Partially used to compensate for local parameters in auxiliary variables, Part of it is a distributed control protocol for achieving leader-follower consistency and global Nash equilibrium optimization; The cost function under the graph game framework is established as follows: In the formula, represents the distributed control protocol of robot i’s neighbors, e ij is the communication parameter between robot i and its neighbors, R i and R j is any positive number, δ ij and Q ij The expressions are as follows: In the formula, In addition, c should also meet the following conditions: In the formula, express The maximum singular value of express The minimum singular value of Denote γ1P1,...,γ N P N The maximum value in max{g i +h i } i∈{1,...,N} represents g1+h1,...,g N +h N The maximum value in is the Laplace matrix corresponding to the network communication structure, G = diag{g1,...,g N } is a positive definite diagonal matrix; According to the cost function (5), the following Hamiltonian function is selected: Where V i The value function corresponding to expression (5) is: Denotes the value function of δ i Partial derivatives of; using extreme value conditions The optimal distributed control protocol can be obtained as: Let the value function be Therefore, we can obtain:
3. The distributed optimal control method for a nonlinear multi-robot cluster system under a graph game framework according to claim 2 is characterized in that: Step 2: For the controller designed in step 1, analyze whether all follower robots can track the motion trajectory of the leader robot, and analyze whether the cost function can achieve the optimal global Nash equilibrium. The specific implementation process is as follows: The robot tracking error system is as follows: Substituting equation (4) and equation (11) into equation (12), we can obtain: Therefore, the closed-loop system expression is: In the formula, Choose the following Lyapunov function: The derivative of the Lyapunov function is: Therefore, δ i →0, that is, y i →y0; When i =y0, we have: In the formula, x 0,1 and its derivatives are all known quantities, and x i,1 is the variable of the differential equation to be solved; let in is a particular solution to the differential equation, is the general solution of the following homogeneous differential equation: Since the parameters θ1,θ2,...,θ n-1 The selection principle is to ensure that the roots of formula (3) all have negative real parts, so Furthermore, In addition, the special solution for: Where m0, m1, ..., m n-1 is the parameter to be solved; according to equations (18) and (20), we can obtain m0=m1=...=m n-2 =0 and So we have: In summary, x i →x0, that is, all follower robots can track the motion trajectory of the leader robot; Substituting formula (11) into formula (9) yields: According to formula (6), Q ij The expression shows: That is, the Hamilton-Jacobi equation holds.
4. The distributed optimal control method for a nonlinear multi-robot cluster system under a graph game framework according to claim 3 is characterized in that: After the Hamilton-Jacobi equation is established, based on the value function V i is twice continuously differentiable, and its partial derivatives Strictly monotonic, the necessary and sufficient condition for the cluster system non-cooperative game optimization problem to converge to a unique Nash equilibrium solution is according to It can be seen that when y i →y0, Therefore, the global Nash equilibrium can be achieved.
5. A nonlinear multi-robot cluster distributed optimal control system under a graph game framework, characterized by: The system has a program module corresponding to the steps of any one of claims 1 to 4 above, and executes the steps of the distributed optimal control method of a nonlinear multi-robot cluster system under a graph game framework during operation.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of a distributed optimal control method for a nonlinear multi-robot cluster system under a graph game framework according to any one of claims 1 to 4 when called by a processor.
Citation Information
Patent Citations
An improved high-order multi-agent cluster distributed non-cooperative game method and system
CN117518833B
Distributed shape control method and agent for intelligent agent cluster without labeling
CN117762151B