A multi-agent differential game control method based on signal preserving matrix
The multi-agent differential game control method based on signal-holding matrix design solves the problem that traditional methods are only applicable to 1-relative order systems, and achieves near-optimal feedback control for 2-relative order systems, which is applicable to the swarm motion control of UAVs and spacecraft.
Patent Information
- Application Number
- CN202211627052.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-12-16
AI Technical Summary
Existing technologies are ineffective in solving differential game problems consisting of 2-relative order systems. Traditional methods are only applicable to 1-relative order systems and lose feedback signals in the construction of dynamic approximate solutions.
A multi-agent differential game control method is designed using a signal-preserving matrix. By constructing a diagonally partitioned algebraic matrix solution and an auxiliary system, an approximately optimal feedback control law is designed, which is applicable to 2-relative order systems.
It achieves an approximate optimal solution to the differential game problem of 2-relative order systems, expands the scope of application, and is applicable to the swarm motion control of UAVs or spacecraft, and has asymptotic stability.
Smart Images

Figure CN116414029B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of multi-agent swarm motion control, and particularly relates to a multi-agent differential game control method based on a signal preserving matrix. BACKGROUND
[0002] Game theory is an effective theory for dealing with multi-person decision-making problems. As a branch of game theory, differential game theory is mainly used for dealing with multi-person decision-making problems in dynamic systems, and has been widely applied in many aspects such as economic management, complex network defense and power system collaborative management in recent years.
[0003] Solving a differential game problem can be attributed to finding a set of equilibrium strategies that make the individual optimization indicators of all participants reach the best possible, which is called a Nash equilibrium strategy. The Nash equilibrium strategy of a differential game depends on solving a set of Hamilton-Jacobi-Isaacs partial differential equations (PDEs) coupled with the control strategy. Unfortunately, it is difficult to obtain the exact analytical solution of the equation set in practical applications, and an approximate solution of the coupled partial differential equation set is usually sought instead.
[0004] Mylvaganam T and Astolfi A et al. gave a dynamic construction method for an approximate solution of the HJI equation set based on the immersion and invariance theory, thereby effectively avoiding the solution of PDEs. This method constructs an exact solution of a class of inequalities by introducing an algebraic matrix solution and an auxiliary system, and regards the difference between the inequalities and the HJI equation to be solved as an error quantity. The auxiliary system is designed through the gradient direction of the error, so that the exact solution of the inequalities continuously approximates the exact solution of the HJI equation set, and it is proved that the differential game system has (local) asymptotic stability at the equilibrium point through the approximate optimal feedback control law. However, this method is difficult to apply to systems with a relative order of 2 when the introduced algebraic matrix solution is required to be in a diagonal block form. SUMMARY
[0005] The application aims to overcome the shortcomings of the prior art and provide a multi-agent differential game control method based on a signal preserving matrix.
[0006] To achieve the above object, the application adopts the following technical scheme:
[0007] A multi-agent differential game control method based on a signal preserving matrix comprises the following steps:
[0008] Step 1: According to the selected agent objects of the multi-agent system, the dynamic models of the agents are established.
[0009] According to the task requirements, the individual performance indicators of the agents are designed.
[0010] Step 2: Based on the agent dynamics model and individual performance indicators of each agent constructed in Step 1, establish a differential game model for the multi-agent system;
[0011] Based on the dynamic models of each agent established in Step 1, design the signal preservation matrix required for each agent;
[0012] Step 3: Construct the algebraic matrix solution for the diagonal blocks, build the auxiliary system, and design the dynamic equations of the auxiliary system with asymptotic stability at the equilibrium point;
[0013] Step 4: Based on the multi-agent system differential game model established in Step 2, the signal preservation matrix of each agent, and the algebraic matrix solution and auxiliary system constructed in Step 3, design the approximate optimal feedback control law for each agent.
[0014] Step 5: Input the near-optimal feedback control laws designed in Step 4 into each agent system and update the system state variables.
[0015] Furthermore, in step one, the number of intelligent agents is... , No. i A smart agent At any time state variables , positive integer The dimension of the state variables; the state variables include position coordinates. and velocity coordinates , positive integer ;
[0016] No. i At the initial moment, each intelligent agent... The state variables at time are , No. i A dynamic model of an intelligent agent in a given Cartesian inertial coordinate system has the following form:
[0017] (1)
[0018] in, For control quantities, positive integers The dimension of the control quantity has an acceleration dimension; Represents the rate of change of the agent's position. about The mapping relationship, Represents the rate of change of the agent's velocity about The mapping relationship; To describe the control quantity rate of change of agent velocity The mapping.
[0019] Further, in step one, the individual performance index of each agent is in the form of:
[0020] (2)
[0021] where the state variable of the multi-agent system is composed of the state variables of the agents, i.e. , , is a scalar index related to the state variable of the multi-agent system , and the matrix is the weight of the control input.
[0022] Further, in step two, the differential game model of the multi-agent system is integrated according to formula (1) into the following form:
[0023] (3).
[0024] Further, in step two, the signal preserving matrix corresponding to each agent designed according to the dynamic model of the agent is:
[0025] (4)
[0026] where , has no all-zero row, , , is a unit directional vector parallel to the state variable corresponding to the all-zero row.
[0027] Further, in step three, the algebraic matrix solution of each agent in the form of a diagonal block is:
[0028] (5)
[0029] Meanwhile, the auxiliary system corresponding to each agent is designed according to the negative gradient as:
[0030] (6)
[0031] where the matrix block represents the matrix block corresponding to the first row and the first column of the algebraic matrix solution of the agent. In the matrix block , is a scalar parameter, is a unit matrix of the same size as the matrix block , with respectively represent the diagonal elements of the matrix block to be designed , is a real number set; is an auxiliary system state variable with the same structure as the multi-agent system state variable .
[0032] Further, is a scalar function.
[0033] Further, in step four, the designed approximate optimal feedback control law of each agent is:
[0034] (7)
[0035] wherein the matrix value function satisfies .
[0036] Compared with the prior art, the present application has the following beneficial effects:
[0037] The application discloses a multi-agent differential game control method based on a signal preserving matrix, solves the problem of losing feedback signals in a traditional differential game system dynamic approximate solution construction method by introducing a signal preserving matrix, avoids the problem that the traditional differential game system dynamic approximate solution construction method is only applicable to a differential game system composed of 1-relative order subsystems, and obtains an approximate optimal feedback control law of a differential game system composed of 2-relative order subsystems through the signal preserving matrix. The application expands the applicable objects of the differential game system dynamic approximate solution construction method to a large range and has important application value. The application can dynamically obtain an approximate optimal solution of the differential game system and is applicable to cluster motion control of a second-order system such as an unmanned aerial vehicle or a spacecraft. BRIEF DESCRIPTION OF DRAWINGS
[0038] Fig. 1 is a flowchart of a multi-agent differential game control method based on a signal preserving matrix according to the application;
[0039] Fig. 2 is a diagram of a multi-agent cooperative control task in a specific embodiment;
[0040] Fig. 3 is a comparison diagram of a single-agent single-channel state response with a traditional method in a specific embodiment;
[0041] Fig. 4 is a state response curve of all agents in a specific embodiment;
[0042] Fig. 5 is a controlled movement curve of multi-agents in a specific embodiment under a LVLH system of a target to be surrounded;
[0043] Fig. 6 is a relative distance change curve between agents in a specific embodiment. DETAILED DESCRIPTION
[0044] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.
[0045] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0046] The present application avoids the conventional differential game system dynamic approximate solution construction method which is only applicable to solving the differential game problem composed of 1-relative order subsystems. The present application introduces a control law design method based on signal preserving matrix to obtain the approximate optimal feedback control law of the differential game problem composed of 2-relative order subsystems.
[0047] The present application will be described in further detail below with reference to the accompanying drawings:
[0048] Referring to Figure 1 , Figure 1 The flowchart of the present application is a multi-agent differential game control method based on signal preserving matrix, comprising the following steps:
[0049] Step 1: According to the selected agent object of the multi-agent system, the dynamic model of each agent is established:
[0050] (1)
[0051] Specifically, the agent , the agent The state variable of the agent at any time includes the position coordinates and the velocity coordinates , and the initial state of each agent is and , is the control input.
[0052] Step 2: According to the task requirements, design the individual performance index of each agent:
[0053] (2)
[0054] Specifically, the multi-agent system state variable should be composed of agent state variables, that is, , is the scalar index related to the multi-agent system state variable , the matrix is the weight of the control input.
[0055] Step 3: According to the agent dynamics model constructed in step 1 and the individual performance index of each agent designed in step 2, establish the differential game model of the multi-agent system:
[0056] (3)
[0057] Step 4: According to the agent dynamics model established in step 1, design the signal holding matrix required by each agent:
[0058] (4)
[0059] Specifically, so that there is no all-zero row, and the constant vector , , is the unit direction vector parallel to the state variable corresponding to the all-zero row.
[0060] Step 5: Construct the diagonal block form of the algebraic matrix corresponding to each agent in the form of diagonal block:
[0061] (5)
[0062] Specifically, is the scalar parameter to be designed. At the same time, according to the negative gradient, the auxiliary system corresponding to each agent is designed as:
[0063] (6)
[0064] Specifically, is the scalar parameter to be designed. is the auxiliary system state variable with the same structure as the real system state variable , and is a scalar function.
[0065] Step seven: according to the multi-agent system differential game model established in step three, the signal preserving matrix of each agent designed in step four, the algebraic matrix solution and auxiliary system constructed in step five, design the approximate optimal feedback control law of each agent:
[0066] (7)
[0067] Specifically, the matrix value function satisfies .
[0068] Step eight: input the approximate optimal feedback control law of each agent designed in step seven into each agent system, and update the system state quantity.
[0069] The following embodiments are used to illustrate the specific calculation process of the multi-agent differential game control method based on signal preserving matrix.
[0070] Step one: give the agent dynamics model formula (1) and initial conditions according to the multi-agent system cooperative control task shown in Figure 1 . As shown in Figure 1 , the dynamics equation of each agent in the LVLH system of the target to be surrounded and the tracking error system are expressed as follows:
[0071] (8)
[0072] (9)
[0073] Wherein, and ; the constant is the orbital angular velocity of the target satellite, the gravitational constant , is the modulus of the position vector of the target satellite in the earth-centered inertial system, and is taken as ; the tracking error state , the initial value and the expected value of the agent state variable are set as follows:
[0074] , (10)
[0075] Step two: according to the task requirement, design the individual performance index for realizing the maneuvering control and collision avoidance of each agent as shown in Figure 1 :
[0076] (11)
[0077] Wherein, the error state variable ; To consider the maneuver control and collision avoidance, and with positive constant parameters of the index term; and each agent safety radius constant related to the collision threat Designed as follows:
[0078] (12)
[0079] Step three: according to the agent dynamics model constructed in step one, the individual performance index of each agent designed in step two, establish multi-agent system differential game model:
[0080] (13)
[0081] Step four: according to the agent dynamics model established in step one, for , take the first Agent required signal holding matrix:
[0082] (14).
[0083] Where, take .
[0084] Step five: construct the diagonal block form of the algebraic matrix corresponding to the diagonal block form of each agent solution is:
[0085] (15)
[0086] Where, At the same time, according to the negative gradient design each agent corresponding to the auxiliary system is:
[0087] (16)
[0088] Where, scalar parameter ; The same structure as the real system state variable Auxiliary system state variable, its initial value is taken as:
[0089] (17)
[0090] Scalar function , take Unit matrix.
[0091] Step seven: according to the multi-agent system differential game model established in step three, the signal preserving matrix of each agent designed in step four, the algebraic matrix solution and the auxiliary system constructed in step five, the approximate optimal feedback control law of each agent is designed:
[0092] (18)
[0093] Step eight: input the approximate optimal feedback control law of each agent designed in step seven into each agent system, and update the system state quantity. By observing the system state quantity and the auxiliary system state quantity, the following equation can be obtained Figures 3-4 ; by observing the first three-dimensional state variables of each agent, the following equation can be obtained Figure 5 ; by observing the modulus value of the difference between the first three-dimensional state variables of any two agents, the following equation can be obtained Figure 6 .
[0094] Referring to Figures 3-6 , the present application obtains the approximate solution of the differential game problem composed of 2-relative order agents, avoiding the limitation of the traditional method which is only applicable to solving the differential game problem composed of 1-relative order agents. At the same time, for the differential game problem composed of 2-relative order agents, the multi-agent differential game control method based on signal preserving matrix involved in the present application can realize the cooperative control task of the multi-agent system, and make the system have asymptotic stability at the equilibrium point.
[0095] The above content is only for illustrating the technical idea of the present application, and cannot limit the protection scope of the present application. Any modification made according to the technical idea of the present application on the basis of the technical scheme falls within the protection scope of the claims of the present application.
Claims
1. A multi-agent differential game control method based on a signal preserving matrix, characterized by, The method comprises the following steps: Step 1: according to the selected agent object of the multi-agent system, establishing the dynamic model of each agent; According to the task requirement, designing the individual performance index of each agent; The number of the intelligent agents is , the i th intelligent agent The state variable at any time The dimension of the state variable is a positive integer The state variables include position coordinates and velocity coordinates , positive integers ; No. i At the initial moment, each intelligent agent... The state variables at time are , No. i A dynamic model of an intelligent agent in a given Cartesian inertial coordinate system has the following form: (1) wherein, is a control quantity, positive integer is a dimension of the control quantity, with the dimension of acceleration; represents the rate of change of the position of the agent about the mapping relationship, represents the rate of change of the velocity of the agent about the mapping relationship; is a description of the control quantity to the rate of change of the velocity of the agent the mapping; Step 2: according to the dynamic model of each agent established in step 1 and the individual performance index of each agent, establishing the differential game model of the multi-agent system; The form of the individual performance index of each agent is: (2) where the agent system state variables are composed of the agent state variables, i.e. , , is a scalar index related to the multi-agent system state variables , the matrix and is the weight of the control input; The differential game model of the multi-agent system is integrated according to formula (1) into the following form: (3); According to the dynamic model of each agent established in step 1, designing the signal maintaining matrix required by each agent; The signal maintaining matrix corresponding to each agent designed according to the dynamic model of each agent is: (4); wherein, such that no all-zero row, , , is a unit direction vector parallel to the state variable corresponding to the mapping all-zero row. Step 3: constructing the diagonal block algebraic matrix solution, constructing the auxiliary system, and designing the dynamic equation of the auxiliary system with asymptotic stability at the equilibrium point; Step 4: according to the differential game model of the multi-agent system established in step 2, the signal maintaining matrix of each agent and the algebraic matrix solution and the auxiliary system constructed in step 3, designing the approximate optimal feedback control law of each agent; Step 5: inputting the approximate optimal feedback control law of each agent designed in step 4 into each agent system to update the system state quantity.
2. The multi-agent differential game control method based on signal preserving matrix according to claim 1, wherein, In step 3, the diagonal block form of the algebraic matrix solution corresponding to each agent is: (5); At the same time, the auxiliary system corresponding to each agent is designed according to the negative gradient as: (6); Among them, matrix partitioning Indicates the first The algebraic matrix solution for each agent The Okay, number Matrix partitioning corresponding to columns; matrix partitioning middle, For scalar parameters, To partition the matrix Identity matrices of the same size and These represent the matrix blocks to be designed. The diagonal elements, It is the set of real numbers; To interact with the state variables of a multi-agent system Auxiliary system state variables with the same structure.
3. The multi-agent differential game control method based on signal preserving matrix according to claim 2, characterized in that, is a scalar function.
4. The multi-agent differential game control method based on signal preserving matrix according to claim 2, wherein, In step 4, the approximate optimal feedback control law of each agent is designed as: (7); where the matrix-valued function satisfies .