A game control method for cooperative target tracking of multiple unmanned vehicles
Through the target state observer and guidance law, compensation input saturation system, optimal cost function and evaluation network, a control strategy that meets Nash equilibrium is designed, and the problems of long convergence time and member interactions are not considered in the coordinated tracking of multiple unmanned boats are solved, and fast and stable target tracking control is achieved.
Patent Information
- Application Number
- CN202510588921.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing multi-unmanned boat collaborative tracking control strategy converges for a long time in complex marine environments, and fails to fully consider the interaction between formation members, resulting in reduced task execution efficiency and accuracy.
The target state observer and guidance law are used for the predetermined time, combined with an auxiliary system that compensates input saturation, an optimal cost function and evaluation network are established, and an approximate optimal control law that meets Nash equilibrium is designed, and a coordinated target tracking of multiple unmanned boats is achieved through adaptive law.
It realizes rapid convergence and stable control of multiple unmanned boat systems in complex marine environments, improves the overall performance of the formation and enhances the resistance to interference.
Smart Images

Figure CN120103871B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cooperative tracking control of unmanned boats, and in particular to a game control method for cooperative target tracking of multiple unmanned boats. Background Art
[0002] With the rapid development of technology, unmanned surface vehicle (USV) technology has evolved from a single platform to multi-platform collaborative and swarm systems. Multi-USV systems significantly improve mission execution efficiency and success rates through information sharing and collaborative decision-making. They have been widely used in fields such as marine resource exploration, environmental monitoring, and near-shore patrols. Collaborative tracking and control, as one of the key technologies, is crucial for ensuring the effective operation of multi-USV systems.
[0003] In complex ocean environments, USVs are subject to interference from a variety of factors, including direct effects from natural factors such as wind, waves, and currents, as well as challenges such as actuator input saturation. These interference factors not only affect the USV's trajectory, speed, and stability, but also increase the difficulty of navigation control, reducing the accuracy and efficiency of mission execution. Furthermore, when multiple USVs perform tracking missions, they need to collaborate with other USVs to achieve overall formation performance. During movement, USVs are influenced not only by their own control inputs but also by the motion states of surrounding USVs. This interaction manifests itself in fluid dynamics as vortices, waves, and currents generated by forces and torques, which can affect USV stability. To ensure the safe, stable, and efficient operation of USVs in complex ocean environments, it is necessary to comprehensively consider various interference factors and implement corresponding design and control strategy optimization measures. Therefore, solving the tracking control problem of multiple unmanned vehicles in complex ocean environments is of great significance.
[0004] Currently, research on the problem of collaborative tracking of multiple unmanned vehicles in complex marine environments has achieved considerable results. However, the following problems still exist in existing strategies:
[0005] Based on the currently widely adopted multi-USV target tracking control scheme, the control convergence time of multi-USV formation systems is generally long, which limits their application potential in missions requiring rapid response. Specifically, existing control strategies often lack a mechanism to pre-set and ensure that the system can reach a stable state or a predetermined formation configuration within a specific time. This lack of convergence speed not only affects the efficiency of formation execution but can also cause mission delays in some cases.
[0006] Furthermore, most existing research exploring multi-USV formation control often examines the control issues of each USV in isolation, neglecting the complex interactions and dynamic influences between formation members. While this simplified approach reduces the complexity of the problem to some extent, it also limits the overall performance of the formation. In reality, each member of a multi-USV formation is influenced not only by its own control strategy but also by the combined effects of the motion states of other members, external environmental factors, and the overall goals of the formation. Therefore, without fully considering these interactions, it is difficult to significantly improve and optimize the overall performance of the formation. Summary of the Invention
[0007] The present invention provides a multi-unmanned boat cooperative target tracking game control method to overcome the above technical problems.
[0008] A multi-unmanned vehicle cooperative target tracking game control method includes the following steps:
[0009] S1: Construct the dynamic model of the unmanned boat and the kinematic model of the unmanned boat considering the disturbance of the marine environment and input saturation;
[0010] S2: establishing a predetermined time target state observer based on the dynamic model and kinematic model of the unmanned boat to obtain an estimated value of the unmanned boat's position and attitude after the tracking target is offset, and an estimated value of the unmanned boat's velocity after the tracking target is offset;
[0011] S3: Establishing a predetermined time guidance law to obtain a longitudinal velocity reference signal and a bow velocity reference signal of the unmanned vehicle based on a differential equation of a position error system according to the unmanned vehicle's estimated values of the tracking target position and attitude after the tracking target position and attitude after the tracking target velocity after the tracking target velocity is offset;
[0012] S4: Establishing an auxiliary system for compensating input saturation to obtain a speed error of the auxiliary system considering the compensation input saturation according to the longitudinal speed reference signal and the bow speed reference signal of the unmanned boat;
[0013] S5: obtaining a transformed error vector based on an error performance function according to the speed error of the auxiliary system considering compensation input saturation;
[0014] S6: establishing an optimal cost function according to the transformed error vector, and obtaining an optimal control law for the unmanned vehicle based on the HJB equation;
[0015] S7: Establish an adaptive law of the evaluation network to obtain an approximate optimal control law that satisfies Nash equilibrium based on the optimal control law of the unmanned boat, so as to realize the control of collaborative target tracking of multiple unmanned boats based on the approximate optimal control law that satisfies Nash equilibrium.
[0016] Beneficial effects: The present invention provides a multi-unmanned boat collaborative target tracking game control method, which obtains the estimated value of the unmanned boat after the tracking target is offset by establishing a predetermined time target state observer, and obtains the reference signal of the unmanned boat through a predetermined time guidance law, and obtains the speed error of the auxiliary system considering the compensation input saturation based on the auxiliary system that compensates for the input saturation, and then obtains the transformed error vector; then, the optimal control law of the unmanned boat is obtained through the established optimal cost function, and the approximate optimal control law that satisfies the Nash equilibrium is obtained based on the adaptive law of the evaluation network, thereby realizing the control of the multi-unmanned boat collaborative target tracking. By establishing the predetermined time guidance law, the upper limit of the convergence time can be set in advance to solve the problem of long convergence time and reduce the dependence of the convergence time on the control parameters. At the same time, in order to improve the performance of the entire system, the influence of neighboring members in the formation is fully considered, and the evaluation network is used to obtain the approximate optimal control law that satisfies the Nash equilibrium, thereby obtaining a solution that satisfies the Nash equilibrium and solving the Nash game problem between multiple unmanned boats in a nonlinear continuous-time system with unknown interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 is a flow chart of the collaborative target tracking game control method of the present invention;
[0019] Figure 2 is a communication topology diagram in an embodiment of the present invention;
[0020] Figure 3 This is a block diagram of the multi-unmanned boat cooperative target tracking game control structure in an embodiment of the present invention;
[0021] Figure 4 This is a tracking effect diagram of three unmanned boats in an embodiment of the present invention;
[0022] Figure 5 is the evaluation network weight change curve of the three unmanned boats in the embodiment of the present invention;
[0023] Figure 6 is the action network weight change curve of the three unmanned boats in the embodiment of the present invention;
[0024] Figure 7 is the speed error curve of USV1 in the embodiment of the present invention;
[0025] Figure 8is the speed error curve of USV2 in an embodiment of the present invention;
[0026] Figure 9 is the speed error curve of USV3 in an embodiment of the present invention;
[0027] Figure 10 is the longitudinal thrust curve of the three unmanned boats in the embodiment of the present invention;
[0028] Figure 11 1 is the bow moment curve of the three unmanned boats in the embodiment of the present invention. DETAILED DESCRIPTION
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0030] This embodiment discloses a multi-unmanned boat cooperative target tracking game control method, including the following steps: Figure 1 As shown:
[0031] Specifically, this embodiment discusses a multi-UAV cooperative target tracking game control method taking into account environmental interference and mutual influence between formation members.
[0032] S1: Construct the dynamic model of the unmanned boat and the kinematic model of the unmanned boat considering the disturbance of the marine environment and input saturation;
[0033] Specifically, in this embodiment, a dynamic model of the unmanned boat is constructed and used as the controlled object, and the position information in the earth coordinate system is obtained through the kinematic model. The dynamic model of the unmanned boat includes the external environmental disturbance caused by the waves.
[0034] Preferably, the dynamic model of the unmanned boat is established as follows:
[0035] (1)
[0036] The kinematic model of the unmanned boat is established as follows:
[0037] ,
[0038] Where: Indicates the The position and attitude matrix of the unmanned boat, where , 、 、 Respectively represent the longitudinal position coordinate, lateral position coordinate and heading angle of the unmanned boat, represents transpose; Represents the coordinate system rotation matrix; Indicates the The velocity matrix of the unmanned boat, where 、 、 represent the longitudinal speed, lateral speed and heading speed respectively; represents the inertia matrix; express The first differential of represents the Coriolis matrix and the centripetal matrix; represents the fluid dynamics damping matrix; represents the external disturbance, where ; They represent longitudinal disturbance, lateral disturbance and heading disturbance respectively; represents a control input with saturation, where , represents the longitudinal thrust with saturation, represents the heading moment with saturation;
[0039] in:
[0040] ,
[0041] ,
[0042] ,
[0043] ,
[0044] ,
[0045] in, It is the mass of the unmanned boat; is the longitudinal additional mass coefficient generated by the longitudinal acceleration; is the longitudinal additional mass coefficient caused by lateral acceleration; is the additional moment of inertia coefficient of heading caused by heading acceleration; The unmanned boat is The moment of inertia of the shaft; is the lateral additional mass coefficient generated by the lateral acceleration; is the longitudinal linear water damping coefficient generated by the longitudinal velocity; is the second-order longitudinal nonlinear water damping coefficient generated by the longitudinal velocity; Indicates absolute value; is the lateral linear water damping coefficient generated by the lateral velocity; is the second-order lateral nonlinear water damping coefficient generated by the lateral velocity; is the linear water damping coefficient of the heading caused by the heading angular velocity; is the second-order heading nonlinear water damping coefficient generated by the heading angular velocity; Indicates the heading angle of the unmanned boat; represents the control input, where , The longitudinal thrust, represents the heading moment with saturation; is the maximum value vector of the control input, where , is the maximum longitudinal thrust, is the maximum heading moment; represents a saturation function.
[0046] S2: Based on the UAV's dynamic model and kinematic model, and considering the situation where the target state is not completely known, a target state observer is established at a predetermined time to obtain the UAV's estimated values of the tracking target's position and attitude after the tracking target is offset, and the UAV's estimated value of the tracking target's velocity after the tracking target is offset, so as to track the estimated value of the target's motion state; wherein, the estimated value of the target's motion state will be used in the subsequent design of the predetermined time guidance law. The motion state includes the UAV's lateral velocity, bow velocity, and longitudinal velocity;
[0047] Preferably, the predetermined time target state observer is constructed as follows:
[0048] (2)
[0049] Where: It is the first The differential of the estimated vector of the target position and attitude after the unmanned boat is offset, where , is the differential of the longitudinal offset position estimate, is the differential of the lateral offset position estimate, is the differential of the offset heading angle estimate; It is the first The differential of the estimated vector of the target velocity after the unmanned boat is offset, where
[0050] , is the second-order differential of the estimated longitudinal position of the tracking target, is the second-order differential of the estimated lateral position of the tracking target, is the second-order differential of the estimated heading angle of the tracking target; It is the first The estimated vector of the target speed after the unmanned boat is offset; , and are the gain coefficients of position error, velocity error and saturation function in the observer respectively; , The scheduled time for position estimation and the scheduled time for speed estimation are respectively; , are all observer position error index parameters, , are all observer velocity error index parameters, and , , , ; , are position error term and velocity error term respectively; It is The second-order differential of the offset between the unmanned boat and the target, where , It is The second-order differential of the longitudinal offset between the unmanned boat and the target, It is The second-order differential of the lateral offset between the unmanned boat and the target, Indicates the The second-order differential of the heading angle offset between the unmanned boat and the target; represents the intermediate calculation parameters, where , is any variable, is a constant; represents a symbolic function;
[0051] in,
[0052] (3)
[0053] Where: , All are index numbers of unmanned boats; represents the total number of unmanned boats; It is Unmanned boat and The weight coefficient between the unmanned boats; It is the first The estimated vector of the target position and attitude after the unmanned boat is offset, where , It is the first The estimated value of the target longitudinal position after the unmanned boat is offset. It is the first The estimated value of the target's lateral position after the unmanned boat shifts, It is the first The estimated value of the heading angle of the unmanned boat to the target after deviation; It is the first The estimated vector after the target position and attitude deviation of the unmanned boat is estimated. , It is the first The estimated value of the target longitudinal position after the unmanned boat is offset. It is the first The estimated value of the target's lateral position after the unmanned boat shifts, It is the first The estimated value of the heading angle of the unmanned boat to the target after deviation; It is The position and attitude offset between the unmanned boat and the tracking target, , It is The longitudinal position offset between the unmanned boat and the tracking target, It is The offset between the lateral position of the unmanned boat and the tracking target, It is The offset of the heading angle between the unmanned boat and the tracking target; It is The position and attitude offset between the unmanned boat and the tracking target, , It is The longitudinal position offset between the unmanned boat and the tracking target, It is The offset between the lateral position of the unmanned boat and the tracking target, It is The offset of the heading angle between the unmanned boat and the tracking target; It is The weight coefficient between the unmanned boat and the tracking target; is the position and attitude vector of the tracking target, , is the longitudinal position of the tracking target, is the lateral position of the tracking target, is the heading angle of the tracking target; It is the first The estimated value of the target speed after the unmanned boat is offset. , It is The differential of the estimated longitudinal velocity of the unmanned boat to the tracked target, It is The differential of the estimated lateral velocity of the unmanned boat to the tracked target, It is The differential of the estimated value of the heading angular velocity of the tracking target by the unmanned boat; It is The differential of the position and attitude offset between the unmanned boat and the tracking target, , It is The differential of the longitudinal position offset between the unmanned boat and the tracking target, It is The differential of the lateral position offset between the unmanned boat and the tracking target, It is The differential of the heading angle offset between the unmanned boat and the tracking target; It is The differential of the position and attitude offset between the unmanned boat and the tracking target, , It is The differential of the longitudinal position offset between the unmanned boat and the tracking target, It is The differential of the lateral position offset between the unmanned boat and the tracking target, It is The differential of the heading angle offset between the unmanned boat and the tracking target; is the velocity vector of the tracking target in the geodetic coordinate system, , is the differential of the longitudinal position of the tracking target, is the differential of the lateral position of the tracking target, is the differential of the heading angle of the tracking target.
[0054] Specifically, in order to enable all USVs in the formation to obtain the motion state of the tracking target, the motion state of the target USV is observed, and the predetermined time theory is used to construct a state observer to estimate the real-time motion state of the target USV. The subsequent design of the predetermined time guidance law will refer to this state;
[0055] S3: Establishing a predetermined time guidance law to obtain a longitudinal velocity reference signal and a bow velocity reference signal of the unmanned vehicle based on a differential equation of a position error system according to the unmanned vehicle's estimated values of the tracking target position and attitude after the tracking target position and attitude after the tracking target velocity after the tracking target velocity is offset;
[0056] Specifically, the estimated values of the tracking target position and attitude after the unmanned boat is offset, and the estimated value of the tracking target speed after the unmanned boat is offset, and the predetermined time guidance law is designed in combination with the motion state of each USV itself, so as to provide the reference longitudinal speed and reference bow angular velocity for achieving target tracking; the reference longitudinal speed and reference bow angular velocity will be used in the calculation of the velocity error system later.
[0057] Specifically, based on the estimated value of the target USV's motion state after offset and the motion state of the formation member USVs, the position error with the formation member USVs is calculated, and a predetermined time guidance law that stabilizes the position error system is obtained. The speed error system that considers control input saturation compensation will use this guidance law as the speed reference value for subsequent calculations.
[0058] Preferably, the differential equation of the position error system is expressed as follows:
[0059] , , (4)
[0060] Where: Indicates the The differential of the error between the real-time longitudinal position of the target unmanned boat and its estimated longitudinal position after the offset from the target unmanned boat; Indicates the The differential of the error between the real-time lateral position of the target unmanned boat and its estimated lateral position after the offset from the target unmanned boat; Indicates the The differential of the error between the real-time heading angle of the unmanned boat and the estimated heading angle after the unmanned boat is offset from the target unmanned boat; Indicates the The position and attitude matrix of each unmanned boat; Indicates the first The estimated value of the longitudinal velocity of the target unmanned boat by the unmanned boat, Indicates the body coordinate system The estimated value of the lateral velocity of the target unmanned boat by the unmanned boat, Indicates the first The estimated value of the bow speed of the target unmanned boat by the unmanned boat, is the target heading angle; Indicates the The error between the real-time heading angle of a UAV and its estimated heading angle after the UAV is offset from the target UAV. Indicates the The error between the real-time lateral position of a UAV and its estimated lateral position after the offset from the target UAV; Indicates the The error between the real-time longitudinal position of the target unmanned boat and its estimated longitudinal position after the offset from the target unmanned boat; Represents the inverse of the coordinate system rotation matrix.
[0061] Preferably, in order to make the position error system converge within a predetermined time, the predetermined time guidance law is established as follows:
[0062] (5)
[0063] in,
[0064] ,
[0065] ,
[0066] Where: Indicates the The longitudinal velocity reference signal of the unmanned boat; Indicates the The error reference signal between the real-time heading angle of the unmanned boat and its estimated heading angle of the target unmanned boat; Indicates the The bow velocity reference signal of an unmanned boat; 、 are the gain coefficients of longitudinal errors; 、 are the gain coefficients of lateral errors; 、 are all gain coefficients of heading angle error; is the scheduled convergence time; is a constant in the guidance law error coefficient, and .
[0067] S4: Establish an auxiliary system that compensates for input saturation to obtain the speed error of the auxiliary system considering the compensation input saturation based on the longitudinal speed reference signal and the bow speed reference signal of the unmanned boat, so as to calculate the real-time speed error of each USV; the performance of the speed error system will be specified later.
[0068] Specifically, to address the issue of control input truncation caused by input saturation, this embodiment introduces an auxiliary system to compensate for input saturation. This auxiliary system is integrated into the real-time speed error between each USV and the target USV. By adjusting the control input in real time, the negative impact of input saturation is mitigated. Subsequently, to ensure that the error is effectively confined to a predetermined region, improve the error convergence process, and enhance system stability and accuracy, the speed error system will undergo a performance transformation.
[0069] Preferably, the auxiliary system for compensating input saturation is constructed as follows:
[0070] (6)
[0071] Where: is the differential of the auxiliary vector; It is an auxiliary vector used to deal with the control input truncation problem caused by input saturation; Indicates the inversion of a matrix or the negation of a function;
[0072] Preferably, the speed error of the auxiliary system considering the compensation input saturation is obtained as follows:
[0073] First, establish the velocity error system differential equation:
[0074] (7)
[0075] Where: represents the differential of the velocity error of the auxiliary system taking into account the compensation input saturation; represents the differential of the reference velocity signal vector, ,in, Indicates the The first-order differential of the longitudinal velocity reference signal of the unmanned boat; Indicates the The first-order differential of the bow velocity reference signal of the unmanned boat; Indicates the auxiliary system number;
[0076] Arranging the differential equation of the velocity error system, we can obtain:
[0077] (8)
[0078] ,
[0079] Where: Represents a dynamic matrix; represents the Coriolis matrix and the centripetal matrix; represents the fluid mechanics damping matrix.
[0080] S5: obtaining a transformed error vector based on an error performance function according to the speed error of the auxiliary system considering compensation input saturation;
[0081] Specifically, based on the calculated speed errors of each unmanned boat, the error transformation is performed in combination with the error performance function to obtain the speed error after the specified performance change, thereby improving the dynamic process of error convergence; subsequently, the dynamic control law will be designed based on the speed error system.
[0082] Specifically, after obtaining the speed error of the auxiliary system considering the compensation input saturation, these errors are transformed using the error performance function to convert the original speed error into a speed error with specific performance characteristics. The changes of these errors are smoother, thereby improving the speed control accuracy and stability of the unmanned boat, and finding a stable control strategy in the subsequent case of considering the interaction between multiple unmanned boats.
[0083] Preferably, the error performance function is established as follows and is designed as the following hyperbolic tangent function:
[0084] , , (9)
[0085] Where: represents the hyperbolic tangent function; represents the transformed error vector, , express is the inverse function of , ; It represents the speed error of the auxiliary system after the speed error of the auxiliary system is constrained by the performance function considering the compensation input saturation; Represents the lower bound coefficient of the error limit region; Represents the upper bound coefficient of the error limit region; is the error performance index function; , are constants, and , ; is a positive constant of the error domain convergence rate; For time; Indicates the maximum value of the error performance index function; Represents the minimum value of the error performance index function;
[0086] Furthermore, the changed error vector is obtained through the following transformed velocity error differential equation:
[0087] , , , (10)
[0088] Where: represents the error differential after transformation; Represents the transformed dynamic matrix; represents the transformed inertia matrix; Represents the inertia matrix transformation coefficient; express The first differential of represents the hyperbolic tangent function.
[0089] S6: establishing an optimal cost function according to the transformed error vector, and obtaining an optimal control law for the unmanned vehicle based on the HJB equation;
[0090] According to the transformed error vector and considering the control law of the neighboring nodes, that is, jThe control law of the unmanned boat is designed, the optimal cost function is designed, and then the Hamilton-Jacobi-Bellman equation (HJB) is designed to calculate the optimal control law of the unmanned boat;
[0091] Specifically, in order to obtain a control law that satisfies Nash equilibrium, the USV's speed error, control law, and the control law of neighboring nodes are introduced into the optimal cost function. According to the optimal cost function of this embodiment, the HJB equation can be obtained, and the optimal control law is obtained based on the HJB equation. The optimal cost function constructed in this embodiment is as follows:
[0092] (11)
[0093] Where: represents the optimal cost function; It is Optimal control law for an unmanned boat; It is Optimal control law for an unmanned boat; Indicates the lower limit of the integral; represents the discount factor; Indicates the integration time; express The transpose of It is Neighbor nodes of unmanned boats; Indicates the strength of the control law of neighboring nodes; represents transpose; represents the differential of the integration time;
[0094] According to the optimal cost function of this embodiment, and by designing an optimal controller that satisfies the HJB equation, all USVs can achieve Nash equilibrium by implementing the optimal control strategy;
[0095] Preferably, the Nash equilibrium of the optimal control law of the unmanned boat can be achieved through the optimal cost function, and the method is as follows:
[0096] Specifically, in order to prove the Nash equilibrium of the optimal control law, this embodiment re-expresses the optimal cost function as follows:
[0097] (12)
[0098] Where: represents the deformation of the cost function after deformation; represents the lower bound of the integral; represents the discount factor in the optimal cost function; Indicates the strength of the control law of neighboring nodes; Represents the gradient of the optimal cost function; represents the initial value of the optimal cost function; Represents the value of the optimal cost function at infinite time;
[0099] Arrange formula (12) and we get:
[0100] (13)
[0101] According to the optimal control law of the unmanned boat, , (13) is expressed as:
[0102] (14)
[0103] According to the optimal control law of the unmanned boat, , (14) is expressed as:
[0104] (15)
[0105] Further expressed as:
[0106] ,
[0107] (16)
[0108] in, Indicates the The optimal cost function when an unmanned boat obtains the best response strategy;
[0109] when , can get , at this time The unmanned boats obtain the best response strategy. When all the unmanned boats obtain the best response strategy, the Nash equilibrium of the optimal control law of the unmanned boats can be achieved.
[0110] Preferably, the HJB equation is expressed as follows:
[0111] (17)
[0112] Where: represents the Hamilton–Jacobi–Bellman equation; Indicates the The position and attitude matrix of each unmanned boat; is the gradient of the optimal cost function; Represents a set of neighbor nodes; represents the discount factor; represents the transformed error vector; Represents the transformed dynamic matrix;
[0113] Then the following optimal control law is obtained:
[0114] (18)
[0115] Where: represents the transformed inertia matrix; represents transpose; It is Optimal control law for an unmanned boat.
[0116] S7: Establishing an adaptive law of the evaluation network to obtain an approximate optimal control law that satisfies Nash equilibrium according to the optimal control law of the unmanned boat;
[0117] Specifically, this embodiment designs an action-evaluation neural network learning structure based on the ideal optimal control law to approximate it.
[0118] Specifically, the optimal control law of the unmanned boat needs to be solved by the gradient of the optimal cost function of the HJB equation. However, the HJB equation has complex nonlinear characteristics, and it is not feasible to obtain a usable optimal solution (18), which brings great difficulties to the control system design. In this case, two neural network-based approximators are deployed to derive the optimal controller that depends on the tracking error signal and can approximate formula (18). The S7 includes giving a game-based evaluation network weight adaptation law, an action network weight adaptation law, and an approximate optimal control law that satisfies Nash equilibrium.
[0119] Preferably, the method for obtaining the approximate optimal control law that satisfies Nash equilibrium is as follows:
[0120] S71: Establishing an evaluation network weight adaptive law to obtain an evaluation network weight adaptive law vector;
[0121] The evaluation network of this embodiment adopts a traditional RBF neural network structure, including an input layer, a hidden layer and an output layer. The input layer and the hidden layer each contain three neurons, and the hidden layer contains one neuron.
[0122] Preferably, the evaluation network weight adaptive law is established as follows:
[0123] (19)
[0124] ,
[0125] Where: represents the adaptive law vector of the evaluation network weight, ,in, , , Both are the differentials of the corresponding weights after evaluating the output of the hidden layer of the network; Indicates time period The difference of the Gaussian basis function vector within ; , represents the Gaussian basis function vector, ,
[0126] , Represent the outputs of the first, second, and third neurons in the hidden layer respectively; represents the network input vector, represents the central constant vector in the hidden layer, ,in, , , are the center point vectors of the Gaussian basis functions of each neuron in the hidden layer, represents the width of the Gaussian basis function; ; is the evaluation network weight vector, , , , They are the corresponding weights after the output of the hidden layer of the evaluation network; is the error gain coefficient; Indicates time period The difference vector of the error norm within ; is the discount factor; is the integration variable; It’s time; is the time period; express Speed error at the moment; express Speed error at the moment;
[0127] S72: Establish an action network weight adaptation law based on the evaluation network weight adaptation law vector to obtain the action network weight adaptation law vector. The formula used is as follows:
[0128] Preferably, the action network weight adaptation law is as follows:
[0129] (20)
[0130] Where: represents the action network weight adaptation law vector, ,in, , , Both are the differentials of the corresponding weights after the output of the hidden layer of the action network; is the gain coefficient; is the action network weight vector, , , , These are the corresponding weights after the output of the hidden layer of the action network; is the basis function gradient vector; represents the discount factor in the action network;
[0131] In this embodiment, the action network also uses the RBF neural network structure;
[0132] S73: According to the action network weight adaptive law vector, an approximate optimal control law satisfying Nash equilibrium is obtained as follows:
[0133] , (twenty one)
[0134] Where: represents the approximate optimal control law that satisfies Nash equilibrium, is the approximation error.
[0135] Specifically, a specific simulation experiment embodiment of the present invention is as follows:
[0136] Communication topology of multiple unmanned boat systems Figure 2 As shown in the figure, three unmanned boats are selected to form the formation members, one USV is the tracking target, and the four USVs together form the communication system. The model parameters of the unmanned boat are:
[0137] , ,
[0138] , ,
[0139] , ,
[0140] , ,
[0141] , ,
[0142] The external disturbances in the longitudinal, transverse and bow directions are:
[0143] ,
[0144] The initial states of the three unmanned boats are expressed as: , , , The motion state of the tracking target is , The target trajectory offset is:
[0145] ,
[0146] in, It is The second-order differential of the position and attitude offset between the unmanned boat and the tracking target;
[0147] The relevant parameters of the target state scheduled time observer are , , , , ;
[0148] , are all observer position error index parameters, , are all observer velocity error index parameters, and , , , ; , The scheduled time for position estimation and the scheduled time for speed estimation are respectively; represents the gain coefficient of the saturation function; 、 Represents the gain coefficient of position error and velocity error in the observer;
[0149] The relevant parameters of the performance function are , , , ;
[0150] Represents the lower bound coefficient of the error limit region; Represents the upper bound coefficient of the error limit region; is a positive constant of the error domain convergence rate; Indicates the maximum value of the error performance index function; Represents the minimum value of the error performance index function;
[0151] The parameters related to the predetermined time guidance law are , =0.03, is the scheduled convergence time; is a constant in the guidance law error coefficient, and ;
[0152] The maximum force and moment for input saturation are , , , ;
[0153] The initial parameters of the evaluation network and action network weights are , .
[0154] Figure 3 This is a block diagram of the coordinated target tracking game control architecture of three unmanned aerial vehicles (UAVs). The UAVs exchange state information via a wireless communication network. Each UAV's control system consists of a timed guidance law, a timed state observer, auxiliary systems, a performance specification system, and an evaluation-action network. As the state of the tracking target UAV changes, each UAV's timed observer estimates its state in real time and calculates the timed guidance law based on the estimated value. Each UAV then uses the evaluation-action network to approximate the optimal Nash equilibrium control law, ensuring that the UAV moves according to the velocity reference signal.
[0155] Figure 4-11 Is a simulation effect diagram, where Figure 4 The collaborative target tracking effect of multiple unmanned boats is demonstrated, and it can be seen that each unmanned boat can accurately track the target according to the corresponding estimated trajectory. Figure 5 and Figure 6 These are the change curves of the evaluation network weight and the action network weight. It can be seen that the weight curves eventually converge. Figure 7 、 Figure 8 and Figure 9 These are the longitudinal speed and bow speed error curves of the three unmanned boats. It can be seen that the speed error curves are all within the specified performance area. Figure 10 and Figure 11 The longitudinal thrust curves and bow moment curves of three unmanned watercraft (USV1, USV2, and USV3) with and without input saturation are shown. The control law with input saturation produces truncation, but due to the introduction of an auxiliary system in the algorithm of this embodiment, good control effect is still achieved.
[0156] This embodiment's multi-UAV cooperative target tracking game control method takes into account factors such as marine environmental interference, the impact between formation members, and input saturation. By combining scheduled time theory, game control methods, and optimal control theory, a target tracking control strategy was designed to achieve multi-UAV target tracking in complex environments.
[0157] Compared with the existing unmanned boat target tracking method, this embodiment takes into account the interference of the marine environment, the influence between the formation members and the saturation of the control input. In order to solve the problem that the state of the tracking target is not globally known, the state of the target USV is reconstructed by the predetermined time state observer, and each USV obtains the accurate target state within the predetermined time. On this basis, a predetermined time guidance law is designed to make the position error system converge within the predetermined time. Then, the speed error is calculated according to the predetermined time guidance law, and the input saturation compensation system is introduced, and the speed error is performance-regulated. Finally, the optimal control law that satisfies the Nash equilibrium is solved through the evaluation-action network. In summary, the present invention can achieve accurate and rapid tracking of targets by multiple unmanned boats in complex environments.
[0158] The multi-UAV collaborative target tracking game control method of this embodiment establishes a two-layer distributed control structure. In an environment where global target information is unknown, the upper-layer control framework uses a predetermined-time target state observer to estimate target information and formulates a predetermined-time guidance law based on this information. This ensures that the target estimated position and attitude tracking errors converge to a neighborhood near the origin within a predetermined time. In the lower-layer control framework, an online reinforcement learning algorithm is designed to solve the control law that satisfies the Nash equilibrium, using an evaluation-action neural network (NN) structure. The control strategies of neighboring members are included in the objective function of each USV. This method is used to solve the multi-UAV Nash game problem of a nonlinear continuous-time system with unknown interference, and takes into account speed error constraints and input saturation compensation. This embodiment significantly improves the overall control performance of the UAV formation and enhances its resistance to various interferences, providing important technical support for the widespread application of UAVs.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-unmanned boat cooperative target tracking game control method, characterized in that: The steps include: S1: Construct the dynamic model of the unmanned boat and the kinematic model of the unmanned boat considering the disturbance of the marine environment and input saturation; S2: establishing a predetermined time target state observer based on the dynamic model and kinematic model of the unmanned boat to obtain an estimated value of the unmanned boat's position and attitude after the tracking target is offset, and an estimated value of the unmanned boat's velocity after the tracking target is offset; S3: Establishing a predetermined time guidance law to obtain a longitudinal velocity reference signal and a bow velocity reference signal of the unmanned vehicle based on a differential equation of a position error system according to the unmanned vehicle's estimated values of the tracking target position and attitude after the tracking target position and attitude after the tracking target velocity after the tracking target velocity is offset; S4: Establishing an auxiliary system for compensating input saturation to obtain a speed error of the auxiliary system considering the compensation input saturation according to the longitudinal speed reference signal and the bow speed reference signal of the unmanned boat; S5: obtaining a transformed error vector based on an error performance function according to the speed error of the auxiliary system considering compensation input saturation; S6: establishing an optimal cost function according to the transformed error vector, and obtaining an optimal control law for the unmanned vehicle based on the HJB equation; S7: Establish an adaptive law of the evaluation network to obtain an approximate optimal control law that satisfies Nash equilibrium based on the optimal control law of the unmanned boat, so as to realize the control of collaborative target tracking of multiple unmanned boats based on the approximate optimal control law that satisfies Nash equilibrium.
2. The method for controlling target tracking by multiple unmanned vehicles in a coordinated game according to claim 1, characterized in that: The predetermined time target state observer is constructed as follows: Where: is the differential of the estimated vector of the target position and attitude offset by the i-th unmanned boat in the geodetic coordinate system, where is the differential of the longitudinal offset position estimate, is the differential of the lateral offset position estimate, is the differential of the offset heading angle estimate; is the differential of the estimated vector of the target velocity after the i-th unmanned boat is offset in the geodetic coordinate system, where is the second-order differential of the estimated longitudinal position of the tracking target, is the second-order differential of the estimated lateral position of the tracking target, is the second-order differential of the estimated heading angle of the tracking target; is the estimated vector of the target velocity offset by the i-th unmanned boat in the geodetic coordinate system; κ1, κ2 and α are the gain coefficients of the position error, velocity error and saturation function in the observer respectively; T1 and T2 are the scheduled time for position estimation and the scheduled time for velocity estimation respectively; m1 and n1 are the observer position error index parameters, m2 and n2 are the observer velocity error index parameters, and 0 <m1<1,0<m2<1,n1> 1, n2>1; z η,i ,z v,i are position error term and velocity error term respectively; is the second-order differential of the offset between the i-th unmanned boat and the target, where is the second-order differential of the longitudinal offset between the i-th unmanned boat and the target, is the second-order differential of the lateral offset between the i-th unmanned boat and the target, represents the second-order differential of the heading angle offset between the i-th unmanned boat and the target; represents the intermediate calculation parameters, where χ is an arbitrary variable, ι1 is a constant; sign(χ) represents the sign function; in, Where: i, j are the index numbers of the unmanned boats; N represents the total number of unmanned boats; a ij is the weight coefficient between the i-th unmanned boat and the j-th unmanned boat; is the estimated vector of the target position and attitude offset by the i-th unmanned boat in the geodetic coordinate system, where is the estimated value of the longitudinal position offset of the target by the i-th unmanned boat in the geodetic coordinate system, is the estimated value of the lateral position offset of the target by the i-th unmanned boat in the geodetic coordinate system, is the estimated value of the heading angle of the i-th unmanned boat relative to the target in the geodetic coordinate system; is the estimated vector of the target position and attitude deviation of the j-th unmanned boat in the geodetic coordinate system, is the estimated value of the longitudinal position offset of the target by the j-th unmanned boat in the geodetic coordinate system, is the estimated value of the lateral position offset of the target by the j-th unmanned boat in the geodetic coordinate system, is the estimated value of the heading angle of the j-th unmanned boat to the target in the geodetic coordinate system; i is the position and attitude offset between the i-th unmanned boat and the tracking target, o x,i is the longitudinal position offset between the i-th unmanned boat and the tracking target, o y,i is the offset between the lateral position of the i-th unmanned boat and the tracking target, is the offset of the heading angle between the i-th unmanned boat and the tracking target; o j is the position and attitude offset between the jth unmanned boat and the tracking target, O x,j is the offset between the jth unmanned boat and the longitudinal position of the tracking target, o y,i is the offset between the jth unmanned boat and the lateral position of the tracking target, is the offset of the heading angle between the jth unmanned boat and the tracking target; l il is the weight coefficient between the i-th unmanned boat and the tracking target; η0 is the position and attitude vector of the tracking target, x0 is the longitudinal position of the tracking target, y0 is the lateral position of the tracking target, is the heading angle of the tracking target; is the estimated value of the target speed after the j-th unmanned boat is offset in the geodetic coordinate system, is the differential of the estimated longitudinal velocity of the j-th unmanned boat to the tracking target, is the differential of the estimated lateral velocity of the j-th unmanned boat to the tracking target, is the differential of the estimated value of the heading angular velocity of the j-th unmanned boat to the tracking target; is the differential of the position and attitude offset between the i-th unmanned boat and the tracking target, is the differential of the longitudinal position offset between the i-th unmanned boat and the tracking target, is the differential of the lateral position offset between the i-th unmanned boat and the tracking target, is the differential of the heading angle offset between the i-th unmanned boat and the tracking target; is the differential of the position and attitude offset between the jth unmanned boat and the tracking target, is the differential of the longitudinal position offset between the jth unmanned boat and the tracking target, is the differential of the lateral position offset between the jth unmanned boat and the tracking target, is the differential of the heading angle offset between the jth unmanned boat and the tracking target; ω0 is the velocity vector of the tracking target in the geodetic coordinate system, is the differential of the longitudinal position of the tracking target, is the differential of the lateral position of the tracking target, is the differential of the heading angle of the tracking target.
3. The method for cooperative target tracking of multiple unmanned vehicles according to claim 2, characterized in that: The dynamic model of the unmanned boat is established as follows: The kinematic model of the unmanned boat is established as follows: Where: η i Represents the position and attitude matrix of the i-th unmanned boat, where x i 、y i 、 They represent the longitudinal position coordinate, lateral position coordinate and heading angle of the unmanned boat respectively, and T represents transposition; represents the coordinate system rotation matrix; v bi =[u i , v i , r i ] T represents the velocity matrix of the i-th unmanned boat, where u i 、v i 、r i Represent longitudinal speed, lateral speed and heading speed respectively; M i represents the inertia matrix; Indicates v bi The first-order differential of i (v bi ) represents the Coriolis matrix and the centripetal matrix; D i (v bi ) represents the fluid dynamics damping matrix; τ id represents the external disturbance, where τ id =[τ uid ,τ vid ,τ rid ] T ; τ uid , τ vid , τ rid represent longitudinal disturbance, lateral disturbance and heading disturbance respectively; μ(τ i ) represents a control input with saturation, where μ(τ i )=[μ(τ ui ),0,μ(τ ri )] T μ(τ ui ) represents the longitudinal thrust with saturation, μ(τ ri ) represents the heading moment with saturation; in: Where m is the mass of the unmanned boat; is the longitudinal additional mass coefficient generated by the longitudinal acceleration; is the longitudinal additional mass coefficient caused by lateral acceleration; I is the additional moment of inertia coefficient of heading caused by heading acceleration; z represents the moment of inertia of the unmanned boat about the z-axis; is the lateral additional mass coefficient caused by lateral acceleration; X u is the longitudinal linear water damping coefficient generated by the longitudinal velocity; X u|u| is the second-order longitudinal nonlinear water damping coefficient generated by the longitudinal velocity; |·| represents the absolute value; Y v Y is the lateral linear water damping coefficient generated by the lateral velocity; v|v| is the second-order lateral nonlinear water damping coefficient generated by the lateral velocity; N r is the linear water damping coefficient of the heading caused by the heading angular velocity; N r|r| is the second-order heading nonlinear water damping coefficient generated by the heading angular velocity; represents the heading angle of the unmanned boat; τ i represents the control input, where τ i =[τ ui ,0,τ ri ] T , τ ui The longitudinal thrust, τ ri represents the saturated heading moment; τ iM is the maximum value vector of the control input, where τ iM =[τ uiM ,0,τ riM] T , τ uiM is the maximum longitudinal thrust, τ riM is the maximum heading moment; tanh(·) represents the saturation function.
4. The method for controlling multi-unmanned vehicle cooperative target tracking game according to claim 3, characterized in that: The differential equation of the position error system is expressed as follows: Where: represents the differential of the error between the real-time longitudinal position of the i-th unmanned boat and its estimated longitudinal position after the offset from the target unmanned boat; represents the differential of the error between the real-time lateral position of the i-th unmanned boat and its estimated lateral position after the offset from the target unmanned boat; represents the differential of the error between the real-time heading angle of the i-th unmanned boat and its estimated heading angle after the deviation from the target unmanned boat; η i represents the position and attitude matrix of the i-th unmanned boat; represents the estimated value of the longitudinal velocity of the i-th unmanned boat to the target unmanned boat in the appendage coordinate system, represents the estimated value of the lateral velocity of the i-th unmanned boat relative to the target unmanned boat in the body coordinate system, represents the estimated value of the bow velocity of the i-th unmanned boat to the target unmanned boat in the appendage coordinate system, is the target heading angle; represents the error between the real-time heading angle of the i-th unmanned boat and its estimated heading angle after the deviation from the target unmanned boat; y i,e represents the error between the real-time lateral position of the i-th unmanned boat and its estimated lateral position after the offset from the target unmanned boat; x i,e represents the error between the real-time longitudinal position of the i-th unmanned boat and its estimated longitudinal position after the offset from the target unmanned boat; Represents the inverse of the coordinate system rotation matrix.
5. The method for controlling cooperative target tracking of multiple unmanned vehicles according to claim 4, characterized in that: The predetermined time guidance law is established as follows: in, Where: α i,u represents the longitudinal velocity reference signal of the i-th unmanned boat; represents the error reference signal between the real-time heading angle of the i-th unmanned boat and its estimated heading angle relative to the target unmanned boat; α i,r represents the bow velocity reference signal of the i-th unmanned boat; are the gain coefficients of longitudinal errors; are the gain coefficients of lateral errors; are the gain coefficients of the heading angle error; T3 is the predetermined convergence time; p is a constant in the guidance law error coefficient, and 0<p<1.
6. The method for controlling cooperative target tracking by multiple unmanned vehicles according to claim 5, characterized in that: In S4, the auxiliary system for compensating for input saturation is constructed as follows: Where: is the differential of the auxiliary vector; h βi is an auxiliary vector used to deal with the control input truncation problem caused by input saturation; (·) -1 Indicates the inversion of a matrix or the negation of a function; The speed error of the auxiliary system considering the compensation input saturation is obtained as follows: First, establish the velocity error system differential equation: Where: represents the differential of the velocity error of the auxiliary system taking into account the compensation input saturation; represents the differential of the reference velocity signal vector, in, represents the first-order differential of the longitudinal velocity reference signal of the i-th unmanned boat; represents the first-order differential of the bow velocity reference signal of the i-th unmanned boat; βi represents the auxiliary system number; Secondly, the velocity error system differential equation is sorted out to obtain: Where: h i represents the dynamic matrix; C i (v bi ) represents the Coriolis matrix and the centripetal matrix; D i (v bi ) represents the fluid mechanics damping matrix.
7. The multi-unmanned vehicle cooperative target tracking game control method according to claim 6, characterized in that: In S5, the error performance function is established as follows: Where: represents the hyperbolic tangent function; represents the transformed error vector, φ i -1 Represents φ i is the inverse function of z i represents the speed error of the auxiliary system after the performance function constraint is taken into account when compensating for input saturation; δ i,min represents the lower bound coefficient of the error limit region; δ i,max represents the upper bound coefficient of the error limit region; λ i (t) is the error performance index function; λ i,0 ,λ i,∞ are all constants, and λ i,0 =λ i (0), λ i,∞ =λ i (∞);c i is a positive constant of the error domain convergence speed; t is time; λ i (0) represents the maximum value of the error performance index function; λ i (∞) represents the minimum value of the error performance index function; Furthermore, the changed error vector is obtained through the following transformed velocity error differential equation: Where: represents the error differential after transformation; Represents the transformed dynamic matrix; represents the transformed inertia matrix; γ i Represents the inertia matrix transformation coefficient; Represents λ i The first differential of (t); represents the hyperbolic tangent function.
8. The multi-unmanned vehicle cooperative target tracking game control method according to claim 7, characterized in that: In S6, the optimal cost function is established as follows: Where: represents the optimal cost function; is the optimal control law of the i-th unmanned boat; is the optimal control law of the jth unmanned boat; t0 represents the lower limit of integration; l i represents the discount factor; u represents the integration time; express The transpose of V i is the neighbor node of the i-th unmanned boat; C ij represents the intensity of the control law of the neighboring nodes; T represents the transposition; d u represents the differential of the integration time; The HJB equation is expressed as follows: Where: represents the Hamilton-Jacobi-Bellman equation; η i represents the position and attitude matrix of the i-th unmanned boat; is the gradient of the optimal cost function; V i - represents the set of neighbor nodes; l i represents the discount factor; represents the transformed error vector; Represents the transformed dynamic matrix; Then the following optimal control law is obtained: Where: represents the transformed inertia matrix; T represents transpose; is the optimal control law for the i-th unmanned boat.
9. The method for controlling multi-unmanned boat cooperative target tracking game according to claim 8, characterized in that: The method used to obtain the approximate optimal control law that satisfies Nash equilibrium is as follows: S71: Establishing an evaluation network weight adaptive law to obtain an evaluation network weight adaptive law vector; The evaluation network weight adaptive law is established as follows: Where: represents the adaptive law vector of the evaluation network weight, in, Both are the differentials of the corresponding weights after the hidden layer output of the evaluation network; ΔS i Represents time period T Δ The difference of the Gaussian basis function vector within ; S i represents the Gaussian basis function vector, S i =[S i,1 ,S i,2 ,S i,3 ] T , S i,1 , S i,2 , S i,3 Represent the outputs of the first, second, and third neurons in the hidden layer respectively; x(t) represents the network input vector, c i represents the central constant vector in the hidden layer, c i =[c 1,i , c 2,i , c 3,i ] T , where c 1,i , c 2,i , c 3,i are the center point vectors of the Gaussian basis functions of each neuron in the hidden layer, b i represents the width of the Gaussian basis function; is the evaluation network weight vector, are the corresponding weights after the output of the hidden layer of the evaluation network; β i is the error gain coefficient; Represents time period T Δ The difference vector of the error norm within l i is the discount factor; u is the integration variable; t is the time; T Δ is the time period; Indicates tT Δ Speed error at the moment; represents the speed error at time t; S72: Establish an action network weight adaptation law based on the evaluation network weight adaptation law vector to obtain the action network weight adaptation law vector. The formula used is as follows: Where: represents the action network weight adaptation law vector, in, are the differentials of the corresponding weights after the output of the hidden layer of the action network; k i,a is the gain coefficient; is the action network weight vector, These are the corresponding weights after the output of the hidden layer of the action network; is the basis function gradient vector; ι represents the discount factor in the action network; S73: According to the action network weight adaptive law vector, an approximate optimal control law satisfying Nash equilibrium is obtained as follows: Where: represents the approximate optimal control law that satisfies the Nash equilibrium, and ε is the approximation error.
Citation Information
Patent Citations
Novel discrete time specified performance reinforcement learning unmanned ship course tracking control method and system
CN116400691A
Unmanned ship trajectory tracking control method with global fixed time stability
CN118938663A