Multi-unmanned ship cooperative target tracking game control method

By constructing dynamic and kinematic models in the coordinated tracking control of multiple unmanned boats, establishing target state observers and guidance laws for predetermined time, and combining the compensating input saturation and evaluation network adaptive laws, the problems of long convergence time and insufficient interaction of multiple unmanned boats in complex marine environments are solved, and fast convergence and efficient control are achieved.

CN120103871AActive Publication Date: 2025-06-06DALIAN MARITIME UNIVERSITY

Patent Information

Application Number
CN202510588921.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing multi-unmanned boat collaborative tracking control scheme converges for a long time in complex marine environments and fails to fully consider the complex interactions between formation members, resulting in limited task execution efficiency and accuracy.

Method used

By constructing an unmanned boat dynamics and kinematics model that considers marine environmental interference and input saturation, a predetermined time target state observer and a predetermined time guidance law are established, the reference signal of the unmanned boat is obtained, and the adaptive law of the auxiliary system that compensates for input saturation and the evaluation network, an approximate optimal control law that satisfies Nash equilibrium is obtained.

Benefits of technology

It realizes rapid convergence and efficient control of coordinated target tracking of multiple unmanned boats in complex marine environments, reducing the risk of mission delays and improving the overall performance of the formation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103871A_ABST
    Figure CN120103871A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-unmanned-ship cooperative target tracking game control method, and relates to the field of unmanned-ship cooperative tracking control, and the method comprises the steps: obtaining an estimation value of an unmanned ship after the offset of a tracking target through building a preset-time target state observer, obtaining a reference signal of the unmanned ship through a preset-time guidance law, and obtaining a target state estimation value of the unmanned ship; obtaining a speed error vector of the auxiliary system considering compensation input saturation; and then obtaining an optimal control law of the unmanned ships through the established optimal cost function, and obtaining an approximation optimal control law satisfying Nash equilibrium based on an adaptive law of the evaluation network, thereby realizing control of cooperative target tracking of the multiple unmanned ships. Through the establishment of the predetermined time guidance law, the problem of relatively long convergence time is solved, and the dependence of the convergence time on control parameters is reduced. And meanwhile, the influence of neighbor members in the formation is fully considered, a solution meeting Nash equilibrium is obtained by adopting an evaluation network, and the multi-unmanned ship Nash game problem of a nonlinear continuous time system with unknown interference is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cooperative tracking control of unmanned boats, and in particular to a game control method for cooperative target tracking of multiple unmanned boats. Background Art

[0002] With the rapid development of technology, unmanned surface vehicle (USV) technology has evolved from a single platform to a multi-platform collaboration and cluster system. Multi-USV systems have significantly improved the efficiency and success rate of mission execution through information sharing and collaborative decision-making, and have been widely used in fields such as marine resource exploration, environmental monitoring, and near-shore patrols. As one of the key technologies, collaborative tracking control is essential to ensure the effective operation of multi-USV systems.

[0003] In a complex marine environment, USVs are subject to interference from a variety of factors, including direct effects of natural environments such as wind, waves, and currents, as well as challenges such as actuator input saturation. These interference factors not only affect the navigation trajectory, speed, and stability of the USV, but also increase the difficulty of its navigation control and reduce the accuracy and efficiency of task execution. In addition, when multiple USVs perform tracking tasks, they need to collaborate with other USVs to achieve the overall performance of the formation. During the movement, the USV is not only affected by its own control input, but also by the motion state of the surrounding USVs. This interaction is manifested in fluid dynamics as vortices, waves, and currents generated by forces and torques, which may affect the stability of the USV. In order to ensure the safe, stable, and efficient operation of USVs in complex marine environments, it is necessary to comprehensively consider various interference factors and take corresponding design and control strategy optimization measures. Therefore, it is of great significance to solve the problem of multi-unmanned boat tracking control in complex marine environments.

[0004] At present, a lot of research has been done on the problem of cooperative tracking of multiple unmanned vehicles in complex marine environments. However, the following problems still exist in the existing strategies: Based on the currently widely adopted multi-USV target tracking control scheme, the control convergence time of the multi-USV formation system is generally long, which limits its application potential in tasks that require rapid response. Specifically, the existing control strategies often lack a mechanism to pre-set and ensure that the system can reach a stable state or a predetermined formation configuration within a specific time. This lack of convergence speed not only affects the efficiency of formation execution, but may also cause delays in mission execution in some cases.

[0005] In addition, most existing studies tend to view the control problem of each USV in isolation when exploring the control of multi-USV formations, while ignoring the complex interactions and dynamic effects between formation members. Although this simplified approach reduces the complexity of the problem to some extent, it also leads to the limitation of the overall performance of the formation. In fact, each member in a multi-USV formation is not only affected by its own control strategy, but also by the combined effects of the motion state of other members, external environmental factors, and the overall goals of the formation. Therefore, if these interactions are not fully considered, it will be difficult to achieve a significant improvement and optimization of the overall performance of the formation. Summary of the invention

[0006] The present invention provides a multi-unmanned boat cooperative target tracking game control method to overcome the above technical problems.

[0007] A multi-unmanned boat cooperative target tracking game control method comprises the following steps: S1: Construct the dynamic model of the unmanned boat and the kinematic model of the unmanned boat considering the marine environment disturbance and input saturation; S2: establishing a predetermined time target state observer according to the unmanned boat's dynamic model and the unmanned boat's kinematic model to obtain the unmanned boat's estimated value of the tracking target's position and attitude after the deviation and the unmanned boat's estimated value of the tracking target's speed after the deviation; S3: Establishing a predetermined time guidance law to obtain a longitudinal velocity reference signal and a bow velocity reference signal of the unmanned boat based on a differential equation of a position error system according to an estimated value of the unmanned boat after the tracking target position and attitude offset and an estimated value of the unmanned boat after the tracking target velocity offset; S4: establishing an auxiliary system for compensating input saturation, so as to obtain a speed error of the auxiliary system considering compensating input saturation according to a longitudinal speed reference signal and a bow speed reference signal of the unmanned boat; S5: obtaining a transformed error vector based on an error performance function according to a speed error of the auxiliary system considering compensation input saturation; S6: establishing an optimal cost function according to the transformed error vector, and obtaining an optimal control law of the unmanned boat based on the HJB equation; S7: Establish an adaptive law of the evaluation network to obtain an approximate optimal control law that satisfies Nash equilibrium according to the optimal control law of the unmanned boat, so as to realize the control of collaborative target tracking of multiple unmanned boats according to the approximate optimal control law that satisfies Nash equilibrium.

[0008] Beneficial effects: A multi-unmanned boat cooperative target tracking game control method of the present invention obtains the estimated value of the unmanned boat after the tracking target is offset by establishing a predetermined time target state observer, and obtains the reference signal of the unmanned boat through the predetermined time guidance law, and obtains the speed error of the auxiliary system considering the compensation input saturation based on the auxiliary system that compensates for the input saturation, and then obtains the transformed error vector; then the optimal control law of the unmanned boat is obtained through the established optimal cost function, and the approximate optimal control law that satisfies the Nash equilibrium is obtained based on the adaptive law of the evaluation network, so as to realize the control of the cooperative target tracking of multiple unmanned boats. By establishing the predetermined time guidance law, the upper limit of the convergence time can be set in advance to solve the problem of large convergence time and reduce the dependence of the convergence time on the control parameters. At the same time, in order to improve the performance of the entire system, the influence of neighbor members in the formation is fully considered, and the approximate optimal control law that satisfies the Nash equilibrium is obtained by using the evaluation network, so as to obtain the solution that satisfies the Nash equilibrium, and solve the Nash game problem between multiple unmanned boats in a nonlinear continuous-time system with unknown interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0010] Figure 1 is a flow chart of the collaborative target tracking game control method of the present invention; Figure 2 is a communication topology diagram in an embodiment of the present invention; Figure 3 It is a block diagram of the multi-unmanned boat cooperative target tracking game control structure in an embodiment of the present invention; Figure 4 is a tracking effect diagram of three unmanned boats in an embodiment of the present invention; Figure 5 is the evaluation network weight change curve of the three unmanned boats in the embodiment of the present invention; Figure 6 is the action network weight change curve of the three unmanned boats in the embodiment of the present invention; Figure 7 is the speed error curve of USV1 in the embodiment of the present invention; Figure 8 is the speed error curve of USV2 in the embodiment of the present invention; Fig. 9 is the speed error curve of USV3 in the embodiment of the present invention; Fig.10is the longitudinal thrust curve of the three unmanned boats in the embodiment of the present invention; Fig.11 1 is the bow moment curve of the three unmanned boats in the embodiment of the present invention. DETAILED DESCRIPTION

[0011] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0012] This embodiment discloses a multi-unmanned boat cooperative target tracking game control method, including the following steps: Figure 1 As shown: Specifically, this embodiment discusses a multi-unmanned boat cooperative target tracking game control method taking into account environmental interference and mutual influence between formation members.

[0013] S1: Construct the dynamic model of the unmanned boat and the kinematic model of the unmanned boat considering the marine environment disturbance and input saturation; Specifically, in this embodiment, a dynamic model of the unmanned boat is constructed and used as a controlled object, and the position information in the earth coordinate system is obtained through a kinematic model, wherein the dynamic model of the unmanned boat includes external environmental disturbances caused by waves; Preferably, the dynamic model of the unmanned boat is established as follows: (1) The kinematic model of the unmanned boat is established as follows: , Where: Indicates The position and attitude matrix of the unmanned boat, where , , , They represent the longitudinal position coordinate, lateral position coordinate and heading angle of the unmanned boat respectively. represents transpose; Represents the coordinate system rotation matrix; Indicates The velocity matrix of the unmanned boat, where , , They represent longitudinal speed, lateral speed and heading speed respectively; represents the inertia matrix; express The first order differential of ; represents the Coriolis matrix and the centripetal matrix; represents the fluid dynamics damping matrix; represents the external disturbance, where ; They represent longitudinal disturbance, lateral disturbance and bow disturbance respectively; represents a control input with saturation, where , represents the longitudinal thrust with saturation, represents the heading moment with saturation; in: , , , , , in, It is the mass of the unmanned boat; is the longitudinal additional mass coefficient caused by the longitudinal acceleration; is the longitudinal additional mass coefficient caused by the lateral acceleration; is the additional moment of inertia coefficient of the heading caused by the heading acceleration; It indicates that unmanned boats are The moment of inertia of the shaft; is the lateral additional mass coefficient caused by the lateral acceleration; is the longitudinal linear water damping coefficient produced by the longitudinal velocity; is the second-order longitudinal nonlinear water damping coefficient generated by the longitudinal velocity; Indicates absolute value; is the transverse linear water damping coefficient generated by the transverse velocity; is the second-order lateral nonlinear water damping coefficient generated by the lateral velocity; is the linear water damping coefficient of the heading caused by the heading angular velocity; is the second-order nonlinear water damping coefficient of the heading caused by the heading angular velocity; Indicates the heading angle of the unmanned boat; represents the control input, where , The longitudinal thrust, represents the heading moment with saturation; is the control input maximum value vector, where , is the maximum longitudinal thrust, is the maximum heading moment; represents a saturation function.

[0014] S2: Based on the dynamic model and kinematic model of the unmanned boat, considering the situation that the target state is not completely known, a target state observer at a predetermined time is established to obtain the estimated value of the unmanned boat's tracking target position and attitude offset and the estimated value of the unmanned boat's tracking target speed offset, so as to track the estimated value of the target motion state; wherein, the estimated value of the target motion state will be used for the design of the subsequent predetermined time guidance law. The motion state includes the lateral speed, bow speed and longitudinal speed of the unmanned boat; Preferably, the predetermined time target state observer is constructed as follows: (2) Where: It is the first The differential of the estimated vector of the target position and attitude offset by the unmanned boat, where , is the differential of the longitudinal offset position estimate, is the differential of the lateral offset position estimate, is the differential of the offset heading angle estimate; It is the first The differential of the estimated vector of the target speed after the unmanned boat is offset, where , is the second-order differential of the estimated longitudinal position of the tracking target, is the second-order differential of the estimated lateral position of the tracking target, is the second-order differential of the estimated heading angle of the tracking target; It is the first The estimated vector of the target speed after the unmanned boat is offset; , and are the gain coefficients of the position error, velocity error and saturation function in the observer respectively; , The estimated time for position and the estimated time for speed are respectively; , are observer position error index parameters, , are observer velocity error exponential parameters, and , , , ; , are the position error term and the velocity error term respectively; It is The second-order differential of the offset between the unmanned boat and the target, where , It is The second-order differential of the longitudinal offset between the unmanned boat and the target, It is The second-order differential of the lateral offset between the unmanned boat and the target, Indicates The second-order differential of the heading angle offset between the unmanned boat and the target; represents the intermediate calculation parameters, where , is any variable, is a constant; represents a symbolic function; in, (3) Where: , All are index numbers of unmanned boats; represents the total number of unmanned boats; It is The unmanned boat and The weight coefficient between the unmanned boats; It is the first The estimated vector of the target position and attitude after the unmanned boat is offset, where: , It is the first The estimated value of the target longitudinal position after the unmanned boat is offset. It is the first The estimated value of the lateral position offset of the target by the unmanned boat, It is the first The estimated value of the heading angle of the unmanned boat to the target after deviation; It is the first The estimated vector after the target position and attitude deviation of the unmanned boat is estimated. , It is the first The estimated value of the target longitudinal position after the unmanned boat is offset. It is the first The estimated value of the lateral position offset of the target by the unmanned boat, It is the first The estimated value of the heading angle of the unmanned boat to the target after deviation; It is The position and attitude offset between the unmanned boat and the tracking target, , It is The longitudinal position offset between the unmanned boat and the tracking target, It is The offset between the lateral position of the unmanned boat and the tracking target, It is The offset of the heading angle between the unmanned boat and the tracking target; It is The position and attitude offset between the unmanned boat and the tracking target, , It is The longitudinal position offset between the unmanned boat and the tracking target, It is The offset between the lateral position of the unmanned boat and the tracking target, It is The offset of the heading angle between the unmanned boat and the tracking target; It is The weight coefficient between the unmanned boat and the tracking target; is the position and attitude vector of the tracking target, , is the longitudinal position of the tracking target, is the lateral position of the tracking target, is the heading angle of the tracking target; It is the first The estimated value of the target speed after the unmanned boat is offset. , It is The differential of the estimated longitudinal velocity of the unmanned boat to the tracked target, It is The differential of the estimated lateral velocity of the unmanned boat to the tracking target, It is The differential of the estimated value of the heading angular velocity of the tracking target by the unmanned boat; It is The differential of the position and attitude offset between the unmanned boat and the tracking target, , It is The differential of the longitudinal position offset between the unmanned boat and the tracking target, It is The differential of the lateral position offset between the unmanned boat and the tracking target, It is The differential of the heading angle offset between the unmanned boat and the tracking target; It is The differential of the position and attitude offset between the unmanned boat and the tracking target, , It is The differential of the longitudinal position offset between the unmanned boat and the tracking target, It is The differential of the lateral position offset between the unmanned boat and the tracking target, It is The differential of the heading angle offset between the unmanned boat and the tracking target; is the velocity vector of the tracking target in the geodetic coordinate system, , is the differential of the longitudinal position of the tracking target, is the differential of the lateral position of the tracking target, is the differential of the heading angle of the tracking target.

[0015] Specifically, in order to enable all USVs in the formation to obtain the motion state of the tracking target, the motion state of the target USV is observed, and the state observer is constructed using the predetermined time theory to estimate the real-time motion state of the target USV. The subsequent design of the predetermined time guidance law will refer to this state; S3: Establishing a predetermined time guidance law to obtain a longitudinal velocity reference signal and a bow velocity reference signal of the unmanned boat based on a differential equation of a position error system according to an estimated value of the unmanned boat after the tracking target position and attitude offset and an estimated value of the unmanned boat after the tracking target velocity offset; Specifically, the estimated value of the tracking target position and attitude after the unmanned boat is offset and the estimated value of the tracking target speed after the unmanned boat is offset, and the predetermined time guidance law is designed in combination with the motion state of each USV itself, so as to provide a reference longitudinal speed and a reference bow angular velocity for achieving target tracking; the reference longitudinal speed and the reference bow angular velocity will be used in the calculation of the speed error system later.

[0016] Specifically, according to the estimated value of the target USV's motion state after the offset and the motion state of the formation member USV, the position error with the formation member USV is calculated, and the predetermined time guidance law that makes the position error system stable is obtained. The speed error system considering the control input saturation compensation will use this guidance law as the speed reference value for calculation in the subsequent calculation; Preferably, the differential equation of the position error system is expressed as follows: , , (4) Where: Indicates The differential of the error between the real-time longitudinal position of the target unmanned boat and its estimated longitudinal position after the target unmanned boat is offset; Indicates The differential of the error between the real-time lateral position of the target unmanned boat and its estimated lateral position after the displacement of the target unmanned boat; Indicates The differential of the error between the real-time heading angle of the unmanned boat and the estimated heading angle after the unmanned boat is offset from the target unmanned boat; Indicates The position and attitude matrix of the unmanned boats; Indicates the coordinate system of the attached The estimated value of the longitudinal velocity of the target unmanned boat by the unmanned boat, Indicates the body coordinate system The estimated value of the lateral velocity of the target unmanned boat by the unmanned boat, Indicates the coordinate system of the attached The estimated value of the bow speed of the target unmanned boat by the unmanned boat, is the target heading angle; Indicates The error between the real-time heading angle of the unmanned boat and its estimated heading angle after the deviation from the target unmanned boat; Indicates The error between the real-time lateral position of the unmanned boat and its estimated lateral position after the displacement of the target unmanned boat; Indicates The error between the real-time longitudinal position of the target unmanned boat and its estimated longitudinal position after the deviation from the target unmanned boat; Represents the inverse of the coordinate system rotation matrix.

[0017] Preferably, in order to make the position error system converge within a predetermined time, the predetermined time guidance law is established as follows: (5) in, , , Where: Indicates The longitudinal velocity reference signal of the unmanned boat; Indicates The error reference signal between the real-time heading angle of the unmanned boat and its estimated heading angle of the target unmanned boat; Indicates The bow speed reference signal of an unmanned boat; , are the gain coefficients of longitudinal errors; , are the gain coefficients of lateral errors; , are the gain coefficients of heading angle errors; is the scheduled convergence time; is a constant in the guidance law error coefficient, and .

[0018] S4: Establish an auxiliary system for compensating input saturation to obtain the speed error of the auxiliary system considering the compensation input saturation according to the longitudinal speed reference signal and the bow speed reference signal of the unmanned boat, so as to calculate the real-time speed error of each USV; the performance of the speed error system will be specified later.

[0019] Specifically, considering the control input truncation problem caused by input saturation, this embodiment introduces an auxiliary system for compensation, that is, an auxiliary system for compensating input saturation, and integrates it into the real-time speed error between each USV and the target USV, and adjusts the control input in real time to reduce the negative impact caused by input saturation. In order to ensure that the error is effectively limited to a predetermined area, improve the error convergence process, and improve the stability and accuracy of the system, the speed error system will be transformed in performance.

[0020] Preferably, the auxiliary system for compensating input saturation is constructed as follows: (6) Where: is the differential of the auxiliary vector; It is an auxiliary vector used to deal with the control input truncation problem caused by input saturation; Indicates the inversion of a matrix or the negation of a function; Preferably, the speed error of the auxiliary system considering the compensation input saturation is obtained as follows: First, establish the velocity error system differential equation: (7) Where: represents the differential of the speed error of the auxiliary system considering the compensation input saturation; represents the differential of the reference velocity signal vector, ,in, Indicates The first-order differential of the longitudinal velocity reference signal of the unmanned boat; Indicates The first-order differential of the bow velocity reference signal of an unmanned boat; Indicates the auxiliary system number; Arrange the differential equation of the velocity error system to obtain: (8) , Where: Represents a dynamic matrix; represents the Coriolis matrix and the centripetal matrix; represents the fluid mechanics damping matrix.

[0021] S5: obtaining a transformed error vector based on an error performance function according to a speed error of the auxiliary system considering compensation input saturation; Specifically, according to the calculated speed errors of each unmanned boat, the error is transformed in combination with the error performance function and the speed error after the specified performance change is obtained, thereby improving the dynamic process of error convergence; subsequently, the dynamic control law will be designed according to the speed error system.

[0022] Specifically, after obtaining the speed error of the auxiliary system considering the compensation input saturation, the error performance function is used to transform these errors, converting the original speed error into a speed error with specific performance characteristics. The changes of these errors are smoother, thereby improving the speed control accuracy and stability of the unmanned boat, and finding a stable control strategy in the subsequent case of considering the interaction between multiple unmanned boats.

[0023] Preferably, the error performance function is established as follows, and is designed as the following hyperbolic tangent function: , , (9)

[0024] Where: represents the hyperbolic tangent function; represents the transformed error vector, , express is the inverse function of , ; It represents the speed error of the auxiliary system after the speed error of the auxiliary system considering the compensation input saturation is constrained by the performance function; Represents the lower bound coefficient of the error limit region; Represents the upper bound coefficient of the error limit region; is the error performance index function; , are constants, and , ; is a positive constant of the error domain convergence speed; For time; Indicates the maximum value of the error performance index function; Represents the minimum value of the error performance index function; Furthermore, the changed error vector is obtained through the following transformed velocity error differential equation: , , , (10) Where: represents the error differential after transformation; Represents the transformed dynamic matrix; represents the transformed inertia matrix; represents the inertia matrix transformation coefficient; express The first order differential of ; represents the hyperbolic tangent function.

[0025] S6: establishing an optimal cost function according to the transformed error vector, and obtaining an optimal control law of the unmanned boat based on the HJB equation; According to the transformed error vector and considering the control law of the neighboring nodes, that is, j The control law of the unmanned boat is designed, the optimal cost function is designed, and then the Hamilton-Jacobi-Bellman equation (HJB) is designed to calculate the optimal control law of the unmanned boat; Specifically, in order to obtain a control law that satisfies the Nash equilibrium, the speed error of the USV, the control law, and the control law of the neighboring nodes are introduced into the optimal cost function. According to the optimal cost function of this embodiment, the HJB equation can be obtained, and the optimal control law is obtained according to the HJB equation. The optimal cost function constructed in this embodiment is as follows: (11) Where: represents the optimal cost function; It is The optimal control law of an unmanned boat; It is The optimal control law of an unmanned boat; represents the lower limit of the integral; represents the discount factor; Indicates the integration time; express The transpose of It is The neighbor nodes of the unmanned boats; Indicates the strength of the control law of neighboring nodes; represents transpose; represents the differential of the integration time; According to the optimal cost function of this embodiment, an optimal controller that satisfies the HJB equation is designed, and all USVs can achieve Nash equilibrium by implementing the optimal control strategy; Preferably, the Nash equilibrium of the optimal control law of the unmanned boat can be achieved through the optimal cost function, and the method is as follows: Specifically, in order to prove the Nash equilibrium of the optimal control law, this embodiment re-expresses the optimal cost function as follows: (12) Where: represents the deformation of the cost function after deformation; represents the lower bound of the integral; represents the discount factor in the optimal cost function; Indicates the strength of the control law of neighboring nodes; Represents the gradient of the optimal cost function; represents the initial value of the optimal cost function; Represents the value of the optimal cost function at infinite time; Rearranging formula (12), we get: (13) According to the optimal control law of the unmanned boat, , (13) is expressed as: (14) According to the optimal control law of the unmanned boat, , (14) is expressed as: (15) Further expressed as: , (16) in, Indicates The optimal cost function for an unmanned boat to obtain the best response strategy; when , can get , at this time The unmanned boats obtain the best response strategy. When all the unmanned boats obtain the best response strategy, the Nash equilibrium of the optimal control law of the unmanned boats can be achieved.

[0026] Preferably, the HJB equation is expressed as follows: (17) Where: represents the Hamilton–Jacobi–Bellman equation; Indicates The position and attitude matrix of the unmanned boats; is the gradient of the optimal cost function; Represents a set of neighbor nodes; represents the discount factor; represents the transformed error vector; Represents the transformed dynamic matrix; Then the following optimal control law is obtained: (18) Where: represents the transformed inertia matrix; represents transpose; It is Optimal control law for an unmanned boat.

[0027] S7: establishing an adaptive law of the evaluation network to obtain an approximate optimal control law that satisfies Nash equilibrium according to the optimal control law of the unmanned boat; Specifically, this embodiment designs an action-evaluation neural network learning structure to approximate the ideal optimal control law.

[0028] Specifically, the optimal control law of the unmanned boat needs to be solved by the gradient of the optimal cost function of the HJB equation. However, the HJB equation has complex nonlinear characteristics, and it is not feasible to obtain the available optimal solution (18), which brings great difficulties to the control system design. In this case, two neural network-based approximators are deployed to derive the optimal controller that depends on the tracking error signal and can approximate formula (18). The S7 includes giving the evaluation network weight adaptive law based on the game, the action network weight adaptive law, and the approximate optimal control law that satisfies the Nash equilibrium.

[0029] Preferably, the method used to obtain the approximate optimal control law that satisfies the Nash equilibrium is as follows: S71: Establishing an evaluation network weight adaptive law to obtain an evaluation network weight adaptive law vector; The evaluation network of this embodiment adopts a traditional RBF neural network structure, including an input layer, a hidden layer and an output layer. The input layer and the hidden layer each contain three neurons, and the hidden layer contains one neuron.

[0030] Preferably, the evaluation network weight adaptive law is established as follows: (19) , Where: represents the adaptive law vector of the evaluation network weights, ,in, , , They are all the differentials of the corresponding weights after evaluating the output of the hidden layer of the network; Indicates time period The difference of the Gaussian basis function vector within ; , represents the Gaussian basis function vector, , , Represent the outputs of the first, second, and third neurons in the hidden layer respectively; represents the network input vector, represents the central constant vector in the hidden layer, ,in, , , are the center point vectors of the Gaussian basis functions of each neuron in the hidden layer. represents the width of the Gaussian basis function; ; is the evaluation network weight vector, , , , They are the corresponding weights after the output of the hidden layer of the evaluation network; is the error gain coefficient; Indicates time period The difference vector of the error norm within ; is the discount factor; is the integration variable; It’s time; It is the time period; express Speed ​​error at the moment; express Speed ​​error at the moment; S72: According to the evaluation network weight adaptive law vector, establish the action network weight adaptive law to obtain the action network weight adaptive law vector, and the adopted formula is as follows: Preferably, the action network weight adaptation law is as follows: (20) Where: represents the action network weight adaptation law vector, ,in, , , They are all the differentials of the corresponding weights after the output of the hidden layer of the action network; is the gain factor; is the action network weight vector, , , , They are the corresponding weights after the output of the hidden layer of the action network; is the basis function gradient vector; represents the discount factor in the action network; In this embodiment, the RBF neural network structure is also used in the action network; S73: According to the action network weight adaptive law vector, an approximate optimal control law satisfying Nash equilibrium is obtained as follows: , (twenty one) Where: represents the approximate optimal control law that satisfies the Nash equilibrium, is the approximation error.

[0031] Specifically, a specific simulation experiment embodiment of the present invention is as follows: Communication topology of multiple unmanned boat systems Figure 2 As shown in the figure, three unmanned boats are selected to form the formation members, one USV is the tracking target, and the four USVs form a communication system. The model parameters of the unmanned boat are: , , , , , , , , , , The external disturbances in the longitudinal, transverse and bow directions are: , The initial states of the three unmanned boats are expressed as: , , , The motion state of the tracking target is , The target trajectory offset is: , in, It is The second order differential of the position and attitude offset between the unmanned boat and the tracking target; The relevant parameters of the target state scheduled time observer are , , , , ; , are observer position error index parameters, , are observer velocity error exponential parameters, and , , , ; , The estimated time for position and the estimated time for speed are respectively; represents the gain coefficient of the saturation function; , Represents the gain coefficient of position error and velocity error in the observer; The relevant parameters of the performance function are , , , ; Represents the lower bound coefficient of the error limit region; Represents the upper bound coefficient of the error limit region; is a positive constant of the error domain convergence speed; Indicates the maximum value of the error performance index function; Represents the minimum value of the error performance index function; The parameters related to the predetermined time guidance law are: , =0.03, is the scheduled convergence time; is a constant in the guidance law error coefficient, and ; The maximum force and moment for input saturation are , , , ; The initial parameters of the evaluation network and action network weights are , .

[0032] Figure 3 This is a block diagram of the coordinated target tracking game control structure of three unmanned boats. The unmanned boats exchange state information through a wireless communication network. The control system of each unmanned boat consists of a predetermined time guidance law, a predetermined time state observer, an auxiliary system, a performance regulation system, and an evaluation-action network. As the state of the tracking target unmanned boat changes, the predetermined time observer of each unmanned boat estimates its state in real time and calculates the predetermined time guidance law based on the estimated value. Each unmanned boat uses the evaluation-action network to approximate the optimal Nash equilibrium control law, so that the unmanned boat moves according to the speed reference signal.

[0033] Figure 4-Figure 11 is a simulation effect diagram, where Figure 4 The collaborative target tracking effect of multiple unmanned boats is demonstrated, and it can be seen that each unmanned boat can accurately track the target according to the corresponding estimated trajectory. Figure 5 and Figure 6 They are the change curves of the evaluation network weight and the action network weight. It can be seen that the weight curves finally converge. Figure 7, Figure 8 and Fig. 9 These are the longitudinal speed and bow speed error curves of the three unmanned boats. It can be seen that the speed error curves are all within the specified performance area. Fig.10 and Fig.11 The longitudinal thrust curves and bow moment curves of three unmanned boats (USV1, USV2, and USV3) with and without input saturation are shown. The control law with input saturation produces a truncation phenomenon, but since an auxiliary system is introduced into the algorithm of this embodiment, a good control effect is still obtained.

[0034] The multi-unmanned boat cooperative target tracking game control method of this embodiment takes into account the interference of the marine environment, the influence between the formation members, the input saturation and other situations, and combines the predetermined time theory, the game control method and the optimal control theory to design the target tracking control strategy, so as to realize the multi-unmanned boat target tracking in a complex environment.

[0035] Compared with the existing unmanned boat target tracking method, this embodiment takes into account the interference of the marine environment, the influence between the formation members and the control input saturation. In order to solve the problem that the state of the tracking target is not globally known, the state of the target USV is reconstructed by the predetermined time state observer, and each USV obtains the accurate target state within the predetermined time. On this basis, a predetermined time guidance law is designed so that the position error system converges within the predetermined time. Then, the speed error is calculated according to the predetermined time guidance law, and the input saturation compensation system is introduced, and the speed error is performance-regulated. Finally, the optimal control law that satisfies the Nash equilibrium is solved through the evaluation-action network. In summary, the present invention can achieve accurate and rapid tracking of targets by multiple unmanned boats in complex environments.

[0036] The multi-unmanned boat cooperative target tracking game control method of this embodiment establishes a two-layer distributed control structure. The upper control framework uses a predetermined time target state observer to estimate the target information in an environment where the global information of the target is unknown, and formulates a predetermined time guidance law based on this. This ensures that the target estimated position and attitude tracking errors converge to the neighborhood near the origin within a predetermined time. In the lower control framework, in order to solve the control law that satisfies the Nash equilibrium, an online reinforcement learning algorithm is designed, and an evaluation-action neural network (NN) structure is adopted. The control strategy of the adjacent members is included in the objective function of each USV, which is used to solve the Nash game problem of multiple unmanned boats in nonlinear continuous-time systems with unknown interference, and considers speed error constraints and input saturation compensation. This embodiment significantly improves the overall control performance of the unmanned boat formation, enhances the resistance to various interferences, and provides important technical support for the widespread application of unmanned boats.

[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-unmanned boat cooperative target tracking game control method, characterized in that: The steps include: S1: Construct the dynamic model of the unmanned boat and the kinematic model of the unmanned boat considering the marine environment disturbance and input saturation; S2: establishing a predetermined time target state observer according to the unmanned boat's dynamic model and the unmanned boat's kinematic model to obtain the unmanned boat's estimated value of the tracking target's position and attitude after the deviation and the unmanned boat's estimated value of the tracking target's speed after the deviation; S3: Establishing a predetermined time guidance law to obtain a longitudinal velocity reference signal and a bow velocity reference signal of the unmanned boat based on a differential equation of a position error system according to an estimated value of the unmanned boat after the tracking target position and attitude offset and an estimated value of the unmanned boat after the tracking target velocity offset; S4: establishing an auxiliary system for compensating input saturation, so as to obtain a speed error of the auxiliary system considering compensating input saturation according to a longitudinal speed reference signal and a bow speed reference signal of the unmanned boat; S5: obtaining a transformed error vector based on an error performance function according to a speed error of the auxiliary system considering compensation input saturation; S6: establishing an optimal cost function according to the transformed error vector, and obtaining an optimal control law of the unmanned boat based on the HJB equation; S7: Establish an adaptive law of the evaluation network to obtain an approximate optimal control law that satisfies Nash equilibrium according to the optimal control law of the unmanned boat, so as to realize the control of collaborative target tracking of multiple unmanned boats according to the approximate optimal control law that satisfies Nash equilibrium.

2. The method for cooperative target tracking of multiple unmanned boats according to claim 1 is characterized in that: The predetermined time target state observer is constructed as follows: , Where: It is the first The differential of the estimated vector of the target position and attitude after the unmanned boat is offset, where: , is the differential of the longitudinal offset position estimate, is the differential of the lateral offset position estimate, is the differential of the offset heading angle estimate; It is the first The differential of the estimated vector of the target speed after the unmanned boat is offset, where , is the second-order differential of the estimated longitudinal position of the tracking target, is the second-order differential of the estimated lateral position of the tracking target, is the second-order differential of the estimated heading angle of the tracking target; It is the first The estimated vector of the target speed after the unmanned boat is offset; , and are the gain coefficients of the position error, velocity error and saturation function in the observer respectively; , The estimated time for position and the estimated time for speed are respectively; , are observer position error index parameters, , are observer velocity error index parameters, and , , , ; , are the position error term and the velocity error term respectively; It is The second-order differential of the offset between the unmanned boat and the target, where , It is The second-order differential of the longitudinal offset between the unmanned boat and the target, It is The second-order differential of the lateral offset between the unmanned boat and the target, Indicates The second-order differential of the heading angle offset between the unmanned boat and the target; represents the intermediate calculation parameters, where , is any variable, is a constant; represents a symbolic function; in, , Where: , All are index numbers of unmanned boats; represents the total number of unmanned boats; It is The unmanned boat and The weight coefficient between the unmanned boats; It is the first The estimated vector of the target position and attitude after the unmanned boat is offset, where: , It is the first The estimated value of the target longitudinal position after the unmanned boat is offset. It is the first The estimated value of the lateral position offset of the target by the unmanned boat, It is the first The estimated value of the heading angle of the unmanned boat to the target after deviation; It is the first The estimated vector after the target position and attitude deviation of the unmanned boat is estimated. , It is the first The estimated value of the target longitudinal position after the unmanned boat is offset. It is the first The estimated value of the lateral position offset of the target by the unmanned boat, It is the first The estimated value of the heading angle of the unmanned boat to the target after deviation; It is The position and attitude offset between the unmanned boat and the tracking target, , It is The longitudinal position offset between the unmanned boat and the tracking target, It is The offset between the lateral position of the unmanned boat and the tracking target, It is The offset of the heading angle between the unmanned boat and the tracking target; It is The position and attitude offset between the unmanned boat and the tracking target, , It is The longitudinal position offset between the unmanned boat and the tracking target, It is The offset between the lateral position of the unmanned boat and the tracking target, It is The offset of the heading angle between the unmanned boat and the tracking target; It is The weight coefficient between the unmanned boat and the tracking target; is the position and attitude vector of the tracking target, , is the longitudinal position of the tracking target, is the lateral position of the tracking target, is the heading angle of the tracking target; It is the first The estimated value of the target speed after the unmanned boat is offset. , It is The differential of the estimated longitudinal velocity of the unmanned boat to the tracked target, It is The differential of the estimated lateral velocity of the unmanned boat to the tracking target, It is The differential of the estimated value of the heading angular velocity of the tracking target by the unmanned boat; It is The differential of the position and attitude offset between the unmanned boat and the tracking target, , It is The differential of the longitudinal position offset between the unmanned boat and the tracking target, It is The differential of the lateral position offset between the unmanned boat and the tracking target, It is The differential of the heading angle offset between the unmanned boat and the tracking target; It is The differential of the position and attitude offset between the unmanned boat and the tracking target, , It is The differential of the longitudinal position offset between the unmanned boat and the tracking target, It is The differential of the lateral position offset between the unmanned boat and the tracking target, It is The differential of the heading angle offset between the unmanned boat and the tracking target; is the velocity vector of the tracking target in the geodetic coordinate system, , is the differential of the longitudinal position of the tracking target, is the differential of the lateral position of the tracking target, is the differential of the heading angle of the tracking target.

3. The method for cooperative target tracking game control of multiple unmanned boats according to claim 1 is characterized in that: The dynamic model of the unmanned boat is established as follows: , The kinematic model of the unmanned boat is established as follows: , Where: Indicates The position and attitude matrix of the unmanned boat, where , , , They represent the longitudinal position coordinate, lateral position coordinate and heading angle of the unmanned boat respectively. represents transpose; Represents the coordinate system rotation matrix; Indicates The velocity matrix of the unmanned boat, where , , They represent longitudinal speed, lateral speed and heading speed respectively; represents the inertia matrix; express The first order differential of represents the Coriolis matrix and the centripetal matrix; represents the fluid dynamics damping matrix; represents the external disturbance, where ; They represent longitudinal disturbance, lateral disturbance and bow disturbance respectively; represents a control input with saturation, where , represents the longitudinal thrust with saturation, represents the heading moment with saturation; in: , , , , , in, It is the mass of the unmanned boat; is the longitudinal additional mass coefficient caused by the longitudinal acceleration; is the longitudinal additional mass coefficient caused by the lateral acceleration; is the additional moment of inertia coefficient of the heading caused by the heading acceleration; It indicates that unmanned boats are The moment of inertia of the shaft; is the lateral additional mass coefficient caused by the lateral acceleration; is the longitudinal linear water damping coefficient produced by the longitudinal velocity; is the second-order longitudinal nonlinear water damping coefficient generated by the longitudinal velocity; Indicates absolute value; is the transverse linear water damping coefficient generated by the transverse velocity; is the second-order lateral nonlinear water damping coefficient generated by the lateral velocity; is the linear water damping coefficient of the heading caused by the heading angular velocity; is the second-order nonlinear water damping coefficient of the heading caused by the heading angular velocity; Indicates the heading angle of the unmanned boat; represents the control input, where , The longitudinal thrust, represents the heading moment with saturation; is the control input maximum value vector, where , is the maximum longitudinal thrust, is the maximum heading moment; represents a saturation function.

4. The method for controlling a multi-unmanned boat coordinated target tracking game according to claim 1, characterized in that: The differential equation of the position error system is expressed as follows: , , , Where: Indicates The differential of the error between the real-time longitudinal position of the target unmanned boat and its estimated longitudinal position after the target unmanned boat is offset; Indicates The differential of the error between the real-time lateral position of the target unmanned boat and its estimated lateral position after the displacement of the target unmanned boat; Indicates The differential of the error between the real-time heading angle of the unmanned boat and the estimated heading angle after the unmanned boat is offset from the target unmanned boat; Indicates The position and attitude matrix of the unmanned boats; Indicates the coordinate system of the attached The estimated value of the longitudinal velocity of the target unmanned boat by the unmanned boat, Indicates the body coordinate system The estimated value of the lateral velocity of the target unmanned boat by the unmanned boat, Indicates the coordinate system of the attached The estimated value of the bow speed of the target unmanned boat by the unmanned boat, is the target heading angle; Indicates The error between the real-time heading angle of the unmanned boat and its estimated heading angle after the deviation from the target unmanned boat; Indicates The error between the real-time lateral position of the unmanned boat and its estimated lateral position after the displacement of the target unmanned boat; Indicates The error between the real-time longitudinal position of the target unmanned boat and its estimated longitudinal position after the deviation from the target unmanned boat; Represents the inverse of the coordinate system rotation matrix.

5. The method for controlling a multi-unmanned boat cooperative target tracking game according to claim 1, characterized in that: The predetermined time guidance law is established as follows: , in, , , Where: Indicates The longitudinal velocity reference signal of the unmanned boat; Indicates The error reference signal between the real-time heading angle of the unmanned boat and its estimated heading angle of the target unmanned boat; Indicates The bow speed reference signal of an unmanned boat; , are the gain coefficients of longitudinal errors; , are the gain coefficients of lateral errors; , are the gain coefficients of heading angle errors; is the scheduled convergence time; is a constant in the guidance law error coefficient, and .

6. The method for controlling a multi-unmanned boat coordinated target tracking game according to claim 1, characterized in that: In S4, the auxiliary system for compensating input saturation is constructed as follows: , Where: is the differential of the auxiliary vector; It is an auxiliary vector used to deal with the control input truncation problem caused by input saturation; Indicates the inversion of a matrix or the negation of a function; The speed error of the auxiliary system considering the compensation input saturation is obtained as follows: First, establish the velocity error system differential equation: , Where: represents the differential of the speed error of the auxiliary system considering the compensation input saturation; represents the differential of the reference velocity signal vector, ,in, Indicates The first-order differential of the longitudinal velocity reference signal of the unmanned boat; Indicates The first-order differential of the bow velocity reference signal of an unmanned boat; Indicates the auxiliary system number; Secondly, the velocity error system differential equation is sorted out to obtain: , , Where: Represents a dynamic matrix; represents the Coriolis matrix and the centripetal matrix; represents the fluid mechanics damping matrix.

7. The method for controlling a multi-unmanned boat cooperative target tracking game according to claim 1, characterized in that: In S5, the error performance function is established as follows: , , , Where: represents the hyperbolic tangent function; represents the transformed error vector, , express is the inverse function of , ; It represents the speed error of the auxiliary system after the speed error of the auxiliary system considering the compensation input saturation is constrained by the performance function; Represents the lower bound coefficient of the error limit region; Represents the upper bound coefficient of the error limit region; is the error performance index function; , are constants, and , ; is a positive constant of the error domain convergence speed; For time; Indicates the maximum value of the error performance index function; Represents the minimum value of the error performance index function; Furthermore, the changed error vector is obtained through the following transformed velocity error differential equation: , , , , Where: represents the error differential after transformation; Represents the transformed dynamic matrix; represents the transformed inertia matrix; represents the inertia matrix transformation coefficient; express The first order differential of represents the hyperbolic tangent function.

8. The method for controlling a multi-unmanned boat coordinated target tracking game according to claim 1, characterized in that: In S6, the optimal cost function is established as follows: , Where: represents the optimal cost function; It is The optimal control law of an unmanned boat; It is The optimal control law of an unmanned boat; represents the lower limit of the integral; represents the discount factor; Indicates the integration time; express The transpose of It is The neighbor nodes of the unmanned boats; Indicates the strength of the control law of neighboring nodes; represents transpose; represents the differential of the integration time; The HJB equation is expressed as follows: , Where: represents the Hamilton–Jacobi–Bellman equation; Indicates The position and attitude matrix of the unmanned boats; is the gradient of the optimal cost function; Represents a set of neighbor nodes; represents the discount factor; represents the transformed error vector; Represents the transformed dynamic matrix; Then the following optimal control law is obtained: , Where: represents the transformed inertia matrix; represents transpose; It is Optimal control law for an unmanned boat.

9. The method for controlling a multi-unmanned boat cooperative target tracking game according to claim 1, characterized in that: The method used to obtain the approximate optimal control law that satisfies Nash equilibrium is as follows: S71: Establishing an evaluation network weight adaptive law to obtain an evaluation network weight adaptive law vector; The evaluation network weight adaptive law is established as follows: , Where: represents the adaptive law vector of the evaluation network weights, ,in, , , They are all the differentials of the corresponding weights after evaluating the output of the hidden layer of the network; Indicates time period The difference of the Gaussian basis function vector within ; , represents the Gaussian basis function vector, , , Represent the outputs of the first, second, and third neurons in the hidden layer respectively; represents the network input vector, represents the central constant vector in the hidden layer, ,in, , , are the center point vectors of the Gaussian basis functions of each neuron in the hidden layer. represents the width of the Gaussian basis function; ; is the evaluation network weight vector, , , , They are the corresponding weights after the output of the hidden layer of the evaluation network; is the error gain coefficient; Indicates time period The difference vector of the error norm within ; is the discount factor; is the integration variable; It’s time; It is the time period; express Speed ​​error at the moment; express Speed ​​error at the moment; S72: According to the evaluation network weight adaptive law vector, establish the action network weight adaptive law to obtain the action network weight adaptive law vector, and the adopted formula is as follows: , Where: represents the action network weight adaptation law vector, ,in, , , They are all the differentials of the corresponding weights after the output of the hidden layer of the action network; is the gain factor; is the action network weight vector, , , , They are the corresponding weights after the output of the hidden layer of the action network; is the basis function gradient vector; represents the discount factor in the action network; S73: According to the action network weight adaptive law vector, an approximate optimal control law satisfying Nash equilibrium is obtained as follows: , , Where: represents the approximate optimal control law that satisfies the Nash equilibrium, is the approximation error.

10. The method for controlling a multi-unmanned boat cooperative target tracking game according to claim 1, characterized in that: Through the optimal cost function, the Nash equilibrium of the optimal control law of the unmanned boat can be achieved, and the proof method is as follows: The optimal cost function is reformulated as follows: , Where: represents the deformation of the cost function after deformation; represents the lower bound of the integral; represents the discount factor in the optimal cost function; Indicates the strength of the control law of neighboring nodes; Represents the gradient of the optimal cost function; represents the initial value of the optimal cost function; Represents the value of the optimal cost function at infinite time; After sorting, we get: , According to the optimal control law of the unmanned boat, ,but: , According to the optimal control law of the unmanned boat, ,but: , Further expressed as: , , in, Indicates The optimal cost function for an unmanned boat to obtain the best response strategy; when , can get , at this time The unmanned boats obtain the best response strategy. When all the unmanned boats obtain the best response strategy, the Nash equilibrium of the optimal control law of the unmanned boats can be achieved.

Citation Information

Patent Citations

  • Unmanned ship path tracking active disturbance rejection control method based on sideslip angle compensation

    CN111580523A

  • Unmanned ship path tracking control method based on sliding mode control

    CN116048078A

  • Novel discrete time specified performance reinforcement learning unmanned ship course tracking control method and system

    CN116400691A

  • Unmanned ship trajectory tracking control method with global fixed time stability

    CN118938663A

  • Method and device for shoreline segmentation in complex environments based on the perspective of an unmanned surface vessel (USV)

    US20250139930A1

Cited By

  • Unmanned ship surrounding hierarchical game control method under unreliable communication

    CN121934396A

  • Unmanned surface vehicle encircle surrounding hierarchical game control method under unreliable communication

    CN121934396B

  • Zero-sum game self-triggering anti-interference reinforcement learning method and device for unmanned ship

    CN122151484A

  • Heterogeneous multi-unmanned ship event triggering formation control method under intermittent communication

    CN122450136A

  • Method for heterogeneous multi-unmanned surface vehicle event-triggered formation control under intermittent communication

    CN122450136B