Cooperative pursuit escape control method and system for multiple unmanned surface ship systems
By using a reinforcement learning algorithm with a dual-executor network architecture, the policy conflict of multiple unmanned surface vessel systems in non-cooperative scenarios was resolved. This enabled efficient solving of the Nash equilibrium in a zero-sum pursuit-escape game, improving control performance and resource utilization efficiency.
Patent Information
- Application Number
- CN202511201480.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-10-31
AI Technical Summary
When facing non-cooperative targets, existing multi-unmanned surface vessel systems struggle to obtain analytical solutions for traditional zero-sum games. Reinforcement learning methods are prone to policy conflicts, resulting in poor control performance and difficulty in efficiently solving the Nash equilibrium of zero-sum pursuit-escape games.
A reinforcement learning algorithm with a dual-executor network architecture is used to construct an augmented system model and performance index function, design an optimal controller, and solve the Nash equilibrium using zero-sum game theory to ensure that the strategies of the pursuer and evasive ships reach a balance in dynamic confrontation.
It improves strategy convergence performance, alleviates conflict issues in single-executor networks, enables efficient collaborative pursuit by multiple unmanned surface vessel systems in adversarial environments, and optimizes control resource consumption.
Smart Images

Figure CN120871875A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of collaborative control technology for multiple unmanned surface vessels, specifically to a collaborative pursuit and escape control method and system for multiple unmanned surface vessel systems. Background Technology
[0002] In recent years, multi-unmanned surface vessels (MUVs) have been widely used in marine exploration, maritime rescue, and military missions due to their efficient collaborative capabilities in complex tasks. However, in practical applications, these systems often face the threat of uncooperative MUVs moving along non-fixed paths. The dynamic evasion behavior of such targets poses a significant challenge to the collaborative control of MUVs.
[0003] Currently, research on collaborative control of multiple unmanned surface vessels focuses primarily on cooperative scenarios, while insufficient attention is paid to the pursuit and escape problems of non-cooperative targets. Among existing technologies, zero-sum games are an important method for solving pursuit and escape problems; however, solving traditional zero-sum games relies on the Hamilton-Jacobi-Bellman equation, making analytical solutions difficult to obtain. Reinforcement learning offers a new approach to solving such complex game problems, but existing reinforcement learning methods often employ a single-agent network architecture, which is prone to policy conflicts when dealing with conflicting targets of pursuit and avoidance, resulting in lengthy training processes and poor convergence performance.
[0004] Therefore, existing control methods cannot be well integrated in non-cooperative scenarios and in efficiently solving the Nash equilibrium of zero-sum pursuit-escape games, resulting in poor control performance. Summary of the Invention
[0005] To address the aforementioned problems, this invention proposes a collaborative pursuit and escape control method and system for multiple unmanned surface vessels. It designs a control method suitable for non-cooperative scenarios that can efficiently solve the Nash equilibrium of zero-sum pursuit and escape games, effectively mitigating the conflicts generated when a single-actor network handles opposing control targets, improving convergence performance, and ensuring good control performance of the system while achieving collaborative pursuit with minimal control resources.
[0006] According to some embodiments, the present invention adopts the following technical solution:
[0007] A collaborative pursuit and escape control method for multiple unmanned surface vessel systems includes:
[0008] Acquire the communication topology of the multiple unmanned surface vessel systems to be controlled;
[0009] Based on the physical characteristics of the pursuit ship and the evasive ship, a dynamic model of the two is established, and by introducing state variables, it is transformed into a state-space model.
[0010] Based on the state space model and communication topology, an augmented system model integrating the states of the pursuing ship and the evading ship is constructed.
[0011] Based on the augmented system model, an optimal controller with a two-actor network is designed through a zero-sum pursuit-escape game.
[0012] During the control process, a reinforcement learning algorithm with an execution-evaluation network architecture is adopted to solve the Nash equilibrium of the zero-sum pursuit-escape game for the optimal controller, so that the dual executor network can approximate the optimal control strategies of the pursuing ship and the evading ship respectively, ensuring that the strategies of both sides reach a balance in the dynamic confrontation.
[0013] According to some embodiments, the present invention adopts the following technical solution:
[0014] A collaborative pursuit and escape control system for multiple unmanned surface vessels includes:
[0015] The topology acquisition module is configured to acquire the communication topology of the multiple unmanned surface vessel systems to be controlled.
[0016] The state modeling module is configured to: establish dynamic models of the pursuer and the evasive ship based on their physical characteristics, and transform them into state-space models by introducing state variables;
[0017] The augmented modeling module is configured to: construct an augmented system model that integrates the states of the pursuer and the evasive ship based on the state space model and the communication topology;
[0018] The controller design module is configured to: design an optimal controller with a two-actor network based on the augmented system model and a zero-sum pursuit-escape game.
[0019] The control optimization module is configured to: during the control process, use a reinforcement learning algorithm with an execution-evaluation network architecture to solve the Nash equilibrium of the zero-sum pursuit-escape game for the optimal controller, so that the dual executor networks respectively approximate the optimal control strategies of the pursuing ship and the evading ship, ensuring that the strategies of both sides reach a balance in the dynamic confrontation.
[0020] According to some embodiments, the present invention adopts the following technical solution:
[0021] A computer program product includes a computer program that, when executed by a processor, implements the aforementioned collaborative pursuit and escape control method for a multi-unmanned surface vessel system.
[0022] According to some embodiments, the present invention adopts the following technical solution:
[0023] A non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the aforementioned collaborative pursuit and escape control method for a multi-unmanned surface vessel system.
[0024] According to some embodiments, the present invention adopts the following technical solution:
[0025] An electronic device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the collaborative pursuit and escape control method of a multi-unmanned surface vessel system.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0027] 1. The technical solution of this invention employs a simplified reinforcement learning algorithm to design an architecture with a dual-executor network. This architecture can not only independently approximate the optimal strategies of the pursuing and evading ships, but also handle conflicting objectives between the two. Due to the dual-executor network structure, the policy convergence performance is effectively improved, making it more suitable for two-player zero-sum pursuit-escape game scenarios.
[0028] 2. The simplified reinforcement learning algorithm proposed in this invention effectively alleviates the conflict problem in single-agent networks. Compared with traditional single-agent reinforcement learning methods, this algorithm shows superior performance in terms of efficiency and stability in solving Nash equilibrium solutions in zero-sum game pursuit and escape tasks.
[0029] 3. The technical solution of this invention proposes a collaborative control strategy based on the fusion of zero-sum game theory and reinforcement learning. It integrates the advantages of game theory and reinforcement learning, and by constructing an optimal performance index function, it ensures that multiple unmanned surface vessels complete collaborative pursuit missions and optimizes control resource consumption. The core is to use Nash equilibrium theory to improve the collaborative control effect, enabling multiple unmanned surface vessels to operate efficiently and collaboratively in adversarial environments. Attached Figure Description
[0030] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0031] Figure 1 The flowchart illustrates the implementation of the multi-unmanned surface vessel collaborative pursuit and escape control method provided in Example 1.
[0032] Figure 2 This is a communication topology diagram between multiple unmanned surface vessels in Example 1;
[0033] Figure 3 This is a diagram showing the movement trajectories of the pursuing and evasive ships from the start to the point of capture in Example 1;
[0034] Figure 4 For example, in Example 1, the pursuit ship and the evasive ship are in x p y p Error diagram in the φ direction;
[0035] Figure 5 For example, in Example 1, the pursuit ship and the evasive ship are Error plots in the v and r directions;
[0036] Figure 6 This is a control input response curve diagram of the pursuit ship 1 in Example 1;
[0037] Figure 7 This is a control input response curve diagram of the pursuit ship 2 in Example 1;
[0038] Figure 8 This is a control input response curve diagram of the evasive ship in Example 1;
[0039] Figure 9 This is a weight convergence curve of the judge network and the dual enforcer network in reinforcement learning in Example 1. Detailed Implementation
[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0042] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0043] Example 1
[0044] One embodiment of the present invention provides a method for coordinated pursuit and escape control of multiple unmanned surface vessels, comprising:
[0045] Step S1: Obtain the communication topology of the multiple unmanned surface vessel systems to be controlled;
[0046] Step S2: Based on the physical characteristics of the pursuit ship and the evasion ship, establish their dynamic models and transform them into state-space models by introducing state variables.
[0047] Step S3: Based on the state space model and communication topology, construct an augmented system model that integrates the states of the pursuer and the evasive ship;
[0048] Step S4: Based on the augmented system model, design an optimal controller with a two-actor network through a zero-sum pursuit-escape game.
[0049] Step S5: During the control process, a reinforcement learning algorithm with an execution-evaluation network architecture is used to solve the Nash equilibrium of the zero-sum pursuit-escape game for the optimal controller, so that the dual-executor network approximates the optimal control strategies of the pursuing ship and the evading ship respectively, ensuring that the strategies of both sides reach a balance in the dynamic confrontation.
[0050] As one embodiment, this invention provides a cooperative pursuit and escape control method for a multi-unmanned surface vessel system. Using reinforcement learning and zero-sum game theory as its design framework, it addresses the cooperative pursuit and control problem of multi-unmanned surface vessel systems against non-cooperative targets moving along non-fixed paths. First, an augmented system model integrating the states of the pursuing and evasive vessels is constructed. This model can uniformly describe the dynamic interaction between the two, providing an overall analytical framework for cooperative strategy design. Then, a simplified reinforcement learning algorithm with an execution-judgment network architecture is employed to solve for Nash equilibrium based on the minimax principle of zero-sum games. The judge network evaluates control performance, and the dual-executor networks approximate the optimal strategies of the pursuing and evasive vessels respectively, ensuring that the strategies of both sides reach equilibrium in dynamic confrontation. Finally, by combining optimal control technology with a cooperative mechanism, efficient cooperative capture of dynamically evasive targets by multiple pursuing vessels is achieved, ensuring the stability of the system and the optimal utilization of control resources in adversarial environments. Figure 1 As shown, the specific steps are as follows:
[0051] Step 1: Communication topology diagram between multiple unmanned surface vessels as shown below Figure 2 As shown, the system includes n pursuing ships and one evasive ship. In this embodiment, n = 2. The communication topology between multiple unmanned surface vessels in the coordinated pursuit and escape control is described by an undirected graph formed by the pursuing ships and a communication matrix between the pursuing ships and the evasive ship. Specifically:
[0052] Construct an undirected graph in, The set of nodes representing the pursuit ships. It is a side set for communication between pursuit ships. Let be an adjacency matrix, satisfying a ij =a ji . Represents the set of neighbors of the pursuing ship i, when When, the corresponding adjacent element a ij =1, otherwise a ij =0; and undirected graph The relevant Laplace matrix is The in-degree matrix
[0053] The communication matrix between the pursuit ship and the evasive ship is B = diag{b1,b2,…,b n}, b i (i = 1, 2, ..., n) represents the communication between the pursuing ship i and the evasive ship.
[0054] Step 2: Establish the dynamic models of the pursuit ship and the evasive ship as follows:
[0055]
[0056] Where, ξ i =[x pi ,y pi ,φ i ] T Includes positional variable (x) pi ,y pi ) and yaw angle φ i These parameters are relative to the geocentric coordinate system; These represent the yaw, pitch, and roll velocities, described relative to the coordinate system of the unmanned surface vessel; the positive definite inertial matrix M... i satisfy D i E i Representing damping force and mooring force respectively, τ i For control input, the thruster configuration matrix Λ can be represented as: Among them, l i (i = 1, ..., 6) is the yaw moment arm, and l1 = l2 are two symmetrical thrusters. It is the azimuth angle.
[0057] Step 3: Based on the dynamic models of the pursuit ship and the evasive ship, transform them into a state-space model by introducing state variables;
[0058] Introducing state variables and The dynamic models of the pursuer and evasive ships are transformed into state-space equations, which are then rewritten as follows:
[0059]
[0060]
[0061] in, It is the dynamic function of the pursuing ship i. This is the dynamic function of the evasive ship. The control matrices of the pursuing ship i and the evasive ship are expressed as follows: and u i =τ i and u e =τ e These represent the control inputs for the pursuing ship i and the evading ship i, respectively. Since the evading ship can dynamically adjust its motion state in real time, its trajectory is not fixed.
[0062] Step 4: Based on the aforementioned state-space equations, construct an augmented system integrating the states of the pursuing ship and the evading ship as follows:
[0063]
[0064] in, (c i =∑a ij +b i The ) indicates the control gain of the pursuit ship. The controller representing the frigate. This represents the control gain of the evasive ship, d = u e This represents the controller of the evasive ship. For ease of writing, g(e(t)), k(e(t)), and f(e(t)) will be represented by g(e), k(e), and f(e) respectively.
[0065] Step 5: Based on simplified reinforcement learning algorithms and pursuit-escape game theory, design the optimal controller for multiple unmanned surface vessels.
[0066] Simplified reinforcement learning has been studied, and its main idea is based on the following formula (15). This method approximates the performance index function by using the gradient term, rather than directly approximating the performance index itself as in traditional reinforcement learning. This approach avoids the partial gradient differentiation process in subsequent calculations, thereby effectively reducing computational complexity. This embodiment introduces a dual-actor network based on simplified reinforcement learning, where one network executes the behavior of the pursuing ship and the other executes the behavior of the evading ship to deal with the control objective of the conflict, specifically:
[0067] first step:
[0068] The performance metric function is constructed as follows:
[0069]
[0070] Among them, η(e,u,d)=e T (t)Qe(t)+uT Ru-γ 2 d T d is the instantaneous or local cost function for computation. and All are positive definite matrices, and γ represents the predetermined interference attenuation level.
[0071] According to equation (5), the value function can be expressed as follows:
[0072]
[0073] The pursuing ship aims to minimize the value function, while the evasive ship strives to maximize it. This constitutes a zero-sum game, and the optimal value function is:
[0074]
[0075] According to equation (6), the Hamiltonian function can be defined as:
[0076]
[0077] in It is the gradient of V(e) with respect to e(t).
[0078] Applying the stability condition of the Hamiltonian function, the Nash policy of the augmented system is obtained as follows:
[0079]
[0080] Based on equations (9) and (10), the following equation can be derived:
[0081]
[0082] Substituting equations (9) and (10) into equation (8), we obtain the Hamilton-Jacobi-Isaac equation:
[0083]
[0084] Substituting formulas (11) and (12) into formula (13), we get:
[0085]
[0086] Step Two:
[0087] To obtain the Nash policy of the augmented system, the gradient... Decomposed into:
[0088]
[0089] Where ρ is a positive constant, and
[0090] Substituting equation (15) into equations (11) and (12) respectively, we get:
[0091]
[0092] Unknown item J ι (e) It can be approximated using a neural network:
[0093] J ι (e)=ω *T S(e)+ε(e) (18)
[0094] in It is the weight matrix of an ideal neural network. Here, nc represents the number of neurons, and ε is the approximation error, whose value is constrained by a constant δ and satisfies ||ε||≤δ.
[0095] Substituting equation (18) into equations (15)-(17) respectively, we obtain the following results:
[0096]
[0097] However, due to ω * Since the optimal controller in formulas (20) and (21) is unknown, it cannot be obtained. In order to obtain the available optimal controller, a judge neural network and two executor neural networks are constructed based on formulas (19)-(21).
[0098] The following evaluator network is used to estimate control performance:
[0099]
[0100] in, It is the output. These are the weights of the evaluation neural network.
[0101] The following actor neural network is used to implement control behavior:
[0102]
[0103] in, and These are the weights of the executor neural network.
[0104] Step Six: Based on gradient descent and Lyapunov function stability theory, design an adaptive update law for the weights of the execution-evaluation neural network:
[0105]
[0106]
[0107] in, It is the weight of the evaluation neural network. and The weights k of the executor neural network c >0 indicates a parameter for evaluating the design of the home network. and These are the executor network design parameters.
[0108] To demonstrate the feasibility, effectiveness, and correctness of this example, the following simulation experiments were conducted:
[0109] In this simulation experiment, for a multi-unmanned surface vessel system with non-cooperative unmanned surface vessels that have non-fixed path movement characteristics, a cooperative controller based on simplified reinforcement learning and zero-sum game is designed to achieve cooperative pursuit control of the pursuing vessel against the evading vessel.
[0110] In the optimal controller design process, the system model parameters are set as follows:
[0111] The model parameters of the unmanned surface vessel are as follows:
[0112] l1=l2=0.0472, l3=0.4108, l4=0.3858, l5=0.4554, l6=0.3373,
[0113]
[0114] The initial positions of the pursuit ship and the evasive ship are [-10, 10, 0.075, 1, 0, 0.95] respectively. T [-20,15,0.60,0,0,0] T and [-30,-10,0,0,7.03,0] T Both the judge's and the executor's neural networks have 24 nodes, with a central μ. i Uniformly spaced within the range [-2,2], with a width μ i =40; Design parameters k of the judge neural network c =6; initial value set to Design parameters of the executor neural network The initial value is set to Other key parameters are given as Q = I 12×12 , R=diag{1,1,0.1,1,0.1,1,0.1,0.1,0.1,1,1,1}, γ=1.1 and ρ=0.1.
[0115] The effectiveness of this simulation is further illustrated by referring to the accompanying diagram:
[0116] Simulation results are shown below Figure 3-9 .like Figure 3 As shown, the movement trajectories of the two pursuing ships and the evading ship are presented intuitively, clearly demonstrating the coordinated pursuit process of the pursuing ships against the evading ships, and reflecting the coordinated pursuit effect of the multi-unmanned surface vessel system. Figure 4 and Figure 5 The following are given in the surge (x) p ), swing (y p ) and yaw angle (φ) and surge The tracking error between the pursuing ship and the evading ship under yaw (v) and yaw speed (r) can be seen from the figure. The error eventually converges to zero, indicating that the control strategy can effectively achieve accurate tracking of dynamically evading targets. Figure 6 and Figure 7 The control input response curves of Pursuit Ship 1 and Pursuit Ship 2 are shown. Figure 8 The control input response curve of the evasive vessel is presented. The changes in the control input are reasonable and stable, which demonstrates the effectiveness of the controller design. Figure 9 The convergence curves of the weights in the judge neural network and the dual-actor neural network are presented. The figures show that the weights of each network eventually stabilize and converge, proving that the dual-actor network architecture based on a simplified reinforcement learning algorithm can effectively approximate the optimal policy. Therefore, the designed cooperative control scheme can achieve efficient cooperative pursuit of non-cooperative targets by multiple unmanned surface vessels. Simulation results verify the effectiveness of the proposed control scheme.
[0117] In summary, all signals in the system are uniformly bounded, and simulation results demonstrate the effectiveness of the proposed cooperative control scheme.
[0118] Example 2
[0119] One embodiment of the present invention provides a collaborative pursuit and escape control system for multiple unmanned surface vessels, comprising:
[0120] The topology acquisition module is configured to acquire the communication topology of the multiple unmanned surface vessel systems to be controlled.
[0121] The state modeling module is configured to: establish dynamic models of the pursuer and the evasive ship based on their physical characteristics, and transform them into state-space models by introducing state variables;
[0122] The augmented modeling module is configured to: construct an augmented system model that integrates the states of the pursuer and the evasive ship based on the state space model and the communication topology;
[0123] The controller design module is configured to: design an optimal controller with a two-actor network based on the augmented system model and a zero-sum pursuit-escape game.
[0124] The control optimization module is configured to: during the control process, use a reinforcement learning algorithm with an execution-evaluation network architecture to solve the Nash equilibrium of the zero-sum pursuit-escape game for the optimal controller, so that the dual executor networks respectively approximate the optimal control strategies of the pursuing ship and the evading ship, ensuring that the strategies of both sides reach a balance in the dynamic confrontation.
[0125] Example 3
[0126] One embodiment of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned collaborative pursuit and escape control method for multiple unmanned surface vessels.
[0127] Example 4
[0128] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided for storing computer instructions. When the computer instructions are executed by a processor, the method for coordinated pursuit and escape control of multiple unmanned surface vessels is implemented.
[0129] Example 5
[0130] One embodiment of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the aforementioned collaborative pursuit and escape control method for multiple unmanned surface vessels.
[0131] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0132] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0133] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for coordinated pursuit and escape control of multiple unmanned surface vessels, characterized in that, include: Acquire the communication topology of the multiple unmanned surface vessel systems to be controlled; Based on the physical characteristics of the pursuit ship and the evasion ship, a dynamic model of the two is established, and by introducing state variables, it is transformed into a state-space model. Based on the state space model and communication topology, an augmented system model integrating the states of the pursuer and the evasive ship is constructed. Based on the augmented system model, an optimal controller with a two-actor network is designed through a zero-sum pursuit-escape game. During the control process, a reinforcement learning algorithm with an execution-evaluation network architecture is adopted to solve the Nash equilibrium of the zero-sum pursuit-escape game for the optimal controller, so that the dual executor network can approximate the optimal control strategies of the pursuing ship and the evading ship respectively, ensuring that the strategies of both sides reach a balance in the dynamic confrontation.
2. The method for coordinated pursuit and escape control of multiple unmanned surface vessels as described in claim 1, characterized in that, The communication topology among multiple unmanned surface vessels in a coordinated pursuit of escaped control is described using an undirected graph composed of pursuing vessels and a communication matrix between pursuing and evading vessels.
3. The method for coordinated pursuit and escape control of multiple unmanned surface vessels as described in claim 1, characterized in that, The state-space model is expressed by the following formula: Where, Γ(χ i ), Γ(χ e ) are the dynamic functions of the pursuing ship i and the evading ship i, respectively. and The control matrices for the pursuing ship i and the evasive ship are respectively, u i and u e represents the control input quantities of the pursuing ship i and the evading ship respectively, and n is the number of pursuing ships.
4. The method for coordinated pursuit and escape control of multiple unmanned surface vessels as described in claim 1, characterized in that, The augmented system model is expressed by the following formula: Where g(e(t)) represents the control gain of the pursuit ship. The controller of the pursuing ship is represented by k(e(t)), which represents the control gain of the evasive ship, and d = u. e This refers to the controller of the evasive vessel.
5. The method for coordinated pursuit and escape control of multiple unmanned surface vessels as described in claim 1, characterized in that, The design of the optimal controller with a dual-actuator network involves the following steps: The design of the value function is as follows: the goal of the pursuit ship is to minimize the value function, while the goal of the evasion ship is to maximize the value function. Based on the above objectives, construct one evaluation neural network and two executor neural networks; Based on gradient descent and Lyapunov function stability theory, an adaptive update law for the weights of an execution-evaluation neural network is designed to solve for the optimal weights of the neural network and obtain the optimal controller.
6. The method for coordinated pursuit and escape control of multiple unmanned surface vessels as described in claim 1, characterized in that, The reinforcement learning algorithm with an executive-judge network architecture is used to solve the Nash equilibrium of the zero-sum pursuit-escape game for the optimal controller. The control performance is estimated by using a judge network, and the optimal control inputs of the pursuer and evasive ships are solved by two executive neural networks to achieve coordinated pursuit-escape control.
7. A collaborative pursuit and escape control system for multiple unmanned surface vessels, characterized in that, include: The topology acquisition module is configured to acquire the communication topology of the multiple unmanned surface vessel systems to be controlled. The state modeling module is configured to: establish dynamic models of the pursuer and the evasive ship based on their physical characteristics, and transform them into state-space models by introducing state variables; The augmented modeling module is configured to: construct an augmented system model that integrates the states of the pursuer and the evasive ship based on the state space model and the communication topology; The controller design module is configured to: design an optimal controller with a two-actor network based on the augmented system model and a zero-sum pursuit-escape game. The control optimization module is configured to: during the control process, use a reinforcement learning algorithm with an execution-evaluation network architecture to solve the Nash equilibrium of the zero-sum pursuit-escape game for the optimal controller, so that the dual executor networks respectively approximate the optimal control strategies of the pursuing ship and the evading ship, ensuring that the strategies of both sides reach a balance in the dynamic confrontation.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the collaborative pursuit and escape control method of a multi-unmanned surface vessel system as described in any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement a collaborative pursuit and escape control method for a multi-unmanned surface vessel system as described in any one of claims 1-6.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a collaborative pursuit and escape control method for a multi-unmanned surface vessel system as described in any one of claims 1-6.