Multi-UUV collaborative hunting strategy based on behavioral cloning and hierarchical reinforcement learning
Through the multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning, the adaptability and robustness problems of UUV collaborative control in complex marine environments are solved, efficient multi-stage mission planning and execution are achieved, and the collaborative capture capability of UUV clusters is improved.
Patent Information
- Application Number
- CN202510992539.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-18
AI Technical Summary
Existing UUV collaborative control methods have problems such as poor scalability, weak adaptability and susceptibility to single point failures when facing complex and changeable marine environments. In addition, there is a lack of in-depth research on the strategy optimization process, resulting in poor system robustness and adaptability.
A multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning was adopted. By establishing a three-dimensional ocean current model, a semi-Markov decision process and a reward function design, combined with a hierarchical reinforcement learning framework, high-level strategies were optimized and converted into specific control instructions, realizing flexible switching of multi-stage tasks and collaborative capture.
It significantly improves the robustness, scalability and adaptability of UUV collaborative capture missions to complex underwater environments, achieves efficient task decomposition and execution, and enhances the stability and adaptability of the system.
Smart Images

Figure CN120491501B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of UUV cluster capture technology, and in particular to a multi-UUV collaborative capture strategy based on behavior cloning and hierarchical reinforcement learning. Background Art
[0002] The widespread use of unmanned underwater vehicles (UUVs) in ocean exploration and military applications has placed higher demands on the coordinated control and strategy optimization of UUVs. A series of studies have been conducted domestically and internationally in the field of multi-UUV coordinated control, including model-based control methods, learning-based control methods, and hybrid control methods. These studies have improved the collaborative performance of UUVs to a certain extent, but most lack in-depth research on the strategy optimization process, and their adaptability and robustness in actual marine environments still need to be improved. Therefore, traditional centralized control methods suffer from poor scalability, weak adaptability, and susceptibility to single points of failure when faced with complex and changing marine environments.
[0003] UUVs face a series of difficulties and challenges in the collaborative capture process. First, the uncertainty and dynamic changes of the marine environment place extremely high demands on the adaptability of UUV collaborative strategies. The collaborative mechanism between UUVs is also key to achieving effective collaborative capture. Second, many research strategies focus solely on the capture itself, failing to fully consider the phased nature of the capture mission. This oversight reduces the adaptability and depth of the strategies. Therefore, how to overcome the limitations of existing capture strategies and design feasible strategy optimization methods based on information technology, especially machine learning, is the key to effectively improving the level of UUV collaborative capture. Summary of the Invention
[0004] In view of the above-mentioned problems that the current UUV collaborative control and strategy optimization are easily affected by the uncertainty and dynamic changes of the marine environment, resulting in poor system robustness, scalability and adaptability to complex underwater environments, the present invention provides a multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning.
[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0006] The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning includes the following steps: S1. Establish a coordinate system with the UUV buoyancy center as the origin and initialize the various physical property parameters of the UUV; S2. Deploy the UUV cluster and obtain the linear velocity, acceleration and Euler angle of each UUV about the X, Y and Z axes; S3. Establish a three-dimensional ocean current model based on viscous lamb vortex to simulate the dynamic characteristics of the real underwater environment and obtain the position control equation of the UUV; S4. Arrange the deployed UUV cluster in an equilateral triangle formation; S5. Model the group state and target state parameters of the UUV cluster through a semi-Markov decision process and output a high-level strategy; S6. Convert the high-level strategy into specific control commands; S7. Introduce behavioral cloning technology in the hierarchical reinforcement learning framework and optimize the high-level strategy using an expert example dataset; S8. Design a reward function based on the encirclement effect, proximity and synchronization indicators of the UUV cluster to guide the update of the reinforcement learning strategy; S9. Construct and simulate the collaborative capture mission of multiple UUVs in a complex underwater environment in a simulation platform.
[0007] Furthermore, in S1, the buoyancy center of the UUV is set as the origin of the fixed coordinate system, and the center of gravity position of the UUV is determined. and the center of buoyancy , and initialize the physical properties of the UUV: mass m, volume V, surface area A, and minimum turning radius R.
[0008] Furthermore, in S2, there are three UUVs in the UUV cluster, and the actual linear velocity of each UUV about the X, Y, and Z axes is obtained. 、 、 With actual acceleration 、 、 , and get the rolling angle of each UUV , pitch angle , yaw angle .
[0009] Furthermore, in S3, first, a three-dimensional ocean current model based on viscous Lamb's vortex is established, and the expression for simulating the dynamic characteristics of the underwater ocean is:
[0010] ;
[0011] Where, Expressed as the velocity component of the water flow along the X axis; Expressed as the velocity component of the water flow along the Y axis; Expressed as the velocity component of the water flow along the Z axis; Expressed as a yaw-angle-dependent intensity parameter (used to describe vortex intensity or disturbance amplitude); Represented as the coordinate of the current position on the X-axis; Expressed as the coordinate of the initial point on the X-axis; Represented as the coordinate of the current position on the Y axis; Expressed as the coordinate of the initial point on the Y axis; Expressed as the distance from the current position to the vortex center; It is expressed as the starting radius of the vortex influence range; Expressed as the base of natural logarithms; It is expressed as a scale parameter that controls the rate at which the eddy spreads or decays;
[0012] Secondly, define the main forces acting on UUV The expression is:
[0013] ;
[0014] Where, Expressed as fluid density; Expressed as the drag coefficient; It is represented as the UUV forward area, and the data is determined by hydrodynamic simulation; It is expressed as the velocity multiplied by its modulus (modulus, that is, the size of the velocity vector); Expressed as a comprehensive resistance constant or parameter, representing the coefficient of the combined term on the right side of the formula (used to simplify the expression);
[0015] Finally, the position control equation of UUV in six degrees of freedom is obtained as follows:
[0016] ;
[0017] Where, It is expressed as the sum of external forces acting on the UUV along the X-axis; It is expressed as the sum of external forces acting on the UUV along the Y-axis; It is expressed as the sum of external forces acting on the UUV along the Z-axis; Expressed as the sum of the rolling moments around the X axis; Expressed as the sum of the rolling moments around the Y axis; Expressed as the sum of the rolling moments about the Z axis.
[0018] Furthermore, in S4, first, the distance between any two adjacent UUVs among the three UUVs is defined as , the coordinates of one UUV are , the coordinates of the other UUV are , and 、 Satisfies the expression:
[0019] ;
[0020] Secondly, calculate the relative phase angle between any two adjacent UUVs , relative phase angle Used to ensure that each UUV is in the correct direction and relative position in the equilateral triangle formation, relative phase angle The expression is:
[0021] .
[0022] Furthermore, in S5 , first, the group state and target state parameters of the UUV cluster are input, including the relative position and three-dimensional velocity of each UUV in the UUV cluster and the current position and velocity of the target UUV;
[0023] Secondly, through the semi-Markov decision process modeling, the round-up strategy is selected by optimizing the cumulative reward function and defining the parameter expected cumulative return , the parameter expected cumulative return expression is:
[0024] ;
[0025] Where, Expressed as a discount factor; Expressed as time steps; Indicates the maximum duration of policy execution; Expressed as time step The current state of the Expressed as time step Execution behavior under Indicates the current state Execution behavior Rewards received; Expressed as mathematical expectation, that is, the average cumulative reward under multiple strategy executions;
[0026] Then, a high-level policy selection is performed, at each time step Based on the current status , select an optimal execution behavior ;
[0027] Finally, output the high-level strategy : Where, Indicates the current state; It is expressed as an act of execution; Indicates the current state Execution behavior Rewards received; Indicates the variable value at which the expression reaches its maximum value.
[0028] Furthermore, in S6, the high-level strategies of lurking, intercepting, encircling, and tracking are converted into specific control instructions;
[0029] In the control method of the latent stage, the Lyapunov guidance vector field is introduced and the Lyapunov function is defined. and its time derivative ,function and its time derivative The expression is:
[0030] ;
[0031] ;
[0032] Where, Expressed as The current position of each UUV; Represented as the target UUV position; is the convergence coefficient, ;
[0033] In order to avoid being detected by the target UUV, an obstacle avoidance item was added , obstacle avoidance item The expression is:
[0034] ;
[0035] Where, Expressed as a coefficient greater than zero, used to adjust the obstacle avoidance item The weight or strength of ; It is expressed as the UUV position within the detection range of the target UUV; Expressed as the detection radius of the target UUV;
[0036] The control method of the interception phase is to adjust the speed , making Always keep constant speed The expression is:
[0037] ;
[0038] Where, Expressed as a proportional coefficient (a control parameter that adjusts the speed variation); It is represented by the actual distance between the current UUV and the target UUV. ;
[0039] The control method of the encirclement phase is to adjust the heading angle deviation and speed modifiers , to achieve the ideal encirclement azimuth ;
[0040] Heading angle deviation The expression is: );
[0041] Where, Represents the current moment; Expressed as The current heading angle of each UUV; Expressed as heading angle control gain coefficient;
[0042] Speed modifier The expression is: );
[0043] Where, It is expressed as the proportional coefficient in speed control; Expressed as The current distance between each UUV and the target; Expressed as the expected target distance;
[0044] The control method in the tracking phase defines the tracking error , by adjusting the speed and heading angle to minimize the tracking error , tracking error The expression is: ;
[0045] Where, Represents the current moment No. The distance from the UUV to the target; Represents the current moment No. The heading angle of the UUV.
[0046] Furthermore, in S7, behavior cloning technology is introduced into the hierarchical reinforcement learning framework. By initializing the high-level policy network with expert example datasets, it accelerates policy convergence, reduces exploration time, and improves the robustness and adaptability of the policy.
[0047] The steps to introduce and optimize behavior cloning into high-level strategies include the following two aspects:
[0048] First, we design a behavioral cloning model. First, we define a high-level policy network. The function form of the high-level policy network is: Where, Represents the behavioral decision output by the high-level policy network; Represented as parameters of the high-level policy network;
[0049] Secondly, define the expert strategy function, which is in the form of: Where, Represents an expert in the current state Optimal decision-making behavior under
[0050] Finally, define the loss function , and define the expert example dataset , including state-behavior pairs , by minimizing the loss function , so that the high-level policy network can imitate the expert decision-making behavior, the loss function The expression is:
[0051] ;
[0052] Where, Represented as a pair of expert example datasets The state-behavior pairs sampled from expected value;
[0053] The second aspect is the joint optimization of behavior cloning and hierarchical reinforcement learning. First, based on the behavior cloning pre-training, combined with the hierarchical reinforcement learning objectives, a joint loss function is designed. , the goal is to maximize the cumulative reward and define the hierarchical reinforcement learning loss function , joint loss function The expression is:
[0054] ;
[0055] Where, satisfy , represents the balance between controlling hierarchical reinforcement learning and behavior cloning weights;
[0056] Secondly, optimize the target by adjusting Implementing switching between hierarchical reinforcement learning and behavior cloning.
[0057] Furthermore, in S8, the reward function is obtained The steps are:
[0058] Considering the surrounding effect, proximity and synchronization, the reward function is defined by weighted summation , the reward function The expression is:
[0059] ;
[0060] Where, Expressed as the weight of the encirclement effect; Expressed as the weight of the proximity; expressed as the weight of synchronicity; It is expressed as a siege effect bonus; Expressed as proximity reward; Expressed as a synchronization reward;
[0061] Encirclement Effect Bonus It is used to ensure that the UUV cluster forms a stable encirclement around the target UUV at a uniform angle. The expression is:
[0062] ;
[0063] Where, It represents the number of each of the three UUVs that perform the encirclement mission; Expressed as the time step The actual azimuth angle of each UUV relative to the target UUV;
[0064] Proximity Reward It is used to ensure that the distance between the UUV and the target UUV is close to the ideal value, avoiding being too close or too far. The expression is:
[0065] ;
[0066] Where, It is expressed as the actual distance between the UUV and the target UUV; Expressed as the ideal enclosing radius;
[0067] Synchronicity Rewards Used to ensure the coordination and consistency between the UUVs in the UUV cluster, synchronize the action completion time, and the synchronized UUV cluster receives positive rewards , otherwise no reward, synchronization reward The expression is:
[0068] ;
[0069] Where, 、 It represents the time it takes for the UUV to complete the action; Indicates the allowed time synchronization error threshold.
[0070] Furthermore, in S9, a dynamic scene of a UUV cluster consisting of three UUVs collaboratively encircling a target UUV was constructed on the Unity3D simulation platform to simulate the collaborative encirclement and capture mission of multiple UUVs in a complex underwater environment.
[0071] The beneficial effects of the present invention are as follows: the present invention decomposes the complex collaborative capture task into high-level strategy planning and low-level control execution through a hierarchical reinforcement learning framework, achieving efficient task decomposition and execution, and the high-level strategy is responsible for overall task planning, such as lurking, intercepting, encircling, and tracking, while the low-level control is responsible for specific execution commands. This hierarchical structure significantly improves the robustness and adaptability of the system. The present invention combines behavioral cloning with hierarchical reinforcement learning, quickly learning the decision-making behavior of experts through behavioral cloning technology, providing an initial strategy for the HRL framework, thereby accelerating the learning process and improving the stability of the capture strategy. The present invention proposes a distributed collaborative capture strategy that utilizes multiple autonomous UUVs to collaboratively encircle a target without centralized control, significantly improving the system's robustness, scalability, and adaptability to complex underwater environments. The present invention achieves flexible switching between multi-stage tasks of lurking, intercepting, encircling, and tracking through a hierarchical decision-making framework. This multi-stage task control method can adapt to different environmental dynamics and target behaviors, significantly improving the adaptability and effectiveness of UUV collaborative capture tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 Shown is a flow chart of the present invention.
[0073] Figure 2 Shown is a schematic diagram of the hierarchical mission's latent phase, interception phase, encirclement phase, and confrontation tracking phase.
[0074] Figure 3 Shown is a diagram showing the changes in the navigation speed and heading angle output of the UUV1 in the present invention.
[0075] Figure 4 Shown is a diagram showing the changes in the navigation speed and heading angle output of the UUV2 in the present invention.
[0076] Figure 5 The above is a diagram showing changes in the navigation speed and heading angle output of the UUV3 in the present invention.
[0077] Figure 6 Shown is a graph showing the changes in distance and angle output of UUV1 in the present invention.
[0078] Figure 7 Shown is a graph showing the changes in distance and angle output of UUV2 in the present invention.
[0079] Figure 8 Shown is a graph showing the changes in distance and angle output of the UUV3 in the present invention.
[0080] Figure 9 Shown is a diagram showing the changing process of the motion trajectories of UUV1, UUV2, UUV3 and the target UUV in the present invention. DETAILED DESCRIPTION
[0081] The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning includes the following steps: S1. Establish a coordinate system with the UUV buoyancy center as the origin and initialize the various physical property parameters of the UUV; S2. Deploy the UUV cluster and obtain the linear velocity, acceleration and Euler angle of each UUV about the X, Y and Z axes; S3. Establish a three-dimensional ocean current model based on viscous lamb vortex to simulate the dynamic characteristics of the real underwater environment and obtain the position control equation of the UUV; S4. Arrange the deployed UUV cluster in an equilateral triangle formation; S5. Model the group state and target state parameters of the UUV cluster through a semi-Markov decision process and output a high-level strategy; S6. Convert the high-level strategy into specific control commands; S7. Introduce behavioral cloning technology in the hierarchical reinforcement learning framework and optimize the high-level strategy using an expert example dataset; S8. Design a reward function based on the encirclement effect, proximity and synchronization indicators of the UUV cluster to guide the update of the reinforcement learning strategy; S9. Construct and simulate the collaborative capture mission of multiple UUVs in a complex underwater environment in a simulation platform.
[0082] Furthermore, in S1, the buoyancy center of the UUV is set as the origin of the fixed coordinate system, and the center of gravity position of the UUV is determined. and the center of buoyancy , and initialize the physical properties of the UUV: mass m, volume V, surface area A, and minimum turning radius R.
[0083] Furthermore, in S2, there are three UUVs in the UUV cluster, and the actual linear velocity of each UUV about the X, Y, and Z axes is obtained. 、 、 With actual acceleration 、 、 , and get the rolling angle of each UUV , pitch angle , yaw angle .
[0084] Furthermore, in S3, first, a three-dimensional ocean current model based on viscous Lamb's vortex is established, and the expression for simulating the dynamic characteristics of the underwater ocean is:
[0085] ;
[0086] Where, Expressed as the velocity component of the water flow along the X axis; Expressed as the velocity component of the water flow along the Y axis; Expressed as the velocity component of the water flow along the Z axis; Expressed as a yaw-angle-dependent intensity parameter (used to describe vortex intensity or disturbance amplitude); Represented as the coordinate of the current position on the X-axis; Expressed as the coordinate of the initial point on the X-axis; Represented as the coordinate of the current position on the Y axis; Expressed as the coordinate of the initial point on the Y axis; Expressed as the distance from the current position to the vortex center; It is expressed as the starting radius of the vortex influence range; Expressed as the base of natural logarithms; It is expressed as a scale parameter that controls the rate at which the eddy spreads or decays;
[0087] Secondly, define the main forces acting on UUV The expression is:
[0088] ;
[0089] Where, Expressed as fluid density; Expressed as the drag coefficient; It is represented as the UUV forward area, and the data is determined by hydrodynamic simulation; It is expressed as the velocity multiplied by its modulus (modulus, that is, the size of the velocity vector); Expressed as a comprehensive resistance constant or parameter, representing the coefficient of the combined term on the right side of the formula (used to simplify the expression);
[0090] Finally, the position control equation of UUV in six degrees of freedom is obtained as follows:
[0091] ;
[0092] Where, It is expressed as the sum of external forces acting on the UUV along the X-axis; It is expressed as the sum of external forces acting on the UUV along the Y-axis; It is expressed as the sum of external forces acting on the UUV along the Z-axis; Expressed as the sum of the rolling moments around the X axis; Expressed as the sum of the rolling moments around the Y axis; Expressed as the sum of the rolling moments about the Z axis.
[0093] Furthermore, in S4, first, the distance between any two adjacent UUVs among the three UUVs is defined as , the coordinates of one UUV are , the coordinates of the other UUV are , and 、 Satisfies the expression:
[0094] ;
[0095] Secondly, calculate the relative phase angle between any two adjacent UUVs , relative phase angle Used to ensure that each UUV is in the correct direction and relative position in the equilateral triangle formation, relative phase angle The expression is:
[0096] .
[0097] Furthermore, in S5 , first, the group state and target state parameters of the UUV cluster are input, including the relative position and three-dimensional velocity of each UUV in the UUV cluster and the current position and velocity of the target UUV;
[0098] Secondly, through the semi-Markov decision process modeling, the round-up strategy is selected by optimizing the cumulative reward function and defining the parameter expected cumulative return , the parameter expected cumulative return expression is:
[0099] ;
[0100] Where, Expressed as a discount factor; Expressed as time steps; Indicates the maximum duration of policy execution; Expressed as time step The current state of the Expressed as time step Execution behavior under Indicates the current state Execution behavior Rewards received; Expressed as mathematical expectation, that is, the average cumulative reward under multiple strategy executions;
[0101] Then, a high-level policy selection is performed, at each time step Based on the current status , select an optimal execution behavior ;
[0102] Finally, output the high-level strategy :
[0103] ;
[0104] Where, Indicates the current state; It is expressed as an act of execution; Indicates the current state Execution behavior Rewards received; Indicates the variable value at which the expression reaches its maximum value.
[0105] Furthermore, in S6, the high-level strategies of lurking, intercepting, encircling, and tracking are converted into specific control instructions;
[0106] In the control method of the latent stage, the Lyapunov guidance vector field is introduced and the Lyapunov function is defined. and its time derivative ,function and its time derivative The expression is:
[0107] ;
[0108] ;
[0109] Where, Expressed as The current position of each UUV; Represented as the target UUV position; is the convergence coefficient, ;
[0110] In order to avoid being detected by the target UUV, an obstacle avoidance item was added , obstacle avoidance item The expression is:
[0111] ;
[0112] Where, Expressed as a coefficient greater than zero, used to adjust the obstacle avoidance item The weight or strength of ; It is expressed as the UUV position within the detection range of the target UUV; Expressed as the detection radius of the target UUV;
[0113] The control method of the interception phase is to adjust the speed , making Always keep constant speed The expression is:
[0114] ;
[0115] Where, Expressed as a proportional coefficient (a control parameter that adjusts the speed variation); It is represented by the actual distance between the current UUV and the target UUV. ;
[0116] The control method of the encirclement phase is to adjust the heading angle deviation and speed modifiers , to achieve the ideal encirclement azimuth ;
[0117] Heading angle deviation The expression is:
[0118] );
[0119] Where, Represents the current moment; Expressed as The current heading angle of each UUV; Expressed as heading angle control gain coefficient;
[0120] Speed modifier The expression is:
[0121] );
[0122] Where, It is expressed as the proportional coefficient in speed control; Expressed as The current distance between each UUV and the target; Expressed as the expected target distance;
[0123] The control method in the tracking phase defines the tracking error , by adjusting the speed and heading angle to minimize the tracking error , tracking error The expression is:
[0124] ;
[0125] Where, Represents the current moment No. The distance from the UUV to the target; Represents the current moment No. The heading angle of the UUV.
[0126] Furthermore, in S7, behavior cloning technology is introduced into the hierarchical reinforcement learning framework. By initializing the high-level policy network with expert example datasets, it accelerates policy convergence, reduces exploration time, and improves the robustness and adaptability of the policy.
[0127] The steps to introduce and optimize behavior cloning into high-level strategies include the following two aspects:
[0128] First, we design a behavioral cloning model. First, we define a high-level policy network. The function form of the high-level policy network is:
[0129] ;
[0130] Where, Represents the behavioral decision output by the high-level policy network; Represented as parameters of the high-level policy network;
[0131] Secondly, define the expert strategy function, which is in the form of:
[0132] ;
[0133] Where, Represents an expert in the current state Optimal decision-making behavior under
[0134] Finally, define the loss function , and define the expert example dataset , including state-behavior pairs , by minimizing the loss function , so that the high-level policy network can imitate the expert decision-making behavior, the loss function The expression is:
[0135] ;
[0136] Where, Represented as a pair of expert example datasets The state-behavior pairs sampled from expected value;
[0137] The second aspect is the joint optimization of behavior cloning and hierarchical reinforcement learning. First, based on the behavior cloning pre-training, combined with the hierarchical reinforcement learning objectives, a joint loss function is designed. , the goal is to maximize the cumulative reward and define the hierarchical reinforcement learning loss function , joint loss function The expression is:
[0138] ;
[0139] Where, satisfy , represents the balance between controlling hierarchical reinforcement learning and behavior cloning weights;
[0140] Secondly, optimize the target by adjusting Implementing switching between hierarchical reinforcement learning and behavior cloning.
[0141] Furthermore, in S8, the reward function is obtained The steps are:
[0142] Considering the surrounding effect, proximity and synchronization, the reward function is defined by weighted summation , the reward function The expression is:
[0143] ;
[0144] Where, Expressed as the weight of the encirclement effect; Expressed as the weight of the proximity; expressed as the weight of synchronicity; It is expressed as a siege effect bonus; Expressed as proximity reward; Expressed as a synchronization reward;
[0145] Encirclement Effect Bonus It is used to ensure that the UUV cluster forms a stable encirclement around the target UUV at a uniform angle. The expression is:
[0146] ;
[0147] Where, It represents the number of each of the three UUVs that perform the encirclement mission; Expressed as the time step The actual azimuth angle of each UUV relative to the target UUV;
[0148] Proximity Reward It is used to ensure that the distance between the UUV and the target UUV is close to the ideal value, avoiding being too close or too far. The expression is:
[0149] ;
[0150] Where, It is expressed as the actual distance between the UUV and the target UUV; Expressed as the ideal enclosing radius;
[0151] Synchronicity Rewards Used to ensure the coordination and consistency between the UUVs in the UUV cluster, synchronize the action completion time, and the synchronized UUV cluster receives positive rewards , otherwise no reward, synchronization reward The expression is:
[0152] ;
[0153] Where, 、 It represents the time it takes for the UUV to complete the action; Indicates the allowed time synchronization error threshold.
[0154] Furthermore, in S9, a dynamic scene of a UUV cluster consisting of three UUVs collaboratively encircling a target UUV was constructed on the Unity3D simulation platform to simulate the collaborative encirclement and capture mission of multiple UUVs in a complex underwater environment.
[0155] like Figure 1 As shown in Figure 2, the multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning specifically includes the following steps:
[0156] The first step is to determine the basic physical parameters of the UUV. Establish a coordinate system with the UUV's buoyancy center as the origin, set the UUV's buoyancy center as the fixed coordinate system origin, and determine the UUV's center of gravity position. and the center of buoyancy , and initialize the physical properties of the UUV: mass m, volume V, surface area A, and minimum turning radius R. These physical property parameters are used for the kinematic modeling of the UUV and the UUV hydrodynamic simulation of the underwater environment.
[0157] The second step is to summarize the required UUV sampling data. The scenario set in this sea trial is that the enemy deploys a target UUV to sneak into our waters for reconnaissance or sabotage missions. We must quickly identify and track the actions of the enemy target UUV, and then use precise interception strategies and distributed collaborative actions to prevent the enemy target UUV from getting closer to the base. Therefore, we deployed a UUV cluster, which includes three UUVs, numbered UUV1, UUV2 and UUV3. A distributed collaborative strategy is adopted between each UUV, and through high-level coordination and collaboration, the enemy target UUV can be effectively surrounded and intercepted. The present invention uses global situational awareness to feedback the actual linear velocity of each UUV in the UUV cluster about the X, Y, and Z axes. 、 、 With actual acceleration 、 、 , and get the rolling angle of each UUV , pitch angle , yaw angle .
[0158] The third step is to initialize the UUV kinematic model in the underwater environment. In order to improve the practicality and robustness of the distributed cooperative encirclement strategy, the invention first establishes a three-dimensional ocean current model based on viscous lambda vortices. The expression for simulating the dynamic characteristics of the real ocean circle underwater is:
[0159] ;
[0160] Where, Expressed as the velocity component of the water flow along the X axis; Expressed as the velocity component of the water flow along the Y axis; Expressed as the velocity component of the water flow along the Z axis; Expressed as a yaw-angle-dependent intensity parameter (generally used to describe vortex intensity or disturbance amplitude); Represented as the coordinate of the current position on the X-axis; Expressed as the coordinate of the initial point on the X-axis; Represented as the coordinate of the current position on the Y axis; Expressed as the coordinate of the initial point on the Y axis; It is expressed as the starting radius of the vortex influence range; Expressed as the base of natural logarithms, it is approximately equal to 2.718; It is expressed as a scale parameter that controls the rate at which the eddy spreads or decays;
[0161] Secondly, define the main forces acting on UUV The expression is: ;
[0162] Where, Expressed as fluid density; Expressed as the drag coefficient; It is represented as the UUV forward area, and the data is determined by hydrodynamic simulation; It is expressed as the velocity multiplied by its modulus (modulus, that is, the size of the velocity vector); Expressed as a comprehensive resistance constant or parameter, representing the coefficient of the combined term on the right side of the formula (used to simplify the expression);
[0163] Finally, the position control equation of UUV in six degrees of freedom is obtained as follows:
[0164] ;
[0165] Where, It is expressed as the sum of external forces acting on the UUV along the X-axis; It is expressed as the sum of external forces acting on the UUV along the Y-axis; It is expressed as the sum of external forces acting on the UUV along the Z-axis; Expressed as the sum of the rolling moments around the X axis; Expressed as the sum of the rolling moments around the Y axis; Expressed as the sum of the rolling moments around the Z axis; Expressed as the mass of the UUV; It is expressed as the actual linear velocity of the UUV about the X axis; Expressed as the first derivative of the actual linear velocity about the X-axis; Expressed as the actual linear velocity of the UUV about the Y axis; Expressed as the first derivative of the actual linear velocity about the Y axis; Expressed as the actual linear velocity of the UUV about the Z axis; Expressed as the first derivative of the actual linear velocity about the Z axis; Expressed as the rolling angle of the UUV; Expressed as the first derivative of the roll angle; Expressed as the pitch angle of the UUV; Expressed as the first derivative of the pitch angle; Expressed as yaw angle; Expressed as the first derivative of the yaw angle.
[0166] The "viscous lamb vortex," also known as the Lamb-Oseen vortex, is a classic vortex solution to the Navier-Stokes equations for incompressible viscous fluids, describing isolated vortices that gradually decay due to viscous diffusion. In fluid mechanics, the Navier-Stokes equations for incompressible viscous fluids are used to describe and predict the transformations of velocity and pressure fields as they evolve over time and space. They are the fundamental equations that bridge the gap between microscopic viscous effects and macroscopic flow patterns. In the third step of this implementation, we generalize this analytical model along the Z-axis to construct a vortex flow field with three-dimensional velocity components.
[0167] The "viscous lamb vortex" has the following advantages: 1. It is analytically closed, has few parameters, and has explicit formulas for both velocity and vorticity. It is used to rapidly generate flow fields and verify the accuracy of numerical algorithms. Without the need for a CFD (computational fluid dynamics) grid, flow fields can be generated within milliseconds within a simulation loop, enabling high-frequency iterations of hierarchical reinforcement learning and behavioral cloning. 2. It has adjustable parameters, with independent control of scale parameters and core radius. It is easy to construct disturbance fields of multiple intensities and scales to test robustness, capable of simulating both deep-sea micro-vortices and the low-disturbance background of inland harbors. 3. It provides a three-dimensional closed-form expression. The generated external torque is directly embedded in the six-degree-of-freedom dynamic equations, and the resulting external torque is embedded in the sum of the external torques acting on the UUV, ensuring consistency in position-attitude coupling. 4. It can realistically reproduce the energy decay of the vortex around the target UUV, providing controllable difficulty during the stages of lurking, interception, and encirclement.
[0168] In summary, the "viscous lamb vortex" not only ensures physical credibility, but also allows the simulation environment and learning framework to achieve a balance between speed, repeatability, and parameterization diversity, providing an efficient and verifiable three-dimensional ocean current model foundation for the multi-UUV collaborative capture strategy proposed in this implementation.
[0169] The advantages of this implementation of a 3D ocean current model based on "viscous lamb vortices" are: 1. It is well-suited to the mission scenario. Local shear and lateral disturbances are the most sensitive factors in multi-UUV capture. "Viscous lamb vortices" provide local vortices of controllable intensity, representing both the low-disturbance background flow found in actual harbors or canyons and simulating complex flow fields by superimposing multiple vortex cores, thus meeting the "low-disturbance, high-fidelity" requirements of simulation experiments. 2. It has low computational complexity. Compared to full CFD methods, "viscous lamb vortices" are at least two orders of magnitude faster. They provide an analytical closed-form 3D velocity field, eliminating the need for grid integration and can be integrated with hierarchical reinforcement learning and behavioral cloning loops on a single GPU, achieving convergence within minutes. 3. It has smooth gradients. While Rankine vortices are simpler to analyze, they are often used in ideal inviscid flows, where velocity jumps occur at the core radius. "Viscous lamb vortices" avoid Rankine-like shear jumps, facilitating analytical differentiation of Lyapunov vector fields and other methods in this implementation, resulting in more stable control. 4. Robustness: Although uniform or linear shear flow offers the simplest formulas and fastest calculations, the lack of local vortices and shear gradients makes it difficult to test the robustness of the trapping strategy to disturbances. The "viscous lamb vortex" employed in this implementation facilitates the construction of disturbance fields of varying intensities and scales for robustness testing.
[0170] In short, other models are either too computationally intensive or their physical or numerical characteristics do not meet the requirements of long-term coordinated capture; only the "viscous lamb vortex" can simultaneously take into account physical realism, analytical derivability and real-time performance. Therefore, this embodiment has significant advantages in using the "viscous lamb vortex" to establish a three-dimensional ocean current model.
[0171] The fourth step is to build a cooperative observation formation of UUVs. The equilateral triangle formation is used as the main observation structure. First, the distance between any two adjacent UUVs among the three UUVs is defined as , the coordinates of one UUV are , the coordinates of the other UUV are , all This is equal to the preset side length of 50 meters to maintain the equilateral triangle structure.
[0172] and 、 Satisfies the expression: ;
[0173] Secondly, calculate the relative phase angle between any two adjacent UUVs , relative phase angle Used to ensure that each UUV is in the correct direction and relative position in the equilateral triangle formation, relative phase angle The expression is: .
[0174] Step 5: High-level decision design and strategy optimization. First, input the group state and target state parameters of the UUV cluster, including the relative position and three-dimensional velocity of each UUV in the UUV cluster and the current position and velocity of the target UUV;
[0175] Secondly, through the semi-Markov decision process modeling, the round-up strategy is selected by optimizing the cumulative reward function and defining the parameter expected cumulative return , the parameter expected cumulative return expression is: ;
[0176] Where, Expressed as a discount factor; Expressed as time steps; Indicates the maximum duration of policy execution; Expressed as time step The current state of the Expressed as time step Execution behavior under Indicates the current state Execution behavior Rewards received; Expressed as mathematical expectation, that is, the average cumulative reward under multiple strategy executions;
[0177] Then, a high-level policy selection is performed, at each time step Based on the current status , select an optimal execution behavior ;
[0178] Finally, output the high-level strategy , and pass it to the low-level execution network to generate detailed execution control commands, high-level strategies for: ;
[0179] Where, Indicates the current state; It is expressed as an act of execution; Indicates the current state Execution behavior Rewards received; Indicates the variable value at which the expression reaches its maximum value.
[0180] The sixth step is to implement the low-level execution network design and control. The high-level strategies of lurking, intercepting, encircling and tracking are converted into specific control instructions, such as Figure 2 As shown in the figure, the green line represents UUV1, the red line represents UUV2, the yellow line represents UUV3, and the blue line represents the target UUV.
[0181] In the control method of the latent stage, the Lyapunov guidance vector field (LGVF) is introduced and the Lyapunov function is defined. and its time derivative ,function and its time derivative The expression is: ;
[0182] ;
[0183] Where, Expressed as The current position of each UUV; Represented as the target UUV position; is the convergence coefficient, .
[0184] In order to avoid being detected by the target UUV, an obstacle avoidance item was added , obstacle avoidance item The expression is:
[0185] ;
[0186] Where, Expressed as a coefficient greater than zero, used to adjust the obstacle avoidance item The weight or strength of ; It is expressed as the UUV position within the detection range of the target UUV; Expressed as the detection radius of the target UUV;
[0187] The control method of the interception phase is to adjust the speed , making Always keep constant speed The expression is:
[0188] Where, Expressed as a proportional coefficient (a control parameter that adjusts the speed variation); It is represented by the actual distance between the current UUV and the target UUV. ;
[0189] The control method of the encirclement phase is to adjust the heading angle deviation and speed modifiers , to achieve the ideal encirclement azimuth ;
[0190] Heading angle deviation The expression is: ); where Represents the current moment; Expressed as The current heading angle of each UUV; Expressed as heading angle control gain coefficient;
[0191] Speed modifier The expression is: ); where It is expressed as the proportional coefficient in speed control; Expressed as The current distance between each UUV and the target; Expressed as the expected target distance;
[0192] The control method in the tracking phase defines the tracking error , by adjusting the speed and heading angle to minimize the tracking error , tracking error The expression is: Where, Represents the current moment No. The distance from the UUV to the target; Represents the current moment No. The heading angle of the UUV.
[0193] The seventh step is to introduce and optimize behavioral cloning in high-level decision-making. Imitation Learning (IL) technology is introduced into the hierarchical reinforcement learning (HRL) framework. By initializing the high-level decision network using expert example datasets, it accelerates policy convergence, reduces exploration time, and improves policy robustness and adaptability.
[0194] The steps to introduce and optimize behavior cloning into high-level strategies include the following two aspects:
[0195] First, we design a behavioral cloning model. First, we define a high-level policy network. The function form of the high-level policy network is: Where, Represents the behavioral decision output by the high-level policy network; Represented as parameters of the high-level policy network.
[0196] Secondly, define the expert strategy function, which is in the form of: Where, Represents an expert in the current state Optimal decision-making behavior under
[0197] Finally, define the loss function , and define the expert example dataset , including state-behavior pairs , by minimizing the loss function , so that the high-level policy network can imitate the expert decision-making behavior, the loss function The expression is: Where, Represented as a pair of expert example datasets The state-behavior pairs sampled from expected value;
[0198] The second aspect is the joint optimization of behavior cloning and hierarchical reinforcement learning. First, based on the behavior cloning pre-training, combined with the hierarchical reinforcement learning objectives, a joint loss function is designed. , the goal is to maximize the cumulative reward and define the hierarchical reinforcement learning loss function , joint loss function The expression is: Where, satisfy , represents the balance between controlling hierarchical reinforcement learning and behavior cloning weights;
[0199] Secondly, optimize the target by adjusting Implementing switching between hierarchical reinforcement learning and behavior cloning.
[0200] The present invention introduces behavioral cloning technology into high-level decision networks to achieve an organic combination of expert knowledge and reinforcement learning, providing a stable and efficient solution for strategy optimization in multi-stage tasks.
[0201] Combine Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 as well as Figure 7 As shown in the figure, the blue line PPO represents a traditional reinforcement learning algorithm, and the green line HRL-IL represents the hierarchical reinforcement learning (HRL) and behavior cloning (IL) combined algorithm of the present invention, and Figure 5 、 Figure 6 as well as Figure 7The 120° red dotted line in the figure represents the ideal phase angle between the three UUVs after they form an encirclement. As can be seen from the above figure, the present invention has the following advantages over the traditional reinforcement learning algorithm PPO: the present invention accelerates strategy initialization by introducing behavioral cloning technology, and combines the hierarchical reinforcement learning framework to achieve efficient task decomposition and execution, effectively improving the performance of multi-UUV collaborative encirclement tasks. The present invention introduces expert example data sets through behavioral cloning technology, shortens the exploration time of reinforcement learning, greatly accelerates the convergence process of the strategy, provides a stable starting point for reinforcement learning, and reduces fluctuations in the early stages of training. The hierarchical decision-making framework proposed in the present invention can effectively decompose complex tasks, allowing UUVs to flexibly switch between different task stages such as lurking, intercepting, encirclement, and tracking, and can adapt to different environmental dynamics and target behaviors. It also comprehensively considers the reward function of encirclement effect, proximity, and synchronization, so that UUVs can achieve multi-objective balance during task execution and form a stable and efficient encirclement formation.
[0202] Step 8: Reward function design and multi-objective optimization. Taking into account the encirclement effect, proximity and synchronization, the reward function is defined in the form of weighted summation. , the reward function The expression is: Where, Expressed as the weight of the encirclement effect; Expressed as the weight of the proximity; expressed as the weight of synchronicity; It is expressed as a siege effect bonus; Expressed as proximity reward; Expressed as a synchronization reward;
[0203] Encirclement Effect Bonus It is used to ensure that the UUV cluster forms a stable encirclement around the target UUV at a uniform angle. The expression is: Where, It represents the number of each of the three UUVs that perform the encirclement mission; Expressed as the time step The actual azimuth angle of each UUV relative to the target UUV;
[0204] Proximity Reward It is used to ensure that the distance between the UUV and the target UUV is close to the ideal value, avoiding being too close or too far. The smaller the distance deviation, the higher the reward, encouraging each UUV to maintain an appropriate encirclement distance. The expression is: Where, It is expressed as the actual distance between the UUV and the target UUV; Expressed as the ideal enclosing radius;
[0205] Synchronicity Rewards Used to ensure the coordination and consistency between the UUVs in the UUV cluster, synchronize the action completion time, and the synchronized UUV cluster will receive positive rewards , otherwise there is no reward, the expression is: Where, Expressed as time steps; 、 It represents the time it takes for the UUV to complete the action; Indicated as positive reward; Indicates the allowed time synchronization error threshold.
[0206] Step 9: Figure 9 As shown in the figure, UUV1, UUV2, and UUV3 in the UUV cluster of the present invention are encircling a target UUV. Using the Unity3D simulation platform, a dynamic scenario was constructed in which three friendly UUVs collaborated to encircle a single enemy UUV, simulating the collaborative capture mission of multiple UUVs in a complex underwater environment. This validated the feasibility, stability, and efficiency of the system's capture strategy, and optimized the algorithm strategy and system parameters based on the experimental results. To further explore the feasibility and performance of this strategy in real-world scenarios and verify their ability to accurately detect targets and effectively collaborate, it was necessary to ensure high-precision simulation of the virtual simulation platform system.
[0207] Example: The first step is to determine the basic physical parameters of the UUV. Establish a coordinate system with the UUV's buoyancy center as the origin, and set the UUV's buoyancy center as the fixed coordinate system origin to determine the UUV's center of gravity position. and the center of buoyancy , and initialize the physical properties of the UUV: mass m = 349.47 kg, volume V = 0.34 m³, surface area A = 3.64 m², and minimum turning radius R = 22.5 m. These physical property parameters are used for the kinematic modeling of the UUV and the UUV hydrodynamic simulation of the underwater environment.
[0208] The second step is to summarize the required UUV sampling data. The scenario set in this sea trial is that the enemy deploys a target UUV to sneak into our waters for reconnaissance or sabotage missions. We must quickly identify and track the actions of the enemy target UUV, and then use precise interception strategies and distributed collaborative actions to prevent the enemy target UUV from getting closer to the base. Therefore, we deployed a UUV cluster, which includes three UUVs, numbered UUV1, UUV2 and UUV3. A distributed collaborative strategy is adopted between each UUV, and through high-level coordination and collaboration, the enemy target UUV can be effectively surrounded and intercepted. The present invention uses global situational awareness to feedback the actual linear velocity of each UUV in the UUV cluster about the X, Y, and Z axes. 、 、 With actual acceleration 、 、 , and get the rolling angle of each UUV , pitch angle , yaw angle .
[0209] The third step is to initialize the UUV kinematic model in the underwater environment.
[0210] Set the vortex center coordinates ( , ) is (200m, 150m), select any point in the vortex ( , ) is (250m, 200m), =1.5m 2 / s, =50m, =100m, , we get:
[0211] ;
[0212] The obtained ocean current disturbance values are within the typical range of low-disturbance background ocean currents such as "underwater lower layer", "deep sea micro-eddies" or "inland harbors" in the real ocean.
[0213] Secondly, define the main forces acting on UUV The expression is: ;
[0214] Where, Expressed as fluid density; Expressed as the drag coefficient; It is expressed as the forward area of the UUV and the projected area of the UUV in the water-facing direction; the data is determined by hydrodynamic simulation; Expressed as the velocity of the UUV relative to the water flow; It is expressed as the velocity multiplied by its modulus (modulus, that is, the size of the velocity vector); Expressed as a comprehensive resistance constant or parameter, representing the coefficient of the combined term on the right side of the formula (used to simplify the expression);
[0215] The speed of UUV in the uniform state is ,but ;
[0216] , , , ,have to N. As Figure 9As shown in Figure 3, initially, the UUV is traveling at a constant speed along the Y-axis during the latent phase, and the resistance in the Y-axis direction is the largest.
[0217] Finally, the position control equation of UUV in six degrees of freedom is obtained as follows: 、 、 、 、 、 ;in, It is expressed as the sum of the external forces acting on the UUV along the X-axis, i.e., -0.088N; Expressed as the sum of the external forces acting on the UUV along the Y-axis, i.e. -114.624N; It is expressed as the sum of the external forces acting on the UUV along the Z-axis, which is 0.002N; Expressed as the sum of the rolling moments around the X axis; Expressed as the sum of the rolling moments around the Y axis; Expressed as the sum of the rolling moments around the Z axis; It is the net force acting at the center of mass, so by default no additional torque is generated.
[0218] Solve the dynamic differential equations simultaneously to obtain the time evolution results of UUV speed, attitude, trajectory, etc., such as Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 as well as Figure 9 shown.
[0219] The fourth step is to build a cooperative observation formation of UUVs. The equilateral triangle formation is used as the main observation structure. First, the distance between any two adjacent UUVs among the three UUVs is defined as , the coordinates of one UUV are , the coordinates of the other UUV are , all This is equal to the preset side length of 50 meters to maintain the equilateral triangle structure.
[0220] and 、 Satisfies the expression: ;when When it can be maintained, a collaborative structure of an equilateral triangle is formed.
[0221] Secondly, calculate the relative phase angle between any two adjacent UUVs , relative phase angle Used to ensure that each UUV is in the correct direction and relative position in the equilateral triangle formation, relative phase angle The expression is: .when When it can be maintained, it further verifies the formation of an equilateral triangle collaborative capture structure.
[0222] Step 5: High-level decision design and strategy optimization. First, input the group state and target state parameters of the UUV cluster, including the relative position and three-dimensional velocity of each UUV in the UUV cluster and the current position and velocity of the target UUV;
[0223] Secondly, through the semi-Markov decision process modeling, the round-up strategy is selected by optimizing the cumulative reward function and defining the parameter expected cumulative return , the parameter expected cumulative return expression is: ;
[0224] Then, a high-level policy selection is performed, at each time step Based on the current status , select an optimal execution behavior ;
[0225] Finally, output the high-level strategy , and pass it to the low-level execution network to generate detailed execution control commands, high-level strategies for: .
[0226] To further illustrate the high-level strategy The selection process now sets the state set , corresponding to the four strategic scenarios of lurking, intercepting, encircling, and tracking, the action set They represent four high-level behaviors: lurking, intercepting, encircling, and tracking. Construct the following reward function table: : ; ; ; ;
[0227] The specific strategy selection results can be calculated as follows: ⇒ "lurking"; ⇒ "intercept"; ⇒ "surround"; ⇒ "track";
[0228] Furthermore, give a more reasonable setting (not necessarily the best) and set the time step , discount factor , each step executes the strategy and receive instant rewards , its cumulative expected return is: .
[0229] The sixth step is to implement the low-level execution network design and control. The high-level strategies of lurking, intercepting, encircling and tracking are converted into specific control instructions, such as Figure 2 As shown in the figure, the green line represents UUV1, the red line represents UUV2, the yellow line represents UUV3, and the blue line represents the target UUV.
[0230] In the control method of the latent stage, the Lyapunov guidance vector field (LGVF) is introduced and the Lyapunov function is defined. and its time derivative ,function and its time derivative The expression is: ; ; is the convergence coefficient, , here it is set to 0.05; the navigation vector is: .
[0231] In order to avoid being detected by the target UUV, an obstacle avoidance item was added , obstacle avoidance item The expression is: ; Here it is set to 2.0; It is expressed as the UUV position within the detection range of the target UUV; Expressed as the detection radius of the target UUV, ;
[0232] according to Figure 9 Randomly select The two-dimensional position of each UUV is: ,at this time .
[0233] , , At this time, A<0, the obstacle avoidance item is adjusted negatively, and the obstacle avoidance mechanism has not yet been activated, meeting the strategic requirement of keeping a distance from the target UUV during the diving phase.
[0234] The control method in the interception phase is to adjust the target UUV speed , making Always keep constant speed The expression is: ; This process is a dynamic adjustment process, the purpose of which is to control the speed of the target UUV by controlling the distance between our UUV and the target UUV, thereby achieving the strategic goal of interception.
[0235] The control method of the encirclement phase is to adjust the heading angle deviation and speed modifiers , to achieve the ideal encirclement azimuth ;
[0236] Heading angle deviation The expression is: ); It is expressed as the heading angle control gain coefficient, which is set to 0.5 here; according to Figure 9 Take the current moment T in the encirclement phase UUVs =127°, At this time, the UUV coordinates are , the target position coordinates are .
[0237] Speed modifier The expression is: ); where It represents the proportional coefficient in speed control, which is set to 0.01 here; It is expressed as the expected target distance, which is 50m here, the same as the detection distance of the target UUV; ,but According to the heading angle deviation and speed correction term obtained at this moment, it can be known that when the heading angle of 127° is greater than the ideal encirclement azimuth angle of 120°, the heading angle is negatively controlled and the speed is decelerated to achieve the ideal encirclement azimuth angle of 120°, meeting the strategic goal of the encirclement phase.
[0238] The control method in the tracking phase defines the tracking error , by adjusting the speed and heading angle to minimize the tracking error , tracking error The expression is: ; During the tracking phase, a UUV The heading angle at this moment is , the target heading is , time step , the current distance to the target is , expected distance , so the tracking error indicator is =24.94, the UUV has achieved stable following and the heading control has basically converged.
[0239] The seventh step is to introduce and optimize behavioral cloning in high-level decision-making. Imitation Learning (IL) technology is introduced into the hierarchical reinforcement learning (HRL) framework. By initializing the high-level decision network using expert example datasets, it accelerates policy convergence, reduces exploration time, and improves policy robustness and adaptability.
[0240] The steps to introduce and optimize behavior cloning into high-level strategies include the following two aspects:
[0241] First, we design a behavioral cloning model. First, we define a high-level policy network. The function form of the high-level policy network is: ; Secondly, define the expert strategy function, the form of the expert strategy function is: ;Finally, define the loss function , and define the expert example dataset , including state-behavior pairs , by minimizing the loss function , so that the high-level policy network can imitate the expert decision-making behavior, the loss function The expression is: ; If the expert action is , and the policy network outputs the softmax probability as ,but
[0242] The second aspect is the joint optimization of behavior cloning and hierarchical reinforcement learning. First, based on the behavior cloning pre-training, combined with the hierarchical reinforcement learning objectives, a joint loss function is designed. , the goal is to maximize the cumulative reward and define the hierarchical reinforcement learning loss function , joint loss function The expression is: Where, satisfy , which represents the balance between controlling hierarchical reinforcement learning and behavior cloning weights; secondly, the optimization goal is to adjust Implement switching between hierarchical reinforcement learning and behavior cloning. , to ensure further strengthening of strategy optimization based on imitation learning.
[0243] The present invention introduces behavioral cloning technology into high-level decision networks to achieve an organic combination of expert knowledge and reinforcement learning, providing a stable and efficient solution for strategy optimization in multi-stage tasks.
[0244] Step 8: Reward function design and multi-objective optimization. Taking into account the encirclement effect, proximity and synchronization, the reward function is defined in the form of weighted summation. , the reward function The expression is:
[0245] ; Take the azimuth angles of the three UUVs at a certain moment as: , after converting to radians , the ideal equilateral angle is The distance between the i-th UUV and the target UUV at a certain moment is: .
[0246] Encirclement Effect Bonus It is used to ensure that the UUV cluster forms a stable encirclement around the target UUV at a uniform angle. The expression is: ;but Where, It represents the number of each of the three UUVs that perform the encirclement mission; Expressed as the time step The actual azimuth angle of each UUV relative to the target UUV;
[0247] Proximity Reward It is used to ensure that the distance between the UUV and the target UUV is close to the ideal value, avoiding being too close or too far. The smaller the distance deviation, the higher the reward, encouraging each UUV to maintain an appropriate encirclement distance. The expression is:
[0248] ;but .
[0249] Synchronicity Rewards Used to ensure the coordination and consistency between the UUVs in the UUV cluster, synchronize the action completion time, and the synchronized UUV cluster will receive positive rewards , otherwise there is no reward, the expression is: ; The current completion time of each UUV is , synchronization error tolerance , the positive reward is set to ,because ,but .
[0250] The weighting coefficient is set as: , the final total reward is: ; This value will serve as an immediate feedback signal for the current high-level strategy and participate in the joint loss optimization process.
[0251] In the ninth step, using the Unity3D simulation platform, a dynamic scenario was constructed in which three friendly UUVs collaborated to capture a single enemy UUV, simulating a multi-UUV coordinated capture mission in a complex underwater environment. This validated the feasibility, stability, and efficiency of the system's capture strategy, and optimized the algorithm strategy and system parameters based on the experimental results. To further explore the feasibility and performance of this strategy in real-world scenarios and verify its ability to accurately detect targets and effectively collaborate, the virtual simulation platform ensured high-precision simulation.
[0252] Combine Figure 3 、 Figure 4 as well as Figure 5 As shown in the figure, the blue line PPO represents a traditional reinforcement learning algorithm, while the green line HRL-IL represents the proposed algorithm combining hierarchical reinforcement learning (HRL) and behavior cloning (IL). The figure shows that the proposed algorithm offers the following advantages over the traditional reinforcement learning algorithm PPO: ① Significantly improved speed control stability: The green dashed line (HRL-IL) maintains a speed of nearly 2.0 m / s after 200 seconds, with minimal fluctuation. PPO, on the other hand, experiences frequent fluctuations and even drops to zero (e.g., between 400 and 600 seconds). In particular, PPO experiences severe deceleration and even freezing in UUV1 and UUV3, while HRL-IL maintains high maneuverability. ② Rapid and stable heading angle adjustment: The HRL-IL strategy converges to the ideal heading angle (e.g., 120°) more quickly. While the PPO strategy experiences repeated directional fluctuations at multiple stages (e.g., between 500 and 700 seconds for UUV2), HRL-IL completes heading adjustments earlier and maintains a stable angle for a longer period. ③ Improved coordination consistency among multiple UUVs: During the course changes of UUVs 1-3, the three UUVs in HRL-IL converged more closely; however, under PPO, there were significant lags in the adjustment timings of multiple UUVs, affecting the stability of the coordinated formation. ④ Higher mission efficiency and greater stability: The HRL-IL curve was smooth overall, with no sudden fluctuations. In particular, its overall performance outperformed PPO in the 700–900 s range; PPO exhibited frequent speed drops and angle jumps, which affected overall mission continuity. Conclusion: The combined hierarchical reinforcement and behavioral cloning mechanism (HRL-IL) introduced in this paper achieves smoother speed regulation, facilitating sustained mission progress. It also achieves faster convergence and more precise adjustments in directional control, reducing path detours and oscillations. It also facilitates synchronization within the multi-agent system, facilitating group encirclement and capture. It can complete encirclement missions with fewer strategy adjustments, resulting in greater mission continuity.
[0253] Combine Figure 5 、 Figure 6 as well as Figure 7As shown in the figure, the blue line PPO represents a traditional reinforcement learning algorithm, while the green line HRL-IL represents the hierarchical reinforcement learning (HRL) and behavior cloning (IL) algorithm of the present invention. As can be seen from the figure, the present invention has the following advantages over the traditional reinforcement learning algorithm PPO: ① Faster target approach, smaller encirclement distance, and greater stability: The green HRL-IL curve in all figures compresses the UUV-target distance to less than 100 meters around 400 seconds; while the PPO scheme only slowly converges after 700 seconds. The HRL-IL curve also exhibits near-zero oscillation in the 400-900 second interval, demonstrating high stability. ② Precise angle control, fast convergence to an equilateral angle of 120°: In all angle sub-figures, the red line represents the ideal angle of 120°; HRL-IL approaches the ideal encirclement angle in 300-400 seconds, while the PPO curve exhibits significant fluctuations, slow adjustment, and ultimately fails to fully converge. ③ Small jitter amplitude, smoother path control: The HRL-IL curve is smoother regardless of distance or angle changes; the PPO curve exhibits frequent oscillations, with amplitudes as high as 30–60 m and 40–60 degrees, indicating that traditional methods are less adaptable to hydrodynamic disturbances or error feedback. ④ Good multi-agent consistency, more coordinated collective collaboration: The HRL-IL curves of all UUVs in Figures 5–7 decrease almost synchronously; the PPO curves show asynchronous or delayed approach of individual UUVs, seriously affecting the final encirclement formation. Conclusion: The proposed HRL-IL strategy comprehensively outperforms traditional reinforcement methods in terms of accuracy, efficiency, and stability, making it particularly suitable for complex underwater capture missions.
[0254] like Figure 9 As shown in the figure, UUV1, UUV2, and UUV3 in the UUV swarm of the present invention exhibit the following advantages when encircling a target UUV: ① The three vessels work together to form a highly symmetrical equilateral triangle: As can be seen in the figure, UUV1 (green), UUV2 (red), and UUV3 (yellow) are evenly distributed around the target UUV (blue). The distances and angles are nearly symmetrical, indicating that the high-level strategy successfully guides each vessel toward the ideal position. The center of gravity of the circular trajectory nearly overlaps with the target trajectory, demonstrating high formation accuracy. ② The encirclement trajectory is smooth and efficient: After entering the encirclement trajectory, the three UUVs form an inward spiral. Compared to traditional tracking strategies (which may follow a straight or broken line), the present strategy forms a natural formation, avoiding conflicts. A blocking curve is formed along the target's path, increasing the probability of successful encirclement. ③ The multi-boat behavior is highly coordinated, and the capture time is highly synchronized: the final "landing point" of each boat is consistent, and the capture of the target is completed almost simultaneously (as can be seen from the end of the wake). This shows that the policy network takes into account the coordination and synchronization reward function between teams, and avoids the capture failure caused by some boats breaking away early or lagging behind.
[0255] The present invention accelerates strategy initialization by introducing behavioral cloning technology, and combines it with a hierarchical reinforcement learning framework to achieve efficient task decomposition and execution, effectively improving the performance of multi-UUV collaborative capture tasks. The present invention introduces expert example data sets through behavioral cloning technology, shortens the exploration time of reinforcement learning, greatly accelerates the convergence process of the strategy, provides a stable starting point for reinforcement learning, and reduces fluctuations in the early stages of training. The hierarchical decision-making framework proposed in the present invention can effectively decompose complex tasks, allowing UUVs to flexibly switch between different mission stages such as lurking, intercepting, encircling, and tracking, and can adapt to different environmental dynamics and target behaviors. It also comprehensively considers the reward function of encirclement effect, proximity, and synchronization, so that UUVs can achieve multi-target balance during mission execution and form a stable and efficient capture formation.
[0256] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.
Claims
1. A multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning, characterized by: The following steps are involved: S1. Establish a coordinate system with the UUV buoyancy center as the origin and initialize the UUV's physical property parameters; S2. Deploy a UUV cluster and obtain the linear velocity, acceleration, and Euler angle of each UUV about the X, Y, and Z axes; S3. Establish a three-dimensional ocean current model based on viscous Lamb's vortex to simulate the dynamic characteristics of the real underwater environment and obtain the position control equation of the UUV; S4. Arrange the deployed UUV cluster in an equilateral triangle formation; S5, modeling the group state and target state parameters of the UUV cluster through a semi-Markov decision process and outputting a high-level strategy; S6, convert high-level strategies into specific control commands; S7. Introducing behavior cloning technology into the hierarchical reinforcement learning framework and optimizing high-level strategies using expert example datasets; S8. Design a reward function based on the encirclement effect, proximity, and synchronization indicators of the UUV cluster to guide the update of the reinforcement learning strategy; S9. Construct and simulate the collaborative capture mission of multiple UUVs in a complex underwater environment in the simulation platform.
2. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 1 is characterized in that: In S1, the buoyancy center of the UUV is set as the origin of the fixed coordinate system, and the center of gravity position of the UUV is determined. and the center of buoyancy , and initialize the physical properties of the UUV: mass m, volume V, surface area A, and minimum turning radius R.
3. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 2 is characterized in that: In S2, there are three UUVs in the UUV cluster, and the actual linear velocity of each UUV about the X, Y, and Z axes is obtained. 、 、 With actual acceleration 、 、 , and get the rolling angle of each UUV , pitch angle , yaw angle .
4. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 3 is characterized in that: In S3, first, a three-dimensional ocean current model based on viscous lambda vortices is established. The expression for simulating the underwater dynamic characteristics of the real ocean is: ; Where, Expressed as the velocity component of the water flow along the X axis; Expressed as the velocity component of the water flow along the Y axis; Expressed as the velocity component of the water flow along the Z axis; Expressed as a yaw-angle-dependent strength parameter; Represented as the coordinate of the current position on the X-axis; Expressed as the coordinate of the initial point on the X-axis; Represented as the coordinate of the current position on the Y axis; Expressed as the coordinate of the initial point on the Y axis; Expressed as the distance from the current position to the vortex center; It is expressed as the starting radius of the vortex influence range; Expressed as the base of natural logarithms; It is expressed as a scale parameter that controls the rate at which the eddy spreads or decays; Secondly, define the main forces acting on UUV The expression is: ; Where, Expressed as fluid density; Expressed as the drag coefficient; It is represented as the UUV forward area, and the data is determined by hydrodynamic simulation; It is expressed as the velocity multiplied by its modulus; Expressed as a comprehensive resistance constant or parameter, representing the coefficient of the combined term on the right side of the formula; Finally, the position control equation of UUV in six degrees of freedom is obtained as follows: ; Where, It is expressed as the sum of external forces acting on the UUV along the X-axis; It is expressed as the sum of external forces acting on the UUV along the Y-axis; It is expressed as the sum of external forces acting on the UUV along the Z-axis; Expressed as the sum of the rolling moments around the X axis; Expressed as the sum of the rolling moments around the Y axis; Expressed as the sum of the rolling moments about the Z axis.
5. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 4 is characterized in that: In S4, first, the distance between any two adjacent UUVs among the three UUVs is defined as , the coordinates of one UUV are , the coordinates of the other UUV are , and 、 Satisfies the expression: ; Secondly, calculate the relative phase angle between any two adjacent UUVs , relative phase angle Used to ensure that each UUV is in the correct direction and relative position in the equilateral triangle formation, relative phase angle The expression is: 。 6. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 5 is characterized in that: In S5, first, the group state and target state parameters of the UUV cluster are input, including the relative position and three-dimensional velocity of each UUV in the UUV cluster and the current position and velocity of the target UUV; Secondly, through the semi-Markov decision process modeling, the round-up strategy is selected by optimizing the cumulative reward function and defining the parameter expected cumulative return , the parameter expected cumulative return expression is: ; Where, Expressed as a discount factor; Expressed as time steps; Indicates the maximum duration of policy execution; Expressed as time step The current state of the Expressed as time step Execution behavior under Indicates the current state Execution behavior Rewards received; Expressed as mathematical expectation, that is, the average cumulative reward under multiple strategy executions; Then, a high-level policy selection is performed, at each time step Based on the current status , select an optimal execution behavior ; Finally, output the high-level strategy : ; Where, Indicates the current state; It is expressed as an act of execution; Indicates the current state Execution behavior Rewards received; Indicates the variable value at which the expression reaches its maximum value.
7. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 6 is characterized in that: In S6, the high-level strategies of lurking, intercepting, encircling, and tracking are converted into specific control instructions; In the control method of the latent stage, the Lyapunov guidance vector field is introduced and the Lyapunov function is defined. and its time derivative ,function and its time derivative The expression is: ; ; Where, Expressed as The current position of each UUV; Represented as the target UUV position; is the convergence coefficient, ; In order to avoid being detected by the target UUV, an obstacle avoidance item was added , obstacle avoidance item The expression is: ; Where, Expressed as a coefficient greater than zero, used to adjust the obstacle avoidance item The weight or strength of ; It is expressed as the UUV position within the detection range of the target UUV; Expressed as the detection radius of the target UUV; The control method of the interception phase is to adjust the speed , making Always keep constant speed The expression is: ; Where, Expressed as a proportionality factor; It is represented by the actual distance between the current UUV and the target UUV. ; The control method of the encirclement phase is to adjust the heading angle deviation and speed modifiers , to achieve the ideal encirclement azimuth ; Heading angle deviation The expression is: ); Where, Represents the current moment; Expressed as The current heading angle of each UUV; Expressed as heading angle control gain coefficient; Speed modifier The expression is: ); Where, It is expressed as the proportional coefficient in speed control; Expressed as The current distance between each UUV and the target; Expressed as the expected target distance; The control method in the tracking phase defines the tracking error , by adjusting the speed and heading angle to minimize the tracking error , tracking error The expression is: ; Where, Represents the current moment No. The distance from the UUV to the target; Represents the current moment No. The heading angle of the UUV.
8. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 7 is characterized in that: In S7, behavior cloning technology is introduced into the hierarchical reinforcement learning framework. By initializing the high-level policy network with expert example datasets, it accelerates policy convergence, reduces exploration time, and improves the robustness and adaptability of the policy. The steps to introduce and optimize behavior cloning into high-level strategies include the following two aspects: First, we design a behavioral cloning model. First, we define a high-level policy network. The function form of the high-level policy network is: ; Where, Represents the behavioral decision output by the high-level policy network; Represented as parameters of the high-level policy network; Secondly, define the expert strategy function, which is in the form of: ; Where, Represents an expert in the current state Optimal decision-making behavior under Finally, define the loss function , and define the expert example dataset , including state-behavior pairs , by minimizing the loss function , so that the high-level policy network can imitate the expert decision-making behavior, the loss function The expression is: ; Where, Represented as a pair of expert example datasets The state-behavior pairs sampled from expected value; The second aspect is the joint optimization of behavior cloning and hierarchical reinforcement learning. First, based on the behavior cloning pre-training, combined with the hierarchical reinforcement learning objectives, a joint loss function is designed. , the goal is to maximize the cumulative reward and define the hierarchical reinforcement learning loss function , joint loss function The expression is: ; Where, satisfy , represents the balance between controlling hierarchical reinforcement learning and behavior cloning weights; Secondly, optimize the target by adjusting Implementing switching between hierarchical reinforcement learning and behavior cloning.
9. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 8 is characterized in that: In S8, the reward function is obtained The steps are: Considering the surrounding effect, proximity and synchronization, the reward function is defined by weighted summation , the reward function The expression is: ; Where, Expressed as the weight of the encirclement effect; Expressed as the weight of the proximity; expressed as the weight of synchronicity; It is expressed as a siege effect bonus; Expressed as proximity reward; Expressed as a synchronization reward; Encirclement Effect Bonus It is used to ensure that the UUV cluster forms a stable encirclement around the target UUV at a uniform angle. The expression is: ; Where, It represents the number of each of the three UUVs that perform the encirclement mission; Expressed as the time step The actual azimuth angle of each UUV relative to the target UUV; Proximity Reward It is used to ensure that the distance between the UUV and the target UUV is close to the ideal value, avoiding being too close or too far. The expression is: ; Where, It is expressed as the actual distance between the UUV and the target UUV; Expressed as the ideal enclosing radius; Synchronicity Rewards Used to ensure the coordination and consistency between the UUVs in the UUV cluster, synchronize the action completion time, and the synchronized UUV cluster receives positive rewards , otherwise no reward, synchronization reward The expression is: ; Where, 、 It represents the time it takes for the UUV to complete the action; Indicates the allowed time synchronization error threshold.
10. The multi-UUV collaborative capture strategy based on behavioral cloning and hierarchical reinforcement learning according to claim 9 is characterized in that: In S9, on the Unity3D simulation platform, a dynamic scene was constructed in which a UUV cluster consisting of three UUVs collaboratively surrounded a target UUV to simulate the collaborative capture mission of multiple UUVs in a complex underwater environment.
Citation Information
Patent Citations
Multi-agent self-organizing synergistic hunting method in non-convex environment
CN117574950A
Underwater multi-agent cooperative hunting method and device based on deep reinforcement learning
CN118153431A