Multi-AUV safe path planning method based on deep reinforcement learning

By introducing safety cost constraints and safety corrections into deep reinforcement learning, and combining the MATD3 algorithm with a centralized training architecture, the problem of collaborative safety of multi-AUV systems in complex underwater environments is solved. This achieves a balance between autonomous decision-making safety and exploratory behavior in multi-AUV systems, and improves the safety and reliability of path planning.

CN121433291APending Publication Date: 2026-01-30HARBIN ENG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511469198.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing path planning algorithms cannot guarantee the collaborative safety of multiple AUV systems in complex underwater operating environments, and the unconstrained exploration in deep reinforcement learning may lead to AUVs entering an unsafe state.

Method used

By introducing safety cost constraints and safety corrections into deep reinforcement learning, and combining them with the MATD3 algorithm, a safety layer is established to limit exploration behavior within a safe range. A centralized training and decentralized decision-making architecture is adopted to optimize the updates of the policy network and the value network.

Benefits of technology

It improves the safety and reliability of path planning in multi-AUV systems, ensures a balance between safety and exploratory decision-making in complex environments, and enhances the collaborative safety of multi-AUV systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121433291A_ABST
    Figure CN121433291A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-AUV safe path planning method based on deep reinforcement learning. According to the method, a deep reinforcement learning algorithm model based on an MATD3 method is adopted, a policy network and a value network are updated by adopting security constraints, the security of policy learning is improved, the expected reward revenue is maximized under the condition that the expected security cost constraints are met, a frequent minimum and maximum optimization process is avoided by adopting first-order penalty optimization, and the security of policy learning is improved. Meanwhile, safety correction based on a safety layer is added in the training process to guarantee safety in the early stage of training, and exploratory and safety balance is brought to strategy optimization of reinforcement learning through safety constraint and safety correction. According to the multi-AUV path planning method, the time cooperation constraint and the space cooperation constraint of multi-AUV path planning are comprehensively considered, the centralized training and decentralized decision-making architecture is applied to multi-AUV path planning, the path planning method capable of ensuring cooperation safety is provided for a multi-AUV system, and the safety and the reliability of path planning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multi-autonomous underwater robot path planning technology, specifically involving a multi-AUV safe path planning method based on deep reinforcement learning. Background Technology

[0002] Autonomous underwater vehicles (AUVs) are intelligent robots capable of independently performing underwater tasks, playing an increasingly important role in marine scientific research and resource exploration. With the rapid development of AUV technology and the continuous expansion of its application scenarios, the flexibility, efficiency, and robustness of multi-AUV systems have led to increasing attention being paid to multi-AUV task collaboration technology.

[0003] Path planning is a core technology for achieving autonomy and flexibility in multi-AUV systems. When multi-AUV systems operate in complex marine environments with numerous obstacles and unclear topography, such as nearshore, near-port, and unknown waters, limitations in sensor performance, AUV dynamic constraints, environmental disturbances, and uncertainties pose significant challenges to the safety, stability, and economy of path planning. Conventional path planning algorithms mostly rely on accurate environmental models, which suffer from response lag when handling path planning tasks in complex marine environments. Deep reinforcement learning algorithms, however, can adapt to unknown environments without any signal guidance, learning online and through trial and error. This allows multi-AUV systems to gradually adapt to complex and changing environments and achieve autonomous decision-making, demonstrating strong generalization capabilities.

[0004] Complex underwater operating environments place high demands on the autonomy and safety of multi-AUV systems. Reinforcement learning learns available strategies by exploring an unknown state space. Through trial and error in interaction with the environment, it can gradually adapt to the environment and make autonomous decisions by setting reasonable reward functions. However, since the policy optimization of reinforcement learning depends on exploration, unconstrained exploration will inevitably lead AUVs to unsafe states, making it impossible to ensure the safety of multi-AUV path planning. Summary of the Invention

[0005] The purpose of this invention is to address the problem that existing path planning algorithms cannot guarantee the collaborative safety of multiple AUV systems in complex underwater operating environments. This invention provides a multi-AUV safe path planning method based on deep reinforcement learning. By treating safety requirements as hard constraints, the method optimizes strategies within the feasible domain, providing a balance for the autonomous decision-making of multi-AUV systems. It ensures the full exploratory nature of the learning process through continuous interaction with the environment, while introducing safety cost constraints to force the exploration behavior to be limited to a predetermined safety range.

[0006] This invention addresses the problem that unconstrained exploration inevitably leads AUVs to unsafe states in conventional deep reinforcement learning-based multi-AUV path planning methods. It adds safety cost constraints based on first-order penalty optimization to the learning and updating of the network, and adds safety correction based on the safety layer to the output action of the policy network, which can effectively ensure the safety of autonomous decision-making in multi-AUV systems.

[0007] The technical solution adopted in this invention is:

[0008] A multi-AUV safe path planning method based on deep reinforcement learning includes the following steps:

[0009] S1. Establish a multi-AUV ranging sonar detection model;

[0010] S2. Establish a path planning simulation environment;

[0011] S3. Establish a reinforcement learning training framework by setting the state space, action space, reward function, and safety cost for multi-AUV path planning;

[0012] S4. Based on the multi-AUV ranging sonar detection model, path planning simulation environment, and reinforcement learning training framework, train the MATD3 algorithm model with added safety cost constraints and safety corrections until the preset training rounds are reached to obtain a trained multi-AUV safe path planning model.

[0013] Furthermore, the specific process of establishing the multi-AUV ranging sonar detection model in S1 is as follows: during the AUV's navigation, beams are continuously emitted in various directions, and the distance between the AUV and the obstacle is calculated based on the beam propagation speed in seawater and the time difference between transmitting and receiving the beams.

[0014] Furthermore, in S1, the detection range of the multi-AUV ranging sonar detection model is in front of the AUV. The maximum detection range is a 90° opening fan-shaped region within the plane. The detection range is divided into 6 regions, each capable of detecting a 15° fan-shaped area. Obstacles within;

[0015] The detected obstacle information is processed by feature normalization, as shown in the following formula:

[0016]

[0017] in, It indicates the first Sonar detection information in a fan-shaped area, This represents the maximum detection range of the ranging sonar. When the ranging sonar does not detect any obstacles within this fan-shaped area, the detection information for that area... It should be the maximum detection range, at which point the detection matrix... When the ranging sonar detects an obstacle within this fan-shaped area, the detection information... The detection matrix is ​​less than the maximum detection range. The closer the AUV is to the obstacle, The closer the value is to 1.

[0018] Furthermore, in S2, a path planning simulation environment is established, including setting the number, position, and speed information of multiple AUVs, the position of the task target point, and the size, position, and number of obstacles; the task target point can be set as one or more.

[0019] Furthermore, the state space in S3 is defined as follows: the state space includes the coordinate positions of all AUVs, the target point position, the position of obstacles, the relative positions between AUVs, and the motion state information of each AUV. For the state space of a multi-AUV system, Indicates the first The status information of each AUV, among which... This indicates the corresponding target location. This indicates the coordinate position of the AUV. Indicates motion state information. This is the sonar detection matrix information for the AUV. Indicates the first Distance information between each AUV and other AUVs.

[0020] Furthermore, the definition of the action space in S3 includes the set of actions that each AUV can possibly take. For multiple AUVs' operating space, For the first Action information of each AUV, among which , , These are the AUV's speed, yaw angle, and pitch angle information, respectively.

[0021] Furthermore, the reward function in S3 is defined as including target distance reward, temporal and spatial reward, and task completion reward.

[0022]

[0023]

[0024]

[0025] in, For the first Target distance bonus for each AUV For the first Temporal and spatial rewards for each AUV Rewards for completing the mission; Indicates the first The coordinates of each AUV at the current moment. Indicates that it takes action The subsequent position coordinates, This indicates the target location corresponding to the AUV. , and The reward coefficients of each reward function, This represents the maximum time step required for the task to complete.

[0026] Furthermore, the definition of safety cost in S3 includes the comprehensive safety cost of obstacle avoidance safety between AUVs and obstacles, and collision avoidance safety between AUVs.

[0027]

[0028]

[0029] in, For the first The distance between the AUV and the obstacle For the first The AUV and the first The distance of an AUV, The set safe distance is n, the number of time steps is C, and the total safety cost of one round is C.

[0030] Furthermore, in S4, a multi-agent reinforcement learning training architecture with centralized training and decentralized decision-making is adopted. Centralized training is used to learn better policies. Each agent has its own policy network. During training, the observations, actions, and rewards of all agents are collected through the central controller to help the agents train their policy networks. After training, each agent makes decisions based on its own observations and its own policy network, and no longer needs to communicate with the central controller.

[0031] Furthermore, in S4, a policy network and a value network are established based on the MATD3 algorithm, which are used to output the action decision of the AUV and evaluate the quality of the action output by the policy network, respectively. The policy network and the value network are updated with safety cost constraints to maximize the expected reward while meeting the expected safety cost constraints. The policy network parameters are optimized with a first-order penalty function to avoid frequent min-max optimization processes. At the same time, a safety correction based on a safety layer is added during the training process to ensure the safety in the early stage of training. Through safety cost constraints and safety correction, a balance between exploration and safety is brought to the policy optimization of reinforcement learning.

[0032] Two value networks and parameters and Updated using the loss function and security cost constraints:

[0033]

[0034]

[0035]

[0036] in, and All are temporal difference objectives, D is empirical replay, L is mean squared loss, and the policy network is... parameters Updated by policy gradient and security cost constraints:

[0037]

[0038] in, For the output of the policy network, the target policy network and the target value network use soft updates to update their parameters:

[0039]

[0040]

[0041] To improve security for soft weight updates, a security-layer-based correction is added to the output action of the policy network.

[0042] ,

[0043] in, As a safety threshold, Based on the single-step safety cost constraint of the previous moment, For safety evaluation of the action.

[0044] Compared with the prior art, the present invention has the following advantages:

[0045] This invention employs a deep reinforcement learning algorithm model based on the MATD3 method. It uses safety cost constraints to update the policy network and value network, thereby improving the safety of policy learning. Under the condition of satisfying the expected safety cost constraints, it maximizes the expected reward. First-order penalty optimization is used to avoid frequent min-max optimization processes. At the same time, safety correction based on the safety layer is added during the training process to ensure safety in the early stage of training. Through safety cost constraints and safety correction, a balance between exploration and safety is achieved in the policy optimization of reinforcement learning.

[0046] This invention comprehensively considers the temporal and spatial coordination constraints of multi-AUV path planning, and applies centralized training and decentralized decision-making architecture to multi-AUV path planning. It provides a path planning method that ensures cooperative safety for multi-AUV systems, improves the safety and reliability of path planning, and provides new possibilities for the application of deep reinforcement learning in multi-AUV cooperative technology. Attached Figure Description

[0047] Figure 1 This is a training diagram of the multi-AUV path planning method based on deep reinforcement learning in this invention;

[0048] Figure 2 This is a schematic diagram of the AUV ranging sonar detection model in this invention;

[0049] Figure 3 This is a schematic diagram illustrating the comprehensive safety cost of obstacle avoidance safety and collision avoidance safety in this invention;

[0050] Figure 4 This is a schematic diagram of a centralized training and decentralized decision-making architecture based on deep reinforcement learning in this invention;

[0051] Figure 5 This is a schematic diagram of the structure of the multi-AUV safe path planning method based on deep reinforcement learning in this invention. Detailed Implementation

[0052] To better understand the purpose, structure, and function of this invention, the invention will be described in further detail below with reference to the accompanying drawings.

[0053] The basic idea of ​​this invention is to combine safety cost constraints and safety corrections with policy exploration using deep reinforcement learning, and to make safety decisions using the trained policy network. Addressing the complexity of the underwater environment, this invention employs a model-free deep reinforcement learning method, allowing AUVs to autonomously explore and learn policies, and adopting a safety-learning training approach to make safe action decisions. To address the collaborative safety problem of multi-AUV path planning arising from unconstrained exploration, safety requirements are treated as hard constraints, and policies are optimized within the feasible region, providing a balance between policy exploration and safety cost constraints for the autonomous decision-making of multi-AUV systems.

[0054] Step 1: Establish a multi-AUV ranging sonar detection model. During the AUV's navigation, continuously emit beams in various directions within a 90° opening angle in front of the AUV. Calculate the distance between the AUV and obstacles based on the beam propagation speed in seawater and the time difference between transmitting and receiving the beams.

[0055] like Figure 2 As shown, the detection range of the multi-AUV ranging sonar detection model is in front of the AUV. The maximum detection range is a 90° opening fan-shaped region within the plane. The detection range is divided into 6 zones, each capable of detecting a 15° fan-shaped area. Obstacles within a certain range. The AUV only needs to use obstacle information from six areas perceived at the current location to make decisions.

[0056] To optimize the retention of obstacle information in the detection matrix and improve the usability of the detection data, the following normalization method is designed:

[0057]

[0058] in, It indicates the first Sonar detection information in a fan-shaped area, This represents the maximum detection range of the ranging sonar. When the ranging sonar does not detect any obstacles within this fan-shaped area, the detection information for that area... It should be the maximum detection range, at which point the detection matrix... When the ranging sonar detects an obstacle within this fan-shaped area, the detection information... The detection matrix is ​​less than the maximum detection range. The closer the AUV is to the obstacle, The closer the value is to 1.

[0059] Step 2: Establish a path planning simulation environment, including setting the number, position and speed information of multiple AUVs, the position of the task target point, and the size, position and number of obstacles; the task target point can be set as one or more, and collisions with obstacles and other AUVs need to be avoided during the path planning process.

[0060] Step 3: Establish a reinforcement learning training framework. By setting the state space, action space, reward function, and safety cost of multi-AUV path planning, a basic framework is provided for training the safety learning algorithm.

[0061] It includes the settings of state space, action space, reward function, and safety cost; where the state space contains the coordinate positions of all AUVs, the target point position, the position of obstacles, the relative positions between AUVs, and the motion state information of each AUV; the action space includes the set of actions that each AUV may take; the reward function includes target distance reward, time and space reward, and mission completion reward; the safety cost is the comprehensive safety cost including obstacle avoidance safety and collision avoidance safety.

[0062] The state space contains the coordinates of all AUVs, the target point's position, the positions of obstacles, the relative positions between AUVs, and the motion state information of each AUV. It is defined as follows: For the state space of a multi-AUV system, Indicates the first The status information of each AUV, among which... This indicates the corresponding target location. This indicates the coordinate position of the AUV. Indicates motion state information. This is the sonar detection matrix information for the AUV. Indicates the first Distance information between each AUV and other AUVs.

[0063] Action space includes the set of actions that each AUV can take, defined as follows: For multiple AUVs' operating space, For the first Action information of each AUV, among which , , These are the AUV's airspeed, yaw angle, and pitch angle information. Due to the inherent dynamic constraints of the AUV, its turning ability is relatively weak. The range of values ​​for the AUV's pitch angle and roll angle is set to... The AUV's sailing speed variation is set in Within a certain range, the AUV's action decision range is set to... The step size.

[0064] The reward function includes target distance reward, time and space reward, and task completion reward:

[0065]

[0066]

[0067]

[0068]

[0069] in, For the first Target distance bonus for each AUV For the first Temporal and spatial rewards for each AUV Rewards for completing the mission; Indicates the first The coordinates of each AUV at the current moment. Indicates that it takes action The subsequent position coordinates, This indicates the target location corresponding to the AUV. , and The reward coefficients of each reward function, This represents the maximum time step required for the task to complete.

[0070] like Figure 3 As shown, the safety cost is the comprehensive safety cost that includes obstacle avoidance safety between AUVs and obstacles, and collision avoidance safety between AUVs:

[0071]

[0072]

[0073] in, For the first The distance between the AUV and the obstacle For the first The AUV and the first The distance of an AUV, The set safe distance is n, the number of time steps is C, and the total safety cost of one round is C.

[0074] Step 4: Based on the multi-AUV ranging sonar detection model, path planning simulation environment, and reinforcement learning training framework, train the MATD3 algorithm model with safety cost constraints and safety correction strategies until the pre-set training rounds are reached to obtain the trained multi-AUV safe path planning model.

[0075] A policy network and a value network are established based on the MATD3 algorithm. The policy network is used to output the AUV's action decisions, and the value network is used to evaluate the quality of the actions output by the policy network. In each training iteration, each AUV agent generates an action based on its own state using the policy network and obtains rewards and new states through interaction with the environment. The value network collects the states and actions of all AUVs and generates corresponding values. The value is used to evaluate the quality of the action output by the policy network, guiding the policy network to make better decisions. Two value networks are used to select the output. Smaller network values ​​can be used to mitigate the overestimation problem.

[0076] like Figure 4 As shown, a multi-agent reinforcement learning training architecture with centralized training and decentralized decision-making is employed. Each AUV is treated as an agent, and centralized training is used for comprehensive policy evaluation, enabling real-time decision-making. Each agent has its own policy network. During training, the central controller collects observations, actions, and rewards from all agents to help train their policy networks. After training, each agent makes decisions based on its own observations and policy networks, eliminating the need for communication with the central controller.

[0077] like Figure 5 As shown, a method combining safety cost constraints and safety corrections is adopted. Safety cost constraints are used to update the policy network and value network to improve the safety of policy learning. Under the premise of satisfying the expected safety cost constraints, the expected reward is maximized. First-order penalty optimization is used to avoid frequent min-max optimization processes. At the same time, safety corrections based on safety layers are added during the training process to ensure safety in the early stage of training. Through safety cost constraints and safety corrections, a balance between exploration and safety is achieved in the policy optimization of reinforcement learning.

[0078] The security layer, based on the existing policy network, uses state features to explicitly represent the sensitivity of actions to changes in security signals, utilizing a parametric linear model:

[0079]

[0080] in, Based on the single-step safety cost constraint of the previous moment, For safety evaluation of actions, based on single-step dynamics, some basic prior knowledge is incorporated. It does not attempt to learn the complete transformation model, but only the constraint functions. Each security signal All use about The linear model is approximated by the coefficients of the model. The features are extracted using a neural network, and the single-step cost function is approximated through supervised training, projecting unsafe actions back into the feasible region:

[0081]

[0082] in, As a safety cost threshold, Let t be the action taken at time t. The security layer consists of a deep policy network that optimizes the action correction process through each forward propagation. The linearized security signal model allows for the closed-form solution... This simplifies to a linear projection. The safety problem of AUV path planning has only one combined safety cost signal, therefore the closed-form solution is:

[0083]

[0084] in, This is the safety threshold. The safety layer differs from standard reward mechanisms in that it is a method for handling state and transient constraints.

[0085] A Constrained Markov Decision Process (CMDP) is a problem that aims to maximize expected reward while satisfying expected safety cost constraints. This leads to the constrained optimization problem.

[0086]

[0087] in, Let t be the single-step reward. The safety cost at time t, Let be the discount factor and d be the safety cost threshold. The Lagrange relaxation method simplifies the constrained optimization problem into an unconstrained optimization problem by adding the objective function to the corresponding Lagrange multiplier-weighted constraint function. Safety reinforcement learning can be viewed as a type of constrained sequential optimization problem:

[0088]

[0089] in, For strategy The evaluation value of state s is used in the policy. and Lagrange multiplier The objective is optimized using alternating gradient descent methods, and the optimization problem is solved by addressing the dual problem:

[0090]

[0091] The timescale for updating the primal variables needs to be faster than that for the Lagrange multipliers. Primitive-dual optimization is applied to alternately update the primal and dual variables:

[0092]

[0093] By approximating the state Lagrange multipliers with an additional neural network, a new state Lagrange function is constructed to solve for the optimal feasible policy and the safest possible policy for infeasible states. Given a policy... Define the cost if in state s Expected returns If the inequality is satisfied, then the state is either safe or feasible:

[0094]

[0095] in Representation Strategy From a given initial state Start generating the expected value of the trajectory. This is called the safe state cost function. A state that is unsafe regardless of the chosen strategy is defined as an infeasible state, while the feasible region is defined as the complementary set of the infeasible regions.

[0096]

[0097] Therefore, the optimization problem with safety cost constraints is:

[0098]

[0099] because Each state in the equation has a constraint, therefore each state has a corresponding Lagrange multiplier, which can be expressed as: The generated Lagrange function is named the original state Lagrange function:

[0100]

[0101] Adjusting the original Lagrange multiplication table to fit the sample-based learning paradigm, and rescaling the state constraints, yields:

[0102]

[0103] Its unique feature lies in the existence of an infinite number of state-dependent Lagrange multipliers, and the update of the multiplier network is shown below:

[0104]

[0105] in, These are the parameters of the multiplier network. The penalty function method can be used to perform first-order penalized optimization on the state Lagrangian method. The basic idea of ​​the penalty function is to transform the constrained problem into an unconstrained problem using a penalty function, and then solve it using unconstrained optimization methods.

[0106]

[0107]

[0108] in: Assuming It is the optimal value of the penalty function constraint problem. Let be the Lagrange multiplier vector corresponding to its dual problem, if If the constraint problem before and after optimization shares the same optimal solution set.

[0109] The penalty function method can construct an equivalent function whose unconstrained minimum point also solves the constrained problem of the Lagrange method. Furthermore, the unconstrained problem can be handled by a single, consistent penalty factor to address multiple constraints. Therefore, using a single penalty factor for first-order penalized optimization of the Lagrange method allows for a single minimization of the original variable with a fixed penalty term, avoiding frequent minimax optimization processes.

[0110]

[0111] The alternative policy network is updated as follows:

[0112]

[0113] For each agent, samples are randomly selected from the experience replay pool, and clipped noise is removed. Introduce target actions to aid value learning:

[0114]

[0115] in It limits the target action value to a valid range. Functions within, two value networks and parameters and Updated using the loss function and security cost constraints:

[0116]

[0117]

[0118] Based on security cost constraints, the value network is updated as follows:

[0119]

[0120]

[0121]

[0122] Policy Network The parameters are Updated by policy gradient and safety cost constraints based on first-order penalty optimization:

[0123]

[0124]

[0125] Where D represents experience replay. The target policy network and target value network use soft updates to update their parameters:

[0126]

[0127]

[0128] in, The weights are updated softly. After training, the trained policy network is used to make safety decisions for multi-AUV cooperative path planning.

[0129] In addition to the above embodiments, the present invention may have other implementation methods. All technical solutions formed by equivalent substitution or equivalent transformation fall within the protection scope claimed by the present invention.

Claims

1. A method for multi-AUV safe path planning based on deep reinforcement learning, characterized in that: The method comprises the following steps: S1. establishing a multi-AUV ranging sonar detection model; S2. establishing a path planning simulation environment; S3. establishing a reinforcement learning training framework by setting a state space, an action space, a reward function and a safety cost of the multi-AUV path planning; S4. training the MATD3 algorithm model with the safety cost constraint and the safety correction based on the multi-AUV ranging sonar detection model, the path planning simulation environment and the reinforcement learning training framework until a preset training round is reached, and obtaining a trained multi-AUV safety path planning model.

2. The method of claim 1, wherein: The specific process of establishing the multi-AUV ranging sonar detection model in S1 is that the AUV continuously emits beams in all directions during navigation, and the distance between the AUV and the obstacle is calculated based on the propagation speed of the beam in seawater and the time difference between the emission and reception of the beam.

3. The method of claim 2, wherein: The detection range of the multi-AUV ranging sonar detection model is the front of the AUV The maximum detection distance of the fan-shaped interval with an opening angle of 90° in the plane is The detection range is divided into 6 regions, and each region can detect obstacles within a 15° fan-shaped region ​ The detected obstacle information is normalized by the following formula: wherein, represents the first sonar detection information of the fan-shaped region, is the maximum detection distance of the ranging sonar, when the ranging sonar does not detect an obstacle in the fan-shaped region, the detection information of the region should be the maximum detection distance, at this time the detection matrix ; when the ranging sonar detects an obstacle in the fan-shaped region, the detection information is less than the maximum detection distance, at this time the detection matrix , the closer the AUV is to the obstacle, the value of the detection matrix 4. The method of claim 1, wherein: The path planning simulation environment in S2 includes setting the number, position and speed information of the multiple AUVs, the position of the task target point, and the size, position and number of the obstacles.

5. The method of claim 1, wherein: The definition of the state space in S3: the state space contains the coordinate position of all AUVs, the position of the target point, the position of the obstacles, the relative position between AUVs and the motion state information of each AUV, and the definition is the state space of the multi-AUV system, represents the state information of the first AUV, wherein, represents the corresponding target position thereof, represents the coordinate position of the AUV, represents the motion state information, is the sonar detection matrix information of the AUV, represents the distance information between the first AUV and other AUVs.

6. The method of claim 1, wherein: Definition of the action space in S3: including the set of actions that each AUV can take, defined as for a multi-AUV action space, for the first AUV, where , , are the AUV's velocity, yaw, and pitch information, respectively.

7. The method of claim 1, wherein: The definition of the reward function in S3 includes target distance reward, time-space reward and task completion reward, wherein, is the target distance reward for the th AUV, is the time-space reward for the th AUV, is the task end reward; represents the coordinate value of the th AUV at the current time, represents the position coordinate after the AUV takes the action , represents the target position corresponding to the AUV, , and are the reward coefficients of the respective reward functions, is the maximum time step of the task end.

8. The method of claim 1, wherein: The definition of the safety cost in S3 includes the comprehensive safety cost of obstacle avoidance safety between the AUV and the obstacle and collision avoidance safety between the AUVs, wherein, is the distance between the th AUV and the obstacle, is the distance between the th AUV and the th AUV, is the set safety distance, n is the number of time steps, and C is the total safety cost for a round.

9. The method of claim 1, wherein: In S4, a multi-agent reinforcement learning training architecture with centralized training and decentralized decision making is adopted, each AUV is regarded as an agent, centralized training is used for comprehensive evaluation of the strategy, each agent has its own strategy network, during training, the central controller collects the observations, actions and rewards of all agents to help the agent train the strategy network; after training, each agent makes decisions according to its own observations using its own strategy network, and no longer needs to communicate with the central controller.

10. The method of claim 9, wherein: In S4, the strategy network and the value network are established based on the MATD3 algorithm, which are used to output the action decision of the AUV and evaluate the quality of the action output by the strategy network respectively, the strategy network and the value network are updated with the safety cost constraint, the expected reward is maximized under the condition of meeting the expected safety cost constraint, the first-order penalty function is used to optimize the strategy network parameters to avoid frequent minimax optimization process, and the safety correction based on the safety layer is added during the training process to ensure the safety in the early training period, the safety cost constraint and the safety correction bring balance between exploration and safety to the strategy optimization of reinforcement learning; Two value networks and parameters of and are updated by loss function and safety cost constraints: where, and are temporal difference targets, D is experience replay, L is mean squared loss, the parameters of the policy network are updated by policy gradients and safety cost constraints: where, As the output of the policy network, the target policy network and the target value network update their parameters using soft updates: wherein, is a soft update weight, and a safety correction based on the safety layer is added on the output action of the policy network to improve safety: , wherein, is a safety threshold, is a single-step safety cost constraint at the previous time instant, is a safety evaluation of the action.

Citation Information

Cited By

  • Transition control methods and systems to assist robots in switching control strategies

    CN122308109A