Autonomous underwater robot three-dimensional dynamic trajectory planning method and system facing real marine environment
By improving the disturbance fluid dynamics and proximal strategy optimization algorithm, combined with real seabed topography and ocean current data, the trajectory planning problem of AUV in complex ocean environments was solved, and efficient and safe three-dimensional dynamic navigation was achieved.
Patent Information
- Application Number
- CN202510865019.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-26
AI Technical Summary
Existing trajectory planning algorithms are difficult to adapt to complex non-convex seabed terrain, multi-scale three-dimensional ocean current disturbances and dynamic obstacles, resulting in unsafe, inefficient and high energy consumption for AUV navigation in real ocean environments.
The improved disturbance fluid dynamics mechanism (IIFDS) is combined with the proximal policy optimization algorithm (PPO), and the real seabed topography and ocean current data are integrated to build a closed-loop collaborative system of state perception, action decision-making and reward evaluation for three-dimensional dynamic trajectory planning.
Significantly improve the navigation intelligence, safety and energy efficiency of AUVs in real environments, effectively handle irregular terrain and complex flow fields, and generate stable and feasible trajectories.
Smart Images

Figure CN120704371A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional trajectory planning for autonomous underwater robots, and more particularly to a planning method and system for a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment. Background Art
[0002] Autonomous underwater vehicles (AUVs) are becoming critical platforms for missions such as marine resource development, seafloor mapping, underwater search and rescue, and national defense reconnaissance. The AUV's operating environment is characterized by complex and variable seafloor topography, presenting a non-convex, irregular, and continuously undulating surface; three-dimensional, unstable ocean currents with strong velocity gradients coupled with depth; and the potential for a variety of dynamic interfering objects. Therefore, efficient, safe, and energy-efficient three-dimensional trajectory planning in a real-world ocean environment is a core challenge for AUV intelligent navigation.
[0003] Traditional trajectory planning algorithms, such as artificial potential field methods, A*, and RRT, are primarily applicable to two-dimensional or regular environments and struggle to handle continuous non-convex terrain and three-dimensional ocean current disturbances. The Improved Influenced Fluid Dynamic System (IIFDS) algorithm, inspired by the "water flowing around rocks" phenomenon, can generate smooth and feasible obstacle avoidance trajectories in continuous space. However, its original form primarily targets standard convex obstacles and fails to consider the synergistic effects of real ocean currents.
[0004] In recent years, reinforcement learning has demonstrated good adaptability in robot trajectory planning. However, most studies are based on idealized or simplified scenarios and lack real terrain and ocean current modeling, making it difficult to deploy training strategies in real ocean environments.
[0005] Therefore, there is an urgent need for a new trajectory planning method that integrates real seabed topography and ocean current information and combines trajectory guidance with strategy optimization mechanism to support the autonomous navigation of AUVs in complex three-dimensional ocean environments. Summary of the Invention
[0006] In view of this, the present invention provides a three-dimensional dynamic trajectory planning method and system for autonomous underwater robots in real ocean environments, which is used to solve the practical problems that existing trajectory planning algorithms cannot fully adapt to complex non-convex seabed terrain, multi-scale three-dimensional ocean current disturbances, and dynamic obstacle interactions. This method introduces an improved model based on the perturbation fluid dynamics mechanism (IIFDS) and combines it with the proximal policy optimization algorithm (PPO) for mission objective optimization to achieve efficient guidance and strategy learning of AUV trajectory behavior. The system fully integrates real terrain and ocean current data, and on this basis realizes closed-loop coordination of disturbance modeling, state perception, action decision-making and reward evaluation, significantly improving the navigation intelligence, safety and energy efficiency of AUVs in real environments.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A three-dimensional dynamic trajectory planning method for an autonomous underwater robot in a real ocean environment, comprising:
[0009] S1: Convert the current environmental data in the latitude, longitude and depth coordinate system to the local Cartesian coordinate system, and simultaneously obtain the seabed elevation matrix and the three-dimensional ocean current velocity field;
[0010] S2: Construct an initial guiding flow field based on the current position of the autonomous underwater vehicle and the target point position;
[0011] S3: collecting current state information of the environment acquired by the autonomous underwater robot to obtain first state information, inputting the first state information as input into a policy network based on a proximal policy optimization algorithm to obtain action parameters, and adjusting the initial guiding flow field according to the action parameters;
[0012] S4: Introducing a terrain disturbance term and an ocean current disturbance term into the adjusted initial guided flow field to obtain a final disturbance velocity vector, integrating the current position and the final disturbance velocity vector to obtain a position of the autonomous underwater vehicle at a next moment, and moving to the position; wherein the terrain disturbance term is calculated based on the seabed elevation matrix, and the ocean current disturbance term is calculated based on the three-dimensional ocean current velocity field;
[0013] S5: Repeat steps S3-S4 until the autonomous underwater robot reaches the target point and completes the three-dimensional dynamic trajectory planning task.
[0014] Preferably, the first state information includes: a vector pointing to the target point; a vector pointing to the surface of the nearest obstacle; a velocity vector of the nearest obstacle; an ocean current velocity vector; an ocean current velocity modulus; a terrain gradient; and a current height above the ground.
[0015] Preferably, the action parameters include: repulsive reaction coefficient ρ; tangential reaction coefficient σ; directional coefficient θ; ocean current fusion coefficient β; and propulsion factor thrust.
[0016] Preferably, the introduction of terrain disturbance terms and ocean current disturbance terms into the adjusted initial guiding flow field specifically includes: superimposing the terrain disturbance terms in the initial guiding flow field, constructing a non-convex terrain disturbance matrix by calculating the seabed terrain gradient at the current position, and synthesizing it with the original velocity vector direction to obtain a comprehensive disturbance velocity vector; after obtaining the comprehensive disturbance velocity vector, introducing the ocean current disturbance terms in proportion to obtain the final disturbance velocity vector.
[0017] Preferably, the terrain disturbance item specifically includes:
[0018] Assume the current position of the autonomous underwater robot is P = (x, y, z);
[0019] Calculate the terrain gradient and estimate the local slope of the terrain using the central difference method, taking the derivative in the x and y directions respectively:
[0020]
[0021] Where elev(i,j) represents the depth value at the i-th row and j-th column of the terrain DEM data, Δlon and Δlat represent the latitude and longitude grid intervals;
[0022] Unit normal vector construction. In three-dimensional space, the normal vector of the terrain patch can be obtained using the following expression:
[0023]
[0024] in, represents the unit normal disturbance direction;
[0025] Adjustment of slope amplitude and disturbance intensity, introducing slope amplitude as the disturbance intensity adjustment factor:
[0026]
[0027] Where s represents the intensity of the terrain slope; w adapt represents the disturbance response coefficient based on slope; ε represents the minimum disturbance threshold; w clip represents the final disturbance intensity;
[0028] The terrain disturbance matrix is constructed by using the unit normal vector to construct a projection matrix that suppresses the velocity. This matrix represents the "velocity offset" along the normal direction:
[0029]
[0030] in, represents the terrain disturbance matrix, represents the outer product of the unit normal vectors.
[0031] Preferably, obtaining the final disturbance velocity vector specifically includes:
[0032]
[0033] in, is a directional velocity vector with a modulus length; Refers to the three-dimensional ocean current velocity vector obtained by querying at the current position; β is the ocean current fusion coefficient; thrust refers to the propulsion factor; stepSize represents the control factor of the unit time step.
[0034] Preferably, the method further includes: collecting second state information, and calculating an instant reward based on the first state information, the action parameter, and the second state information;
[0035] The quadruple of the first state information, the action parameter, the immediate reward, and the second state information is stored in an experience replay buffer pool, and the policy network is periodically updated and trained.
[0036] Preferably, the instant reward includes: obstacle avoidance reward, terrain crossing reward, ground buffer zone reward, ocean current coordination reward, energy consumption reward and target distance reward, and all reward items are combined to obtain the final total reward function;
[0037] A reward function framework is constructed, and instant rewards are calculated through the reward function framework.
[0038] A three-dimensional dynamic trajectory planning system for autonomous underwater robots in real ocean environments, including:
[0039] The basic acquisition module converts the current environmental data in the latitude, longitude and depth coordinate system into the local Cartesian coordinate system, and simultaneously obtains the seabed elevation matrix and the three-dimensional ocean current velocity field;
[0040] The guidance flow field construction module constructs the initial guidance flow field according to the current position of the autonomous underwater robot and the position of the target point;
[0041] an adjustment module, collecting current state information of the environment acquired by the autonomous underwater robot to obtain first state information, inputting the first state information as input into a policy network based on a proximal policy optimization algorithm to obtain action parameters, and adjusting the initial guiding flow field according to the action parameters;
[0042] a position calculation module that introduces terrain disturbance terms and ocean current disturbance terms into the adjusted initial guided flow field, obtains a final disturbance velocity vector, integrates the current position and the final disturbance velocity vector, obtains the next position of the autonomous underwater vehicle, and moves to that position; wherein the terrain disturbance term is calculated based on the seabed elevation matrix, and the ocean current disturbance term is calculated based on the three-dimensional ocean current velocity field;
[0043] The mission planning module repeats the operations of the adjustment module and the position calculation module until the autonomous underwater robot reaches the target point and completes the three-dimensional dynamic trajectory planning task.
[0044] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method and system for planning the three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment. The present invention has the following beneficial effects:
[0045] This is the first time that a deep reinforcement learning trajectory planning system has been integrated with real-world ocean environment modeling. Based on high-precision seafloor topography (DEM) and three-dimensional vector ocean current data (.nc files), this system achieves unified modeling of irregular terrain structures, complex spatial flow fields, and dynamic obstacle interactions. By integrating this with deep trajectory planning strategies, it effectively enhances the AUV's ability to perceive and adapt to the real natural environment.
[0046] Constructing a trajectory generation mechanism that separates disturbance guidance from propulsion: Building on the original IIFDS algorithm, this paper innovatively decouples the disturbance term (for obstacle avoidance and guidance) from the ocean current dynamic term (for energy utilization) and introduces a current fusion coefficient β to dynamically control the fusion strength. This mechanism maintains the stability of flow-induced obstacle avoidance while preventing the ocean current term from interfering with the strategy learning process, enhancing the system's controllability and convergence efficiency.
[0047] Design of a 16-dimensional state space and 5-dimensional action space for real-world perception tasks: This invention is the first to construct a state and action expression structure specifically for real-world ocean trajectory planning tasks. The state space comprehensively integrates 16-dimensional information such as target pointing, obstacle perception, ocean current speed, terrain slope, and turbulence intensity; the action space is expanded to 5 dimensions, covering disturbance field control parameters, ocean current fusion coefficients, and propulsion factors, to achieve joint optimization control of energy consumption and trajectory feasibility.
[0048] Establish a multi-objective-oriented dynamic reward function structure: This invention designs a composite reward mechanism based on the multi-objective characteristics of actual marine missions (mission completion, obstacle avoidance safety, energy consumption control, and environmental coordination), introduces multiple reward components such as propulsion energy consumption, ocean current coordinated propulsion direction, terrain buffer distance, and endpoint hit, and sets a segmentation strategy to improve learning stability, effectively solving problems such as sparse rewards and local optimality.
[0049] Construct a high-fidelity, multi-obstacle disturbance simulation environment library: To support algorithm training and generalization evaluation, the present invention designs and implements multiple multi-dynamic obstacle scenario combination environments, which simultaneously contain real terrain, real ocean currents, and multiple interference obstacle models. It has stronger real-world adversarial and strategy generalization capabilities, significantly improving the adaptability and deployability of the algorithm in complex real-sea environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0051] Figure 1A schematic flow chart of the method provided by the present invention;
[0052] Figure 2 The real seabed topographic map provided by the present invention;
[0053] Figure 3 Overlay map of real ocean current data provided by the present invention;
[0054] Figure 4 A diagram of the reward function during the training process provided by the present invention;
[0055] FIG5( a ) is a three-dimensional trajectory diagram of the AUV provided by the present invention in a terrain-ocean current coupled environment;
[0056] FIG5( b ) is a top view of the geographical coordinates provided by the present invention;
[0057] FIG6( a ) is a diagram showing the three-dimensional trajectory generation result provided by an embodiment of the present invention;
[0058] FIG6( b ) is a diagram showing the trajectory evolution of an AUV in a two-dimensional terrain map under multi-obstacle conditions provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0060] Problems based on existing technologies:
[0061] The applicability of the algorithm is limited to regular convex obstacles: Reaction field algorithms represented by IIFDS are originally designed to construct repulsive force fields based on regular geometric shapes (such as spheres and ellipsoids). Their obstacle avoidance mechanism is relatively effective in standard convex obstacle scenarios; however, when faced with non-convex seabed terrain with irregular boundaries, sharp protrusions or deep concave structures (such as trenches, slopes, and ridges), this type of algorithm cannot accurately generate appropriate reaction vectors, and is prone to falling into local minima or producing trajectory oscillations, and lacks the ability to adapt to complex terrain disturbances.
[0062] Lack of modeling and control mechanisms for three-dimensional ocean current disturbances: Traditional trajectory planning methods or deep reinforcement learning frameworks often fail to incorporate real ocean current data as a basis for dynamic environment modeling. Even some methods that consider disturbance factors often use simplified two-dimensional fields or static vector fields, which fail to reflect the dynamic characteristics of ocean currents over time and space. The lack of a coordinated propulsion mechanism makes it difficult for AUVs to navigate or actively avoid currents, resulting in low propulsion efficiency and abnormal energy consumption.
[0063] Incomplete expression of state space and action space: In existing deep reinforcement learning trajectory planning research, state space design often focuses on the distance between the target point and the obstacle, failing to fully express the interactive relationship with ocean environmental variables such as ocean currents, terrain, depth, and turbulence. This results in insufficient strategy generalization ability and difficulty in migrating to unstructured or complex disturbance environments.
[0064] The reward function structure is simple and the strategy training efficiency is low: Most deep reinforcement learning trajectory planning algorithms use sparse or single-goal-oriented reward functions, which make it difficult to cover the multi-dimensional goals of task execution, such as propulsion energy consumption control and ocean current coordinated propulsion direction, thereby affecting strategy stability and convergence speed.
[0065] The simulation environment is idealized and lacks real-world validation: The training and verification environments of existing trajectory planning algorithms are often simplified two-dimensional models, or scenes with a single structure and static obstacles. These environments cannot reflect the interference factors of dynamic obstacles (such as underwater work platforms and other underwater vehicles) in real ocean missions, limiting the adaptability and practicality of the algorithms in real environments.
[0066] The embodiment of the present invention discloses a method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment. Figure 1 Shown, including:
[0067] S1: Convert the current environmental data in the latitude, longitude and depth coordinate system to the local Cartesian coordinate system, extract the seabed elevation matrix from the .tif format seabed topography map, and extract the three-dimensional ocean current velocity field from the .nc format ocean current data;
[0068] S2: Based on the current position of the autonomous underwater vehicle and the position of the target point, the improved disturbed fluid dynamic system (IIFDS) is used to construct the initial guiding flow field;
[0069] S3: collecting current state information of the environment acquired by the autonomous underwater vehicle to obtain first state information, inputting the first state information as input into a policy network based on a proximal policy optimization algorithm (PPO), obtaining action parameters, and adjusting the initial guiding flow field according to the action parameters;
[0070] S4: Introducing a terrain disturbance term and an ocean current disturbance term into the adjusted initial guided flow field to obtain a final disturbance velocity vector, integrating the current position and the final disturbance velocity vector to obtain a position of the autonomous underwater vehicle at a next moment, and moving to the position; wherein the terrain disturbance term is calculated based on the seabed elevation matrix, and the ocean current disturbance term is calculated based on the three-dimensional ocean current velocity field;
[0071] S5: Repeat steps S3-S4 until the autonomous underwater robot reaches the target point and completes the three-dimensional dynamic trajectory planning task.
[0072] The IIFDS algorithm is based on simulating the characteristics of natural fluids. In the absence of disturbances, the initial flow streamlines point directly to the target point. The IIFDS algorithm is used to calculate the initial flow velocity, including:
[0073] In the initial flow field, the fluid vector velocity U(P) flowing through each position point P is expressed as:
[0074]
[0075] Among them, U(P) represents the initial flow field velocity vector of the current position of the autonomous underwater robot P = (x, y, z), V0 represents the virtual velocity constant, which is used to set the flow intensity. d =(x d ,y d ,z d ) represents the target point position, d(P,P d ) represents the current position P of the autonomous underwater robot and the target point P d The Euclidean distance of . That is:
[0076]
[0077] In a specific embodiment, the AUV obtains environmental information through sensors and constructs the current state. The first state information includes: a vector pointing to the target point Vector pointing to the surface of the nearest obstacle The velocity vector of the nearest obstacle Ocean current velocity vector Ocean current velocity mode length Terrain gradient Current height above ground z h .
[0078] In a specific embodiment, the action parameters include: a repulsive reaction coefficient ρ; a tangential reaction coefficient σ; a directional coefficient θ; a current fusion coefficient β; and a propulsion factor thrust.
[0079] In a specific embodiment, the introduction of terrain disturbance terms and ocean current disturbance terms into the adjusted initial guiding flow field specifically includes: superimposing the terrain disturbance terms in the initial guiding flow field, constructing a non-convex terrain disturbance matrix by calculating the seabed terrain gradient at the current position, and synthesizing it with the original velocity vector direction to obtain a comprehensive disturbance velocity vector; after obtaining the comprehensive disturbance velocity vector, introducing the ocean current disturbance terms in proportion to obtain the final disturbance velocity vector.
[0080] In a specific embodiment, the trajectory points at the next moment obtained by IIFDS planning include:
[0081] (1) Constructing the initial guiding flow field
[0082] In the initial pilot flow field constructed using the improved disturbed fluid dynamic system (IIFDS), the fluid vector velocity U(P) flowing through each position point P is expressed as:
[0083]
[0084] Among them, U(P) represents the initial flow field velocity vector of the current position of the autonomous underwater robot P = (x, y, z), V0 represents the virtual velocity constant, which is used to set the flow intensity. d =(x d ,y d ,z d ) represents the target point position, d(P,P d ) represents the current position P of the autonomous underwater robot and the target point P d The Euclidean distance of .
[0085]
[0086] (2) Standard convex obstacle disturbance modeling
[0087] When standard convex obstacles exist in the environment, they will cause disturbances to the previously defined initial flow field, causing the flow direction of the fluid to change. The effect of standard convex obstacles on the initial flow field is modeled by the perturbation matrix. The perturbation matrix is defined as:
[0088]
[0089] Among them, M obs (P) represents the total perturbation matrix of the standard convex obstacle, ω k (P) represents the weight coefficient of the kth obstacle, M k (P) represents the perturbation matrix of the kth obstacle, and K represents the total number of obstacles.
[0090] The weight coefficient of each obstacle is defined as follows:
[0091]
[0092] Among them, φ k (P) represents the obstacle surface equation, which defines the internal, surface and external positions of the obstacle.
[0093] Perturbation matrix M k(P) is defined by the combination of the exclusion term and the tangential term:
[0094]
[0095] Where I represents the third-order unit matrix, also called the attraction matrix, which has a function similar to the gravitational function in the artificial potential field method. k represents the normal vector of the obstacle surface, t k represents the tangent vector, ρ k Represents the repulsion coefficient, which controls the strength of the obstacle repulsion. The larger its value, the earlier the disturbed fluid can avoid obstacles in the environment. k It represents the tangential reaction coefficient, which controls the intensity of the tangential flow of the fluid. T represents the distance from the center of the obstacle to the current position, normalized to the obstacle radius.
[0096] In a three-dimensional environment with obstacles, the appearance of obstacles causes the original flow field trajectory to change, forming a disturbed flow field. This disturbed flow field can converge at the end point while bypassing the obstacles.
[0097] (3) Tangential vector modeling
[0098] In practical applications of the IFDS algorithm, streamlines generated for obstacle avoidance are typically confined to a single plane, which can cause the generated trajectory to become trapped or stagnant at specific points. To address this issue, the IIFDS algorithm introduces the concept of a tangent matrix, which frees streamlines from being confined to a single plane and allows them to flow in any direction around standard convex obstacles. When approaching a stagnation point, the IIFDS algorithm adds a tangential matrix to provide tangential momentum along the obstacle surface, effectively preventing the AUV from becoming stuck at the stagnation point.
[0099] In the normal vector n k On the tangent plane defined by k Generated by:
[0100] t k =R k t' k ;
[0101] Among them, t' k represents the tangent vector in the local tangent coordinate system, based on the tangent angle θ, R k Represents the rotation matrix from the local tangent basis vector to the global coordinate system, specifically:
[0102]
[0103] Here, θ controls the rotation angle in the tangential direction and determines the specific direction to bypass the standard convex obstacle.
[0104] (4) Impact of dynamic obstacles
[0105] The velocity of a dynamic obstacle affects the flow field through the following formula:
[0106]
[0107] Among them, v obs Indicates the speed impact of the dynamic obstacle on the current position, T represents the normalized distance between the current point and the obstacle, λ represents the attenuation factor, which is used to control the impact range of the obstacle speed on the flow field, V obs Represents the dynamic obstacle velocity vector.
[0108] (5) Influence of non-convex real seabed topography
[0109] The steps for the non-convex true seabed topography disturbance term are as follows:
[0110] Assume the current position of the autonomous underwater robot is P = (x, y, z);
[0111] Calculate the terrain gradient and estimate the local slope of the terrain using the central difference method, taking the derivative in the x and y directions respectively:
[0112]
[0113] Where elev(i,j) represents the depth value at the i-th row and j-th column of the terrain DEM data, Δlon and Δlat represent the latitude and longitude grid intervals, which are used to calculate the terrain slope;
[0114] Unit normal vector construction. In three-dimensional space, the normal vector of the terrain patch can be obtained using the following expression:
[0115]
[0116] The numerator vector is the local slope direction of the terrain gradient structure, with a negative sign indicating "pointing from the ground upwards"; the denominator is the Euclidean norm (modulus) of the vector, which is used for normalization.
[0117] represents the unit normal disturbance direction, i.e., the local vertical direction of the terrain surface;
[0118] Adjustment of slope amplitude and disturbance intensity, introducing slope amplitude as the disturbance intensity adjustment factor:
[0119]
[0120] Where s represents the intensity of the terrain slope (dimensionless), the larger the value, the steeper the slope; w adaptrepresents the disturbance response coefficient based on slope, which is used to smoothly control the disturbance intensity; ε represents the minimum disturbance threshold to prevent the disturbance from being too small to cause failure; w clip represents the final disturbance intensity;
[0121] The terrain disturbance matrix is constructed by using the unit normal vector to construct a projection matrix that suppresses the velocity. This matrix represents the "velocity offset" along the normal direction:
[0122]
[0123] in, represents the terrain disturbance matrix, a real symmetric matrix, It represents the outer product of the unit normal vector, that is, the velocity projection component in this direction. The negative sign before the matrix and the disturbance weight together determine that it is an "inhibition effect", that is, the velocity decreases in the normal propulsion direction.
[0124] (6) Calculation of comprehensive velocity after disturbance by standard convex obstacles and non-convex real seabed terrain
[0125] The calculation formula of the comprehensive speed of AUV is as follows:
[0126]
[0127] (7) Ocean current disturbances
[0128] The final disturbance velocity vector calculation formula after adding the ocean current disturbance term is as follows:
[0129]
[0130] in, It refers to the directional velocity vector with a modulus, where the modulus represents the "suggested propulsion velocity magnitude" and the direction is the "direction after disturbance". The "disturbance-guided velocity vector" calculated by the improved IIFDS is composed of the terrain disturbance matrix in target attraction, obstacle detour, and terrain disturbance. Refers to the three-dimensional ocean current velocity vector obtained by querying the current position; β refers to the "ocean current fusion coefficient" learned by the improved PPO, which controls whether to follow the current (β→1) or ignore the flow field (β→0); thrust refers to the propulsion factor, which is learned by the improved PPO strategy and can adjust the speed to save energy or avoid obstacles; stepSize represents the control factor of the unit time step.
[0131] In a specific embodiment, the method further includes: collecting second state information, and calculating an instant reward based on the first state information, the action parameter, and the second state information;
[0132] The quadruple of the first state information, the action parameter, the immediate reward, and the second state information is stored in an experience replay buffer pool, and the policy network is periodically updated and trained.
[0133] In a specific embodiment, the instant reward includes: obstacle avoidance reward, terrain crossing reward, ground buffer zone reward, ocean current coordination reward, energy consumption reward, and target distance reward. All reward items are combined to obtain the final total reward function;
[0134] A reward function framework is constructed, and instant rewards are calculated through the reward function framework.
[0135] (8) Trajectory update
[0136] Based on the comprehensive speed of calculation Will Integrate and obtain the trajectory point at the next moment according to the following formula:
[0137]
[0138] Specifically, the calculation of instant rewards includes:
[0139] Construct a reward function framework, which includes obstacle avoidance reward, terrain traversal reward, ground clearance buffer reward, ocean current coordination reward, energy consumption reward, and target distance reward. Combine all reward items to obtain the final total reward function.
[0140] Through the reward function framework, the immediate reward is calculated.
[0141] Specifically, the present invention proposes an improved reward function framework that aims to provide more accurate guidance for dynamic trajectory planning of AUVs. This reward function includes six main components: obstacle avoidance reward, terrain traversal reward, ground clearance buffer reward, ocean current coordination reward, energy consumption reward, and target distance reward, as follows:
[0142] (1) Obstacle Avoidance Reward
[0143] Obstacle avoidance rewards encourage the AUV to maintain a safe distance from obstacles. This is achieved by modeling a buffer zone around obstacles using a layered linear function. By applying a piecewise increasing penalty when approaching an obstacle and a fixed strong penalty upon actual collision, the policy can quickly learn obstacle avoidance behaviors during the early stages of reinforcement learning.
[0144]
[0145] Among them, d obs Indicates the distance from the current position of the AUV to the center of the dynamic obstacle, R obs represents the dynamic obstacle radius, δ bufIndicates the buffer radius.
[0146] (2) Terrain Crossing Rewards
[0147] Terrain traversal bonuses are used to enforce the prevention of AUVs from traversing the seafloor.
[0148]
[0149] Among them, z<z seafloor It indicates that the AUV currently crosses the terrain surface, that is, below the seabed DEM.
[0150] (3) Off-ground buffer zone reward
[0151] The off-ground buffer zone reward and penalty mechanism encourages the AUV to stabilize in the appropriate water layer, that is, not sticking to the seabed.
[0152]
[0153] Where m = zz seafloor This design penalizes ground-based navigation near the seabed to avoid bottom collision risk, while also providing positive rewards for moderate ground-based altitudes to encourage trajectory generation within a reasonable terrain layer.
[0154] (4) Ocean Current Synergy Reward
[0155] The current coordination bonus measures the consistency between the AUV's current direction and the current. It rewards forward motion (downstream) and imposes a severe penalty on reverse propulsion (against the current) when the current fusion coefficient β is large. This design enhances the strategy's adaptability and coordination to currents, making it particularly suitable for energy-efficient navigation strategies.
[0156]
[0157] in, It represents the cosine of the angle between the current direction and the ocean current direction. β∈[0,1] represents the ocean current fusion coefficient, which is output by PPO.
[0158] (5) Energy consumption reward
[0159] Punish high energy consumption behaviors and encourage AUVs to use energy-saving movement methods.
[0160] r thrust =-λ·thrust 2 ;
[0161] Where thrust∈[0,1] represents the thrust factor, which is output by PPO. λ represents the thrust cost weight.
[0162] (6) Successful Ending Point Rewards
[0163]
[0164] This item is used to give a one-time large reward when the AUV successfully reaches the target point (xy two-dimensional plane distance is less than 100m, depth error is less than 10m), thereby strengthening the destination guidance.
[0165] (7) Total Reward Function
[0166] All reward items are combined to obtain the final total reward function, which comprehensively guides the AUV to complete the task and ensure smooth, safe and efficient movement.
[0167] R=r avoid +r below +r margin +r ocean +r thrust +r hit ;
[0168] A three-dimensional dynamic trajectory planning system for autonomous underwater robots in real ocean environments, including:
[0169] The basic acquisition module converts the current environmental data in the latitude, longitude and depth coordinate system into the local Cartesian coordinate system, and simultaneously obtains the seabed elevation matrix and the three-dimensional ocean current velocity field;
[0170] The guidance flow field construction module constructs the initial guidance flow field according to the current position of the autonomous underwater robot and the position of the target point;
[0171] an adjustment module, collecting current state information of the environment acquired by the autonomous underwater robot to obtain first state information, inputting the first state information as input into a policy network based on a proximal policy optimization algorithm to obtain action parameters, and adjusting the initial guiding flow field according to the action parameters;
[0172] a position calculation module that introduces terrain disturbance terms and ocean current disturbance terms into the adjusted initial guided flow field, obtains a final disturbance velocity vector, integrates the current position and the final disturbance velocity vector, obtains the next position of the autonomous underwater vehicle, and moves to that position; wherein the terrain disturbance term is calculated based on the seabed elevation matrix, and the ocean current disturbance term is calculated based on the three-dimensional ocean current velocity field;
[0173] The mission planning module repeats the operations of the adjustment module and the position calculation module until the autonomous underwater robot reaches the target point and completes the three-dimensional dynamic trajectory planning task.
[0174] Specifically, the key to the three-dimensional dynamic trajectory planning method of the autonomous underwater robot proposed in the embodiment of the present invention lies in the training process, and the most important part of the training is the construction of a standardized simulation environment. The selected real seabed terrain area ranges from 127° to 128° east longitude and 29° to 30° north latitude. The real seabed terrain area map is as follows Figure 2 Then the ocean current data in this range is interpolated. The processed real topographic ocean current map is as follows: Figure 3 As shown in the figure. Furthermore, considering the uncertainty of dynamic obstacle motion in real-world mission scenarios, dynamic obstacles with varying movement rates, impact radius, and trajectory are introduced into the simulation environment. During training, at the beginning of each episode, a random starting and ending point is selected near the initial starting and ending points, and a random dynamic obstacle is selected from the set of dynamic obstacles.
[0175] The results of the training are as follows Figure 4 The figure shows the cumulative reward curve during trajectory planning training for a three-dimensional dynamic trajectory planning method for an autonomous underwater robot in a real ocean environment. The horizontal axis represents the number of training episodes, and the vertical axis represents the total reward value obtained by the autonomous underwater robot during each trajectory planning task.
[0176] The following key features can be observed from the training results:
[0177] The initial stage is highly exploratory: Within the first 50 rounds, the strategy is still in the exploratory learning stage, with unstable trajectory quality and significant reward fluctuations. The lowest reward value once dropped to approximately -600, reflecting that the unoptimized strategy has difficulty in completing effective obstacle avoidance and target advancement tasks in complex environments.
[0178] The strategy gradually converges and stabilizes: With the increase of training rounds, especially after the 100th round, the cumulative reward shows a continuous upward trend, indicating that the policy network gradually learns to effectively adjust trajectory behavior and can comprehensively balance multiple factors such as target propulsion, obstacle avoidance, ocean current coordination and terrain adaptation.
[0179] The final performance was stable and robust: From rounds 150 to 370, the strategy's overall reward remained within a range of 300–500, demonstrating that the strategy had developed a control mode with stable propulsion, good obstacle avoidance, and environmental adaptability. In particular, after round 300, the reward peaked at nearly 580, demonstrating that the strategy was capable of high-quality three-dimensional trajectory planning under complex seabed terrain and turbulent ocean currents.
[0180] Next, the trained model was tested in a single dynamic obstacle environment, as shown in Figure 5. The starting point was set to [127.89, 29.1, -800], and the ending point was set to [127.2, 29.60, -10]. Figure 5(a) shows a three-dimensional global trajectory diagram of trajectory planning using the trained model in a single dynamic obstacle environment. Although ocean currents are not shown in this figure, the planned trajectory takes them into account. Figure 5(a) shows the three-dimensional trajectory of the AUV in a terrain-current coupled environment (blue line). The terrain is constructed using real DEM data, showing significant slope fluctuations and non-convex shapes. Red spheres represent dynamic obstacles and their impact ranges, and red lines represent the dynamic obstacle trajectory. The starting point and ending point are marked with blue and orange solid dots, respectively. During the AUV's motion, by modeling the perturbation of the IIFDS flow field and dynamically adjusting the parameters of the PPO strategy, a continuous and feasible trajectory was successfully generated that strikes a balance between target direction and obstacle avoidance. Figure 5(b) shows the corresponding top-down view (top view in geographic coordinates). The background shows the depth distribution of the actual seafloor topography, with darker colors representing deeper seafloor depths. The superimposed black arrows represent the ocean current velocity field in the corresponding area. In the figure, the blue line represents the AUV's three-dimensional trajectory in the terrain-current coupled environment. The red sphere represents the impact range of the dynamic obstacle, and the red line represents the trajectory traversed by the dynamic obstacle. The blue sphere represents the starting point, the yellow star represents the end point, and the blue triangle represents the AUV's current position. It can be observed that the AUV's trajectory significantly changes direction when it reaches the dynamic obstacle, demonstrating its effective obstacle avoidance strategy. The trajectory is smooth overall, without jagged edges or abnormal behavior crossing terrain or obstacle core areas. The trajectory fully considers terrain slope, seafloor boundaries, and ocean current direction, avoiding blind advance against the current and promoting energy-efficient navigation. After passing through the obstacle interference area, the trajectory quickly returns to the target direction, demonstrating the trajectory recovery capability of the disturbance guidance mechanism.
[0181] To more comprehensively verify the robustness and generalization of the three-dimensional dynamic trajectory planning method for an autonomous underwater vehicle in a real ocean environment proposed in an embodiment of the present invention, this embodiment simulates a dynamic and complex underwater environment in reality and tests the model in an environment containing multiple dynamic obstacles. As shown in Figure 6, the starting point is set to [127.89, 29.1, -800], and the ending point is set to [127.2, 29.60, -10]. Four dynamic obstacles are randomly selected from the dynamic obstacle set to form a dynamic obstacle combination environment. The results of trajectory planning using the trained model are shown in Figure 6. Figure 6(a) shows a three-dimensional global trajectory view of trajectory planning using the trained model in the dynamic obstacle combination environment. Although ocean currents are not shown in this figure, the planned trajectory takes them into account. The blue line represents the three-dimensional trajectory of the AUV in the terrain-current coupled environment. The terrain is constructed using real DEM data. The red spheres represent the influence range of the dynamic obstacles. The red line represents the trajectory traversed by the dynamic obstacles. The blue and orange solid spheres represent the starting and ending points, respectively. Figure 6(b) shows the corresponding top-down view (top view in geographic coordinates). The background shows the depth distribution of the actual seafloor topography, with darker colors representing deeper seafloor. The superimposed black arrows represent the ocean current velocity field in the corresponding area. The blue line represents the AUV's 3D trajectory in the terrain-current coupled environment. The red sphere represents the impact range of the dynamic obstacle, and the red line represents the trajectory of the dynamic obstacle. The blue sphere represents the starting point, the yellow star represents the end point, and the blue triangle represents the AUV's current position. Figure 6(a) shows the 3D trajectory generated for the AUV in the real terrain and current interference environment. The trajectory originates from the deep sea area at the bottom of the figure, traverses multiple dynamic obstacle motion zones, and ultimately reaches the target point above. As can be seen from the figure, the AUV trajectory continuously circumvents obstacles without collision, demonstrating that the designed perturbation superposition mechanism and PPO strategy can effectively achieve obstacle avoidance and trajectory correction in complex dynamic scenarios. The overall trajectory is continuous and smooth, and does not cross terrain boundaries, indicating that the terrain perturbation modeling effectively avoids collision risks. The final trajectory maintains a reasonable propulsion direction and speed, balancing obstacle avoidance safety and navigation efficiency. Figure 6(b) shows the trajectory evolution of the AUV on a two-dimensional terrain map under multiple obstacle conditions using latitude and longitude coordinates, with information on the depth of the seabed terrain and the direction of the ocean current superimposed. It can be seen that the AUV advances from the deep sea along a curved trajectory toward the target, avoiding multiple red dynamic obstacle trajectories along the way. Although the trajectory deflects to a certain extent, it always remains close to the target direction and does not experience violent oscillations or regression, demonstrating the PPO optimization strategy's good stability and directional control capabilities. Furthermore, the trajectory distribution as a whole avoids high and steep slopes and strong ocean current reversal zones, demonstrating the adaptability and robustness of the trajectory generation mechanism under the dual constraints of terrain and ocean currents.
[0182] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0183] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A three-dimensional dynamic trajectory planning method for an autonomous underwater robot in a real ocean environment, characterized by: include: S1: Convert the current environmental data in the latitude, longitude and depth coordinate system to the local Cartesian coordinate system, and simultaneously obtain the seabed elevation matrix and the three-dimensional ocean current velocity field; S2: Construct an initial guiding flow field based on the current position of the autonomous underwater vehicle and the position of the target point; S3: collecting current state information of the environment acquired by the autonomous underwater robot to obtain first state information, inputting the first state information as input into the improved policy network based on the proximal policy optimization algorithm to obtain action parameters, and adjusting the initial guiding flow field according to the action parameters; S4: Introducing a terrain disturbance term and an ocean current disturbance term into the adjusted initial guided flow field to obtain a final disturbance velocity vector, integrating the current position and the final disturbance velocity vector to obtain a position of the autonomous underwater vehicle at a next moment, and moving to the position; wherein the terrain disturbance term is calculated based on the seabed elevation matrix, and the ocean current disturbance term is calculated based on the three-dimensional ocean current velocity field; S5: Repeat steps S3-S4 until the autonomous underwater robot reaches the target point and completes the three-dimensional dynamic trajectory planning task.
2. The method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to claim 1, characterized in that: The first state information includes: a vector pointing to the target point; a vector pointing to the surface of the nearest obstacle; a velocity vector of the nearest obstacle; an ocean current velocity vector; an ocean current velocity modulus; a terrain gradient; and a current height above the ground.
3. The method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to claim 1, characterized in that: The action parameters include: repulsive reaction coefficient ρ; tangential reaction coefficient σ; directional coefficient θ; ocean current fusion coefficient β; and propulsion factor thrust.
4. The method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to claim 1, characterized in that: The method of introducing terrain disturbance terms and ocean current disturbance terms into the adjusted initial guiding flow field specifically includes: superimposing the terrain disturbance terms into the initial guiding flow field, constructing a non-convex terrain disturbance matrix by calculating the seabed terrain gradient at the current position, and synthesizing it with the original velocity vector direction to obtain a comprehensive disturbance velocity vector; after obtaining the comprehensive disturbance velocity vector, introducing the ocean current disturbance terms in proportion to obtain the final disturbance velocity vector.
5. The method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to claim 4, characterized in that: The terrain disturbance items specifically include: Assume the current position of the autonomous underwater robot is P = (x, y, z); Calculate the terrain gradient and estimate the local slope of the terrain using the central difference method, taking the derivative in the x and y directions respectively: Where elev(i,j) represents the depth value at the i-th row and j-th column of the terrain DEM data, Δlon and Δlat represent the latitude and longitude grid intervals; Unit normal vector construction. In three-dimensional space, the normal vector of the terrain patch can be obtained using the following expression: in, represents the unit normal disturbance direction; Adjustment of slope amplitude and disturbance intensity, introducing slope amplitude as the disturbance intensity adjustment factor: In clip =max(in adapt ,ε); Where s represents the intensity of the terrain slope; w adapt represents the disturbance response coefficient based on slope; ε represents the minimum disturbance threshold; w clip represents the final disturbance intensity; The terrain disturbance matrix is constructed by using the unit normal vector to construct a projection matrix that suppresses the velocity. This matrix represents the "velocity offset" along the normal direction: in, represents the terrain disturbance matrix, represents the outer product of the unit normal vectors.
6. The method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to claim 5, characterized in that: The obtaining of the final disturbance velocity vector specifically includes: in, is a directional velocity vector with a modulus length; Refers to the three-dimensional ocean current velocity vector obtained by querying at the current position; β is the ocean current fusion coefficient; thrust refers to the propulsion factor; stepSize represents the control factor of the unit time step.
7. The method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to claim 1, characterized in that: Also includes: collecting second state information, and calculating an instant reward based on the first state information, the action parameter, and the second state information; The quadruple of the first state information, the action parameter, the immediate reward, and the second state information is stored in an experience replay buffer pool, and the policy network is periodically updated and trained.
8. The method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to claim 7, characterized in that: The instant reward includes: obstacle avoidance reward, terrain crossing reward, ground buffer zone reward, ocean current coordination reward, energy consumption reward and target distance reward. All reward items are combined to obtain the final total reward function; A reward function framework is constructed, and instant rewards are calculated through the reward function framework.
9. A system for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment, applying the method for planning a three-dimensional dynamic trajectory of an autonomous underwater robot in a real ocean environment according to any one of claims 1 to 8, characterized in that: include: The basic acquisition module converts the current environmental data in the latitude, longitude and depth coordinate system into the local Cartesian coordinate system, and simultaneously obtains the seabed elevation matrix and the three-dimensional ocean current velocity field; The guidance flow field construction module constructs the initial guidance flow field according to the current position of the autonomous underwater robot and the position of the target point; an adjustment module, collecting current state information of the environment acquired by the autonomous underwater robot to obtain first state information, inputting the first state information as input into a policy network based on a proximal policy optimization algorithm to obtain action parameters, and adjusting the initial guiding flow field according to the action parameters; a position calculation module that introduces terrain disturbance terms and ocean current disturbance terms into the adjusted initial guided flow field, obtains a final disturbance velocity vector, integrates the current position and the final disturbance velocity vector, obtains the next position of the autonomous underwater vehicle, and moves to that position; wherein the terrain disturbance term is calculated based on the seabed elevation matrix, and the ocean current disturbance term is calculated based on the three-dimensional ocean current velocity field; The mission planning module repeats the operations of the adjustment module and the position calculation module until the autonomous underwater robot reaches the target point and completes the three-dimensional dynamic trajectory planning task.