AUV (Autonomous Underwater Vehicle) three-dimensional path planning method for dynamic obstacle environment

By combining the improved DDQN algorithm with the Kalman filter and the Singer model, the discrete path nodes of the AUV are obtained and processed continuously, which solves the problem of low path planning efficiency of under-actuated AUV in dynamic environments and realizes efficient and safe three-dimensional path planning.

CN120628102APending Publication Date: 2025-09-12HUNAN UNIV OF SCI & TECH SANYA RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510750745.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When an under-actuated AUV is performing a mission far away from its mother ship, it is difficult to generate a safe and efficient trajectory in real time in a three-dimensional dynamic environment due to its limited navigation and control computing power.

Method used

An improved dual deep Q-network (DDQN) algorithm is used in combination with the extended Kalman filter and the Singer model. The discrete path nodes of the AUV are obtained through state prediction and path planning models, and the path is continuous and smoothed through the basis spline function.

Benefits of technology

The three-dimensional path planning of AUV in dynamic obstacle environment is realized, which improves the efficiency and safety of path planning and ensures that AUV can perform tasks efficiently and safely in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120628102A_ABST
    Figure CN120628102A_ABST
Patent Text Reader

Abstract

The invention relates to an AUV (Autonomous Underwater Vehicle) three-dimensional path planning method for a dynamic obstacle environment, which comprises the following steps of: setting a three-dimensional state space, updating a motion state of a dynamic obstacle in the three-dimensional state space, and acquiring a future motion state of the dynamic obstacle based on the motion state; inputting the target position parameter and the future motion state of the three-dimensional state space into a path planning model, and obtaining discrete path nodes of the unmanned underwater vehicle; the path planning model is obtained by combining a training set with a comprehensive reward function; and performing path serialization and smoothing processing on the discrete path nodes to obtain a three-dimensional path planning result. The method shows obvious advantages in the aspects of path length, obstacle avoidance success rate and calculation efficiency, and effectively solves the problem of three-dimensional path planning in a dynamic obstacle environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of path planning for autonomous underwater vehicles (AUVs), and in particular to a three-dimensional path planning method for AUVs in dynamic obstacle environments. Background Art

[0002] An AUV is an unmanned underwater vehicle (AUV) that uses its own energy to navigate autonomously and can perform a variety of tasks depending on the payload it carries. As an essential tool for exploring and developing marine resources, it boasts a wide range of operations and high maneuverability, making it widely used in seabed environmental mapping, marine resource exploration, hydrographic data collection, and marine rescue operations. Currently, most AUVs used for deep-sea exploration are underactuated, typically consisting of only tail thrusters, pitch and yaw rudders to enable the AUV to navigate and steer. When performing cruising missions away from a mother ship or shore-based platform, the AUV's actual effectiveness is often limited by the measurement and computational capabilities of its navigation and control systems.

[0003] Autonomous navigation and intelligent control have become core technologies for the future development of AUV applications. Path planning is a crucial component of these two key technologies. It aims to design an optimal trajectory from a starting point to a destination within a specified mission area, while simultaneously meeting mission requirements and environmental constraints. As a key indicator of an AUV's ability to interact with the external environment, the quality of path planning is directly related to its mission efficiency and navigation safety. Summary of the Invention

[0004] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to propose a three-dimensional path planning method for AUVs in dynamic obstacle environments, which shows obvious advantages in path length, obstacle avoidance success rate and computational efficiency, and effectively solves the three-dimensional path planning problem in dynamic obstacle environments.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A three-dimensional path planning method for an AUV in a dynamic obstacle environment, comprising:

[0007] Setting a three-dimensional state space, updating a motion state of a dynamic obstacle in the three-dimensional state space, and obtaining a future motion state of the dynamic obstacle based on the motion state;

[0008] Inputting the target position parameters of the three-dimensional state space and the future motion state into a path planning model to obtain discrete path nodes of the unmanned underwater vehicle; the path planning model is obtained by training using a training set combined with a comprehensive reward function;

[0009] Path continuity and smoothing processing are performed on the discrete path nodes to obtain a three-dimensional path planning result.

[0010] Optionally, the three-dimensional state space includes: the AUV current position coordinates, the target position coordinates, the dynamic obstacle position coordinates, the fixed obstacle position coordinates and the ocean current vector.

[0011] Optionally, updating the motion state of the dynamic obstacle in the three-dimensional state space includes:

[0012] Determine the maneuver acceleration time-dependent function of the dynamic obstacle:

[0013]

[0014] in, is the variance of acceleration, τ is the time constant;

[0015] The acceleration of the dynamic obstacle is obtained by combining the Gaussian white noise with the target mean and target variance in the Singh model with the maneuver acceleration time correlation function:

[0016]

[0017] Where a(t) is the acceleration of the obstacle; α is the acceleration attenuation coefficient, and w(t) is Gaussian white noise.

[0018] Optionally, obtaining the future motion state of the dynamic obstacle includes:

[0019] The motion state is input into the state prediction model to obtain the future motion state of the dynamic obstacle:

[0020]

[0021] in, is the state at the next moment, is the current state estimate, and F1 is the state transfer matrix.

[0022] Optionally, in the process of inputting the motion state into the state prediction model to predict the future motion state, a process noise matrix is ​​introduced, and the observation data of the obstacle motion state is corrected. The measurement residual between the future motion state and the corrected observation data is calculated and combined with the Kalman gain to update the current state estimate:

[0023]

[0024] Among them, y k is the measurement residual, R is the measurement noise covariance matrix, K k is the Kalman gain, zk is the measurement vector, H is the measurement matrix, is the predicted state vector, is the forecast error covariance matrix, is the transpose of H.

[0025] Optionally, the discrete path nodes are corresponding optimal actions for avoiding the dynamic obstacle;

[0026] The action combination includes: continuous forward movement and simultaneous left and right yaw and up and down pitching.

[0027] Optionally, the path planning model includes: after the input layer of the dual-depth Q network model, a multi-layer fully connected structure is used to replace the Double-DQN value network to extract target features, and a Leaky ReLU activation function is used after each layer of full connection; the multi-layer fully connected structure is a fully connected layer with different numbers of neurons, and the input result of the last layer of fully connected structure is the same as the dimension of the action space, then the output layer outputs a vector with the same dimension as the action space, that is, the corresponding optimal action.

[0028] Optionally, obtaining the corresponding optimal action includes:

[0029]

[0030] Among them, Q(s t ,a t ) is the corresponding optimal action, γ is the discount factor, r t For instant rewards, Q target is the Q value output by the target network, Indicates using the policy network to select the next state s t+1 The optimal action under s t is the current state, a t is the current action, a′ is the candidate action in the next state, and θ is the policy network parameter.

[0031] Optionally, the comprehensive reward function includes:

[0032] Target distance bonus items:

[0033] R d =-k1d

[0034] Among them, R d represents the reward based on the target distance, and d is the current AUV position P r and the target point position P t , the Euclidean distance between them, k1 represents the step penalty factor;

[0035] Ocean Current Bonus:

[0036] R o =k2·cosθ

[0037]

[0038] Where k2 is the ocean current scaling factor, cosθ is the angle between the ocean current directions, c is the fixed vector in the direction of the ocean current, m·c represents the dot product of the two vectors, and ||m|| and ||c|| are the moduli of the two vectors respectively.

[0039] Bonus item for cosine similarity:

[0040] R o =k2·cosθ

[0041] Among them, k2 is the ocean current expansion and contraction factor;

[0042] Rewards for triggering collision events:

[0043] R c =-500

[0044] Among them, R c A reward for triggering a collision event when the distance between the unmanned underwater vehicle and the dynamic obstacle is less than a target value;

[0045] Terminal rewards:

[0046] Rt=1000

[0047] Among them, Rt is the reward for the unmanned underwater vehicle reaching the end point.

[0048] Optionally, obtaining the three-dimensional path planning result includes:

[0049] The target point on the Basis spline curve is determined by performing weighted averaging based on the discrete path nodes and the corresponding basis functions:

[0050]

[0051] Among them, P i is the discrete path node, n is the number of discrete control nodes;

[0052] Based on the target point on the Basis spline curve, the three-dimensional path planning result is obtained; in the process of determining the target point on the Basis spline curve, each discrete path node is adjusted by the basis function weight corresponding to the node:

[0053]

[0054] Among them, u is the element of the non-uniform periodic vector, u i and u i+pare the i-th and i+p-th nodes in the node vector, respectively, N i+1,p-1 (u) is the B-spline basis function of order p-1 on the i+1th segment.

[0055] The beneficial effects of the present invention are:

[0056] To address the technical difficulties faced by under-actuated AUVs in generating safe and efficient trajectories in real time in three-dimensional dynamic environments due to limited navigation and control computing power when performing missions far away from the mother ship, the present invention adopts a technical chain consisting of "improved DDQN-deep reinforcement learning + IMM-EKF dynamic obstacle prediction + Basis spline smoothing": first, through dual-network over-estimation and efficient reward design, the computational burden of the high-dimensional state-action space is compressed; second, the Singer model and extended Kalman filter are used to jointly predict the position, velocity and acceleration of surrounding dynamic targets to achieve forward avoidance of random maneuvering obstacles; then, the AUV is guided to sail downstream with the help of ocean current direction coupling rewards, and the discrete decision sequence is smoothed with a spline function to meet the continuous posture constraints of the under-actuated AUV. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0058] Figure 1 This is a flow chart of a three-dimensional path planning method for an AUV in a dynamic obstacle environment according to an embodiment of the present invention;

[0059] Figure 2 A three-dimensional dynamic obstacle trajectory prediction map according to an embodiment of the present invention;

[0060] Figure 3 This is a distribution diagram of position prediction deviations in the three directions of XYZ according to an embodiment of the present invention;

[0061] Figure 4 Results of dynamic obstacle avoidance tests; (a) is the simulation result of the AUV encountering a static obstacle, (b) is the simulation result of the AUV encountering a dynamic obstacle that crosses laterally, (c) is the simulation result of the AUV encountering a static obstacle that moves in the same direction, and (d) is the simulation result of the AUV encountering a static obstacle that moves relative to the AUV.

[0062] Figure 5These are the test results of AUV three-dimensional path planning; (a) is the test result of AUV three-dimensional path planning based on the PRM algorithm, (b) is the test result of AUV three-dimensional path planning based on the RRT algorithm, (c) is the test result of AUV three-dimensional path planning based on the APF algorithm, and (d) is the test result of AUV three-dimensional path planning based on the DDQN algorithm. DETAILED DESCRIPTION

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0064] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] like Figure 1 As shown, this embodiment discloses a three-dimensional path planning method for an AUV in a dynamic obstacle environment, including: setting a three-dimensional state space, updating the motion state of a dynamic obstacle in the three-dimensional state space, and obtaining the future motion state of the dynamic obstacle based on the motion state; inputting the target position parameters and the future motion state of the three-dimensional state space into a path planning model to obtain discrete path nodes of the unmanned underwater vehicle; the path planning model is obtained by training using a training set combined with a comprehensive reward function; and the discrete path nodes are subjected to path continuity and smoothing processing to obtain a three-dimensional path planning result.

[0066] Specifically, this embodiment discloses a three-dimensional path planning method for an AUV in a dynamic obstacle environment, including: step 1: constructing a three-dimensional state space; step 2: defining an action space; step 3: designing a comprehensive reward function; step 4: using an extended Kalman filter (EKF) and a Singer model to predict the future motion state of a dynamic obstacle and generate an obstacle prediction trajectory; step 5: constructing a deep reinforcement learning network model based on an improved dual deep Q network (DDQN) algorithm; step 6: using the trained DDQN model to select the optimal action according to the real-time environmental state to realize real-time path planning of the AUV; step 7: using the basis spline function to perform path continuity and smoothing on the discrete path nodes output by the DDQN algorithm.

[0067] Furthermore, the three-dimensional state space includes: the AUV current position coordinates, the target position coordinates, the dynamic obstacle position coordinates, the fixed obstacle position coordinates and the ocean current vector.

[0068] Specifically, a three-dimensional state space is constructed, which includes the AUV current position coordinates, target position coordinates, dynamic obstacle position coordinates, fixed obstacle position coordinates and ocean current vector, and is expressed as:

[0069]

[0070] If the number of dynamic obstacles and fixed obstacles is insufficient, zero padding is used to ensure that the dimension of the state vector is consistent.

[0071] The current position coordinates of the AUV are expressed as:

[0072] P r =[x r ,y r ,z r ]

[0073] Used to describe the specific position of the AUV in the environment space at the current moment, x r ,y r ,z r Represent the xyz axis coordinates respectively.

[0074] The target position coordinates are expressed as:

[0075] P t =[x r ,y r ,z r ]

[0076] The dynamic obstacle position coordinates are expressed as:

[0077]

[0078] Where i = 1, 2, ..., N d , N d The maximum number of dynamic obstacles that can be detected. If the number of obstacles detected is less than N d , the state vector is padded with the zero vector [0,0,0] to ensure the consistency of the data dimension.

[0079] The coordinates of the fixed obstacle position are expressed as:

[0080]

[0081] Where i = 1, 2, ..., N f , N f Indicates the number of fixed obstacles in the environment.

[0082] The ocean current vector is a fixed-direction ocean current with a force direction of [-1,0,0], indicating that the ocean current exerts a constant force from the right side to the left side of the environment.

[0083] Furthermore, updating the motion state of the dynamic obstacle in the three-dimensional state space includes: determining a maneuver acceleration time correlation function of the dynamic obstacle, and using Gaussian white noise with a target mean and a target variance in the Singer model combined with the maneuver acceleration time correlation function to obtain the acceleration of the dynamic obstacle.

[0084] Furthermore, obtaining the future motion state of the dynamic obstacle includes: inputting the motion state into a state prediction model to obtain the future motion state of the dynamic obstacle.

[0085] Furthermore, in the process of inputting the motion state into the state prediction model to predict the future motion state, the process noise matrix is ​​introduced, and the observation data of the obstacle motion state is corrected at the same time. The measurement residual between the future motion state and the corrected observation data is calculated and combined with the Kalman gain to update the current state estimate.

[0086] Specifically, the extended Kalman filter and Singer model are used to predict the future motion state of dynamic obstacles and generate obstacle prediction trajectories.

[0087] In the Extended Kalman Filter (EKF), the state of an obstacle is described as a three-dimensional vector, consisting of position, velocity, and acceleration, denoted as s = [x, v, a]. To dynamically predict the obstacle's trajectory, the EKF constructs state transition equations and measurement update equations based on the kinematic model. Its core computational modules include a state prediction model, a measurement update model, and a state covariance update.

[0088] The state prediction model is implemented through the state transfer matrix F1, which is in the form of:

[0089]

[0090] Where Δt is the time step, which represents the discrete update period of the obstacle motion state. Based on this model, the state of the next moment is predicted

[0091]

[0092] in, Estimated value for the current state.

[0093] In order to consider the random disturbances in the environment, the process noise matrix Q is introduced into the prediction process, which is defined as:

[0094] Q=diag(σ 2 ,σ 2 ,σ 2 )

[0095] Where σ is the standard deviation of the process noise, which is used to simulate the uncertainty of the obstacle motion.

[0096] The measurement update model is based on the observation data z=[x m ,v m ,a m ,] for correction; the measurement equation adopts linear form, and the measurement matrix is:

[0097] H=I 3×3

[0098] By extending the Kalman filter algorithm, the measurement residual yk between the predicted state and the observed data is calculated and combined with the Kalman gain K k Update the state estimate:

[0099]

[0100]

[0101] Wherein, R is the measurement noise covariance matrix. In the present invention, R=diag(0.01, 0.01, 0.01) is set to reflect the measurement capability of the high-precision sensor on the obstacle status.

[0102] The state covariance update is used to quantify the uncertainty of the state estimate. The extended Kalman filter also dynamically updates the covariance matrix P:

[0103] P k =(IK k H)P k∣k-1

[0104] This updating process further improves the accuracy of state estimation by reducing the impact of measurement noise on the estimation results.

[0105] The Singer model is used to describe dynamic obstacles. In the Singer model, the maneuver acceleration time-dependent function of a dynamic obstacle is usually expressed in an exponential decay form, as follows:

[0106]

[0107] in, is the variance of acceleration, which indicates the uncertainty of acceleration; τ is the time constant, which controls the rate of acceleration decay. α is taken as 1 / τ. According to experience, α is 1 / 60 for smooth turns and α is 1 / 20 for complex turns.

[0108] Combined with the relevant function formula, in the Singh model, the mean is 0 and the variance is The Gaussian white noise w(t) is used to represent the acceleration of the obstacle. The specific model expression is as follows:

[0109]

[0110] Where: a(t) is the acceleration of the obstacle; α is the acceleration attenuation coefficient, which controls the attenuation characteristics of the acceleration over time.

[0111] The discrete state update formula of the Singh model is:

[0112]

[0113] The state transfer matrix F2 is:

[0114] Furthermore, the discrete path nodes are corresponding optimal actions for avoiding dynamic obstacles; wherein the action combination includes: continuous forward movement and simultaneous left and right yaw and up and down pitch.

[0115] Specifically, the action space is defined as a set of discrete actions, including continuous forward movement and simultaneous left and right yaw and up and down pitching, defined as: A = {[1,0,0], [1,0,1], [1,0,-1], [1,1,0], [1,1,1], [1,1,-1], [1,-1,0], [1,-1,1], [1,-1,-1]}. The components of the action vector [x, y, z] have clear physical meanings: x represents the movement along the AUV's forward direction, which is always 1, indicating continuous forward movement; y represents the AUV's yaw action, with values ​​of 1 or -1 corresponding to right or left yaw, respectively, and 0 indicating no yaw; z represents the AUV's pitch action, with values ​​of 1 or -1 indicating upward or downward pitch, respectively, and 0 indicating no pitch.

[0116] The action vector corresponds to a discrete set, and the AUV performs the corresponding displacement operation according to the selected action. Specifically, the displacement of each step is calculated by the following vector: deep neural network.

[0117] Δp=[x·Δd,y·Δd,z·Δd]

[0118] Among them, Δd is the unit displacement size, and the value of the action [x, y, z] directly determines the movement direction and amplitude of the AUV in three-dimensional space at the current moment.

[0119] Furthermore, the path planning model includes: after the input layer of the dual-depth Q-network model, a multi-layer fully connected structure is used to extract target features, and a Leaky ReLU activation function is used after each fully connected layer; the multi-layer fully connected structure is a fully connected layer with different numbers of neurons, and the input result of the last layer of the fully connected structure is the same as the dimension of the action space, and the output layer outputs a vector with the same dimension as the action space, that is, the corresponding optimal action.

[0120] Specifically, a deep reinforcement learning network model is constructed based on the improved dual deep Q network (DDQN) algorithm. In order to adapt to the large amount of state information in the three-dimensional environment, the following optimizations are made to the network structure: first, the input layer receives the state vector S from the environment; secondly, for situations where the state dimension is high and the features are more complex, a three-layer fully connected structure is adopted, with 256 neurons in the first layer, 128 neurons in the second layer, and the dimension of the third output layer is the same as the action space. The LeakyReLU activation function is used after each layer is fully connected. This can improve the model's ability to extract nonlinear features while maintaining a moderate network depth, and the Leaky ReLU retains a small gradient in the negative interval, avoiding the dead zone problem that causes the network to be unable to update. The output layer outputs a vector Q(s) of the same size as the action space, where each element represents the Q value of taking the corresponding action in a given state. The Q value update formula of DDQN is as follows:

[0121]

[0122] Where γ is the discount factor, r t It is an instant reward, Q target is the Q value output by the target network, Indicates using the policy network to select the next state s t+1 The target network evaluates the value of the action.

[0123] Furthermore, the comprehensive reward function includes: a target distance reward item, an ocean current reward item, a cosine similarity reward item, a collision event triggering reward item, and a terminal reward item.

[0124] Specifically, a comprehensive reward function is designed, which consists of a target distance reward term, an ocean current reward term, an obstacle collision penalty term, and a terminal reward term.

[0125] The reward function is expressed as:

[0126] R=R d +R c +R o +R t

[0127] The target distance bonus term is expressed as:

[0128] R d =-k1d

[0129] Among them, R d represents the reward based on the target distance, and the current AUV position is P r , the target point position is P t , the Euclidean distance between the two is d = ||P r -P t||, k1 represents the step penalty factor.

[0130] The current bonus is expressed as:

[0131] R o =k2·cosθ

[0132] Among them, k2 is the ocean current scaling factor, which is used to enhance the impact of the ocean current direction on the reward. If the AUV's movement direction is consistent with the ocean current direction, that is, the cosine value is close to 1, it means that the path length chosen by the AUV is closer to the straight-line distance at this time, and the reward will be higher; conversely, if the direction is opposite (that is, the cosine value is close to -1), the reward will be reduced. The movement direction corresponding to the current action of the agent is vector m, and its angle with the ocean current direction is calculated by cosine similarity:

[0133]

[0134] The direction of the ocean current is a fixed vector c = [-1, 0, 0], indicating that the ocean current exerts a constant force from the right side to the left side of the environment. m·c represents the dot product of the two vectors, and ||m|| and ||c|| are the moduli of the two vectors, respectively.

[0135] The reward term based on cosine similarity is defined as:

[0136] R o =k2·cosθ

[0137] k2 is the current scaling factor, which is used to enhance the effect of current direction on rewards. If the AUV's movement direction is consistent with the current direction (i.e., the cosine value is close to 1), the AUV will receive a higher reward if the path length chosen is closer to the straight line distance. Conversely, if the direction is opposite (i.e., the cosine value is close to -1), the reward will be reduced.

[0138] The obstacle collision penalty is that when the distance between the AUV and the dynamic obstacle is less than 20 within the detection range of the agent, a collision event is triggered, and the reward is set to:

[0139] R c =-500

[0140] The terminal reward item is R t =1000, used to improve the driving force of AUV to complete the task.

[0141] Furthermore, obtaining the three-dimensional path planning result includes: performing weighted averaging based on discrete path nodes and corresponding basis functions to determine the target point on the Basis spline curve; obtaining the three-dimensional path planning result based on the target point on the Basis spline curve; in the process of determining the target point on the Basis spline curve, each discrete path node is adjusted by the basis function weight corresponding to the node.

[0142] Specifically, the discrete path nodes output by the DDQN algorithm are continuous and smoothed using the Basis spline function to obtain a continuous and stable AUV navigation path. The Basis spline basis function is defined recursively. For a B-spline basis function of order 0 (i.e. linear), it is defined as follows:

[0143]

[0144] For basis functions of order p>0, the recursive relation is as follows:

[0145]

[0146] Where u is the element of the non-uniform periodic vector, u i and u i+p are the i-th and i+p-th nodes in the node vector respectively.

[0147] By discrete control nodes P0, P1, ..., P n , the Basis spline curve C(u) is defined as:

[0148]

[0149] That is, the points on the curve are determined by weighted average of the control points and the corresponding basis functions. Each control point is represented by its corresponding basis function weight N i,p (u) is adjusted, and the weight of the basis function changes with the change of parameter u.

[0150] Table 1 shows the proposed algorithm applied to AUV 3D path planning in a simulated environment. The algorithm is compared with other algorithms, such as PRM, RRT, and APF, in terms of path length and navigation time to verify the feasibility and robustness of the proposed algorithm. The data in the table show that the DDQN-based AUV 3D path planning algorithm performs best in terms of trajectory length, with the shortest path length reaching 1742.93 meters. This demonstrates the algorithm's high efficiency and effectiveness in path planning.

[0151] Table 1

[0152]

[0153] Figure 2 This is the three-dimensional dynamic obstacle trajectory prediction diagram of this embodiment. It is a simulation diagram of the Kalman filter algorithm in the present invention predicting the random trajectory of a dynamic obstacle, which shows the effectiveness of the EKF algorithm in processing dynamic obstacle prediction.

[0154] Figure 3The following table shows the distribution of prediction deviations in the X, Y, and Z directions, with the largest deviation in the Y-axis direction, reaching 1.24 meters. Considering the AUV's minimum turning radius of 15 meters and its local detection range of 50 meters, from a spatial perspective, this deviation accounts for only 8% of the minimum turning radius, well within the AUV's safe obstacle avoidance margin. From a temporal perspective, a 50-meter detection range provides a 25-40 second warning window. Combined with the continuous updates of the EKF, the error converges to approximately 0.5 meters. Therefore, the EKF algorithm can achieve sufficiently accurate dynamic obstacle predictions in practical applications, effectively supporting the AUV's obstacle avoidance decisions and path planning.

[0155] Figure 4 In unknown water environments, when an autonomous underwater vehicle (AUV) does not detect an obstacle, the deployed obstacle avoidance behavior network can output heading angle and speed instructions to guide the AUV to navigate towards the target position; once the sensor detects a static or dynamic obstacle threat, the obstacle avoidance behavior network generates an avoidance action based on the real-time dynamic planning results and immediately adjusts the AUV's heading angle and navigation speed, thereby achieving dynamic avoidance of the obstacle. Figure 4 (a)-(d) respectively show the simulation results when the AUV encounters a static obstacle, a dynamic obstacle crossing laterally, a static obstacle moving in the same direction, and a static obstacle moving relative to each other.

[0156] Figure 5 (a)-(d) show the application of the proposed algorithm to AUV three-dimensional path planning in a simulated environment, and are compared with PRM, RRT, APF and other algorithms in terms of path length and navigation time to verify the feasibility and robustness of the proposed algorithm.

[0157] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A three-dimensional path planning method for AUV in a dynamic obstacle environment, characterized by: include: Setting a three-dimensional state space, updating a motion state of a dynamic obstacle in the three-dimensional state space, and obtaining a future motion state of the dynamic obstacle based on the motion state; Inputting the target position parameters of the three-dimensional state space and the future motion state into a path planning model to obtain discrete path nodes of the unmanned underwater vehicle; the path planning model is obtained by training using a training set combined with a comprehensive reward function; Path continuity and smoothing processing are performed on the discrete path nodes to obtain a three-dimensional path planning result.

2. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 1 is characterized in that: The three-dimensional state space includes: the AUV current position coordinates, the target position coordinates, the dynamic obstacle position coordinates, the fixed obstacle position coordinates and the ocean current vector.

3. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 1 is characterized in that: Updating the motion state of the dynamic obstacle in the three-dimensional state space includes: Determine the maneuver acceleration time-dependent function of the dynamic obstacle: in, is the variance of acceleration, τ is the time constant; The acceleration of the dynamic obstacle is obtained by combining the Gaussian white noise with the target mean and target variance in the Singh model with the maneuver acceleration time correlation function: Where a(t) is the acceleration of the obstacle; α is the acceleration attenuation coefficient, and w(t) is Gaussian white noise.

4. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 1 is characterized in that: Obtaining the future motion state of the dynamic obstacle includes: The motion state is input into the state prediction model to obtain the future motion state of the dynamic obstacle: in, is the state at the next moment, is the current state estimate, and F1 is the state transfer matrix.

5. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 4 is characterized in that: In the process of inputting the motion state into the state prediction model to predict the future motion state, a process noise matrix is ​​introduced, and the observation data of the obstacle motion state is corrected at the same time. The measurement residual between the future motion state and the corrected observation data is calculated and combined with the Kalman gain to update the current state estimate: Among them, y k is the measurement residual, R is the measurement noise covariance matrix, K k is the Kalman gain, z k is the measurement vector, H is the measurement matrix, is the predicted state vector, is the forecast error covariance matrix, is the transpose of H.

6. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 1 is characterized in that: The discrete path nodes are corresponding optimal actions for avoiding the dynamic obstacles; The action combination includes: continuous forward movement and simultaneous left and right yaw and up and down pitching.

7. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 1 is characterized in that: The path planning model includes: replacing the Double-DQN value network with a multi-layer fully connected structure after the input layer of the dual-depth Q network model to extract target features, and using a Leaky ReLU activation function after each fully connected layer; the multi-layer fully connected structure is a fully connected layer with different numbers of neurons, and the input result of the last fully connected layer is the same as the dimension of the action space, and the output layer outputs a vector with the same dimension as the action space, that is, the corresponding optimal action.

8. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 7 is characterized in that: Obtaining the corresponding optimal action includes: Among them, Q(s t ,a t ) is the corresponding optimal action, γ is the discount factor, r t For instant rewards, Q target is the Q value output by the target network, Indicates using the policy network to select the next state s t+1 The optimal action under s t is the current state, a t is the current action, a′ is the candidate action in the next state, and θ is the policy network parameter.

9. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 1, characterized in that: The comprehensive reward function includes: Target distance bonus items: R d =-k1d Among them, R d represents the reward based on the target distance, and d is the current AUV position P r and the target point position P t , the Euclidean distance between them, k1 represents the step penalty factor; Ocean Current Bonus: R o =k2·cosθ Where k2 is the ocean current scaling factor, cosθ is the angle between the ocean current directions, c is the fixed vector in the direction of the ocean current, m·c represents the dot product of the two vectors, and ||m|| and ||c|| are the moduli of the two vectors respectively. Bonus item for cosine similarity: R o =k2·cosθ Among them, k2 is the ocean current expansion and contraction factor; Rewards for triggering collision events: R c =-500 Among them, R c A reward for triggering a collision event when the distance between the unmanned underwater vehicle and the dynamic obstacle is less than a target value; Terminal rewards: Rt=1000 Among them, Rt is the reward for the unmanned underwater vehicle reaching the end point.

10. The AUV three-dimensional path planning method for dynamic obstacle environments according to claim 1, characterized in that: Obtaining the three-dimensional path planning result includes: The target point on the Basis spline curve is determined by performing weighted averaging based on the discrete path nodes and the corresponding basis functions: Among them, P i is the discrete path node, n is the number of discrete control nodes; Based on the target point on the Basis spline curve, the three-dimensional path planning result is obtained; in the process of determining the target point on the Basis spline curve, each discrete path node is adjusted by the basis function weight corresponding to the node: Among them, u is the element of the non-uniform periodic vector, u i and u i+p are the i-th and i+p-th nodes in the node vector, respectively, N i+1,p-1 (u) is the B-spline basis function of order p-1 on the i+1th segment.