AUV (Autonomous Underwater Vehicle) docking path planning method and system under ocean current disturbance
Through the deep reinforcement learning method combined with the priority experience playback mechanism, the path planning problem of AUV in the ocean current disturbance and obstacle environment is solved, safe, efficient and energy-saving path planning is achieved, and the autonomy and stability of AUV are enhanced.
Patent Information
- Application Number
- CN202510672125.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-26
AI Technical Summary
The existing AUV path planning methods are difficult to achieve real-time and accurate path planning in the environment of current perturbations and complex obstacles. The traditional methods lack the ability to adapt dynamic environments, and the training efficiency of deep Q learning algorithms is low and the convergence is poor.
The deep reinforcement learning method that combines the Double DQN mechanism and the priority experience replay PER mechanism is used to build a path planning model, combine the first-order Gaussmarkov process to simulate current perturbation, design a comprehensive reward function, and perform path post-processing through the Bezier curve.
It improves the path planning efficiency of AUV in dynamic and complex marine environments, generates safe, smooth and efficient docking paths, reduces energy consumption, and enhances autonomy and independence.
Smart Images

Figure CN120540358A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an AUV return-docking path planning method and system, and in particular to an AUV return-docking path planning method and system under ocean current disturbance, belonging to the technical field of AUV path planning. Background Art
[0002] With the widespread application of autonomous underwater vehicles (AUVs) in ocean exploration, deep-sea exploration, and military missions, efficient path planning in dynamic ocean currents and complex obstacle environments has become a key issue in AUV development. Currently, the main challenge in AUV path planning lies in the complexity of the operating environment, especially under the influence of turbulent ocean currents and multiple obstacles. Traditional path planning methods often struggle to provide real-time, accurate solutions.
[0003] Existing path planning methods are primarily based on manually designed rules or pre-defined paths, lacking the ability to adapt to dynamic environments (such as fluctuating ocean currents and the random distribution of obstacles) and real-time perception. Traditional optimization methods, such as single-objective optimization based on weighted sums, while effective in static environments, often fail to meet the requirements of real-time performance and efficiency in dynamic environments, particularly those affected by dynamic ocean currents and complex obstacles.
[0004] With the rise of deep reinforcement learning (DRL) technology, AUVs can autonomously learn and optimize their control strategies through interaction with their environment, overcoming the limitations of traditional path planning methods. However, in practical applications, existing deep Q-learning algorithms (such as basic DQN) still suffer from low training efficiency and poor convergence, especially under the influence of multiple obstacles and dynamic ocean currents. Summary of the Invention
[0005] Purpose of the invention: The purpose of the present invention is to provide a method and system for AUV docking path planning under ocean current disturbances that can improve the efficiency of AUV path planning.
[0006] Technical solution: The present invention provides a method for planning an AUV's path back to dock under ocean current disturbances, comprising:
[0007] S1: Obtain obstacle and ocean environment data, build an underwater 3D environment model and introduce ocean environment disturbances;
[0008] S2: Based on the underwater 3D environment model, determine the initial state and target state of the AUV according to the mission requirements, and construct a path planning reward function;
[0009] S3: Using the Dueling DQN architecture, we build a path planning model based on deep reinforcement learning. We train the path planning model by combining the DoubleDQN mechanism, the target network, and the Prioritized Experience Replay (PER) mechanism.
[0010] S4: Perform path planning based on the path planning reward function and the trained path planning model.
[0011] Furthermore, the step S1 includes:
[0012] Obtain the locations of obstacles in the underwater environment and construct an underwater three-dimensional environment model based on the locations of the obstacles;
[0013] The marine environment data is introduced to improve the underwater three-dimensional environmental model. The marine environment disturbance is simulated and controlled according to the first-order Gauss-Markov process to optimize the underwater three-dimensional environmental model. The first-order Gauss-Markov process is:
[0014] X t+1 =AX t +W t
[0015] Among them, X t is the state vector, A is the state transfer matrix, W t is a zero-mean Gaussian white noise model, and t is a discrete time step.
[0016] Furthermore, the reward function R(s,a) in step S2 consists of six terms:
[0017] R(s,a)=k1R dist +k2R obs +k3R cur +k4R step +k5R energy +T done R goal
[0018] Among them, R dist is the distance reward; R obs is the collision penalty; R cur Is the ocean current reward; R step is the time step penalty; R energy is the energy consumption penalty; R goal is the target reward; k1, k2, k3, k4 and k5 are weight coefficients; T done is the termination mark. When the round ends and the target is reached, T done =1, otherwise 0.
[0019] Furthermore, the distance reward R dist The definition is as follows:
[0020] Rdist =L euc (P t ,P goal )-L euc (P t+1 ,P goal )
[0021] Among them, L euc (P t ,P goal ) represents the distance between the current position and the target point at time t, P t is the current position at time t, P goal is the target location;
[0022] The collision penalty R obs The definition is as follows:
[0023]
[0024] Among them, C is the penalty factor; d min,i is the distance between the agent and the edge of the closest obstacle i; k is the adjustment parameter of exponential decay; d safe is the safety distance threshold;
[0025] The current rewards R cur The definition is as follows:
[0026]
[0027] Among them, V cur Indicates the speed of ocean current; V AUV represents the navigation speed of AUV, θ c It represents the horizontal angle between the direction of the ocean current and the direction of movement of the AUV.
[0028] Furthermore, step S3 includes:
[0029] Discretize the steering angle and speed control parameters of the AUV from a continuous space to a finite action space;
[0030] The action space consists of 12 discrete actions, which are mathematically defined as:
[0031] A={±e s |s∈{1,2,...,12}}
[0032] Among them, e s is the 6-DOF standard basis vector, corresponding to the 3 linear velocity degrees of freedom (Δv x ,Δv y ,Δv z ) and three angular velocity degrees of freedom (Δω x ,Δωy ,Δω z ) unit increment adjustment; using Dueling DQN architecture to build path planning network;
[0033] The path planning network is trained using the Double DQN mechanism, which uses two neural networks: an online network for action selection and a target network for evaluating the Q value of the best action selected in the next state. The target value calculation process of the target network is:
[0034]
[0035] Among them, θ online is the online network parameter, It represents the optimal action selected by the online network in state S, which is the Q value function output by the target network. target is the target network parameter, γ is the discount factor, s′ is the next state, a′ is the executed action, and r is the immediate reward;
[0036] The parameters of the target network are synchronized from the online network through the Polyak soft update strategy, and the update formula is:
[0037] θ target ←τθ online +(1-τ)θ target
[0038] Where τ is the soft update coefficient, which is used to ensure that the target network parameters smoothly follow the online network changes;
[0039] The Prioritized Experience Replay (PER) mechanism is used to calculate the TD error and use the target network to estimate the maximum Q value in the next state:
[0040]
[0041] Where r is the immediate reward for taking an action in the current state; γ is the discount factor; Q(s′,a′) is the Q value of the next state s′ performing action a′; Q(s,a) is the Q value of the next state s performing action a;
[0042] The priority experience replay PER mechanism is used to improve sample utilization. The priority calculation formula is:
[0043] p i =(|δ i |+∈) α
[0044] Among them, p i Indicates the priority of the i-th sample; δ irepresents the TD error of the i-th sample; ∈ is used to smooth and avoid zero division errors or numerical instability when the error is zero; α is the power of the control error, which is used to determine the speed at which the priority increases as the error increases;
[0045] Calculate sampling probability and introduce importance sampling weights to correct bias;
[0046] During the training process, the exploration rate ∈ t , the update formula is:
[0047] ∈ t =∈ min +(∈ max -∈ min )×e -λ×episode
[0048] Among them, ∈ max is the initial exploration rate, ∈ min is the final exploration rate, λ is the exploration rate attenuation factor, and episode is the current training round.
[0049] Furthermore, the path planning network includes an input layer, a shared feature extraction layer and an output layer;
[0050] The shared feature extraction layer outputs two branches, one of which is a value flow network for estimating the global value of the state, and the other is an advantage flow network for evaluating the local advantage of each discrete action;
[0051] The outputs of the value flow network and the advantage flow network are synthesized into the final Q-value output layer through the Q-value addition layer:
[0052]
[0053] in, represents the size of the action space, A(s,a) represents the advantage function, V(s) represents the advantage function, and Q(s,a) represents the advantage of each action relative to the average level;
[0054] Each layer uses the Leaky ReLU activation function, which is defined as:
[0055]
[0056] Among them, x is the signal passed through the previous layer in the network; α is a small constant.
[0057] Furthermore, the step S4 includes:
[0058] Collect sample data through the interaction between AUV and the environment, and perform training based on the sample data;
[0059] Dynamically update reward function parameters according to task requirements;
[0060] After obtaining the preliminary planned path, the Bezier curve formula is used for post-processing to complete the path planning.
[0061] Furthermore, the post-processing using the Bezier curve formula includes:
[0062] Select the starting point, end point and inflection point from the preliminary planned path as the control points of the Bezier curve;
[0063] The formula for defining the n-order Bezier curve is as follows:
[0064]
[0065] Among them, B(k) is the n-order Bezier curve, P i is the control point;
[0066] During the smoothing process, collision detection is performed on the generated curve to ensure that the smoothed path does not collide with obstacles or task boundaries. If necessary, local adjustments are made to the control points or curves.
[0067] Based on the same inventive concept, the present invention also provides an AUV docking path planning system under ocean current disturbances, comprising:
[0068] Initialization module, used to obtain obstacle and ocean environment data, build underwater 3D environment model and introduce ocean environment disturbance;
[0069] The reward function module is used to determine the initial state and target state of the AUV based on the underwater 3D environment model and the mission requirements, and to construct a path planning reward function;
[0070] The model building module is used to build a path planning model based on deep reinforcement learning using the Dueling DQN architecture. It combines the Double DQN mechanism, the target network, and the Prioritized Experience Replay (PER) mechanism to train the path planning model.
[0071] The planning module is used to perform path planning based on the path planning reward function and the trained path planning model.
[0072] Based on the same inventive concept, the present invention also provides a computing device, comprising: one or more processors, one or more memories, and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the AUV docking path planning method under ocean current disturbances are implemented according to any one of the above items.
[0073] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: 1. The present invention solves the problems of traditional AUV path planning methods that are prone to falling into local optimality, uneven path, and high energy consumption in environments with ocean current disturbances and complex obstacles; 2. The present invention adopts an improved deep reinforcement learning strategy to enable AUV to autonomously learn and optimize path planning based on real-time perception information in dynamic and complex ocean environments, combining Dueling DQN and Double The DQN architecture is integrated with the PER mechanism, which significantly improves the network's convergence speed and decision robustness, enabling the AUV to quickly adapt to environmental changes and generate a safe and efficient return path. 3. The reward function adopted by the present invention comprehensively considers factors such as distance, collision risk, energy consumption, ocean current compliance, and time step, effectively guiding the intelligent agent to form a planning strategy that goes with the flow and avoids detours, which not only reduces energy consumption but also optimizes task execution efficiency. After training on a large number of state-action pairs, the present method has acquired strong generalization ability and can achieve stable and effective path planning even in unknown or changeable ocean environments. 4. The present invention uses Bezier curves to post-process the preliminary planned path, making the output path continuous and smooth at turns, greatly improving the navigation stability and safety of the AUV, while reducing dependence on external navigation systems and enhancing the autonomy and independence of the AUV. 5. Through the organic combination of multiple technical means, efficient, smooth, and energy-saving AUV return path planning is achieved in complex ocean current disturbances and obstacle environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is a flow chart of a method according to an embodiment of the present invention;
[0075] Figure 2 Schematic diagram of obstacle distribution and ocean current disturbance in an underwater three-dimensional environment according to an embodiment of the present invention;
[0076] Figure 3 Schematic diagram of an algorithm framework based on deep reinforcement learning according to an embodiment of the present invention;
[0077] Figure 4 This is a flowchart of the deep reinforcement learning simulation training according to an embodiment of the present invention;
[0078] Figure 5 Schematic diagram of the convergence of the reward function during the training process of an embodiment of the present invention;
[0079] Figure 6 This is a schematic diagram of the route planned after considering the influence of ocean currents in an embodiment of the present invention. DETAILED DESCRIPTION
[0080] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0081] Example 1, as attached Figure 1 As shown, the AUV return-to-docking path planning method under ocean current disturbance in this embodiment includes:
[0082] S1: Obtain obstacle and ocean environment data, build an underwater 3D environment model and introduce ocean environment disturbances;
[0083] S2: Based on the underwater 3D environment model, determine the initial state and target state of the AUV according to the mission requirements, and construct a path planning reward function;
[0084] S3: Using the Dueling DQN architecture, we build a path planning model based on deep reinforcement learning. We train the path planning model by combining the DoubleDQN mechanism, the target network, and the Prioritized Experience Replay (PER) mechanism.
[0085] S4: Perform path planning based on the path planning reward function and the trained path planning model.
[0086] Specifically, in step S1, in order to achieve autonomous docking of the AUV in a complex ocean environment, this embodiment uses underwater sensors to collect obstacle location information, constructs an underwater three-dimensional environment model through multi-beam forward-looking sonar, introduces real ocean measurement data, and uses a first-order Gauss-Markov process to model the ocean flow field disturbance. Specifically, the discrete form of this process is expressed as
[0087] X t+1 =AX t +W t (1)
[0088] In formula (1), X t is the state vector, A is the state transfer matrix, W t is a zero-mean Gaussian white noise model. By assigning the corresponding ocean current velocity vector (v x ,v y ), this embodiment realizes the true reproduction of the obstacle distribution and dynamic ocean current disturbance in the actual sea area, such as Figure 2 As shown, this provides a solid simulation foundation for subsequent path planning.
[0089] In step S2, the initial and target states of the AUV are determined, and a path planning reward function is established.
[0090] The docking path planning in this embodiment is performed in a three-dimensional underwater environment. The initial and target states consider not only the AUV's position but also its attitude. Specifically, during the mission, the AUV sends a recovery request and its own status information to the recovery platform, which in turn transmits the preset recovery position information. To meet actual docking requirements, this embodiment requires that the initial state matches the AUV's actual position and attitude, while the target state is set to a certain distance from the recovery platform, with its attitude pointing toward the platform's center.
[0091] To this end, this embodiment constructs a state space that incorporates multi-source observation information. The state vector not only contains the relative position between the AUV and the target, but also integrates ocean current disturbance data at the current position and the obstacle distance matrix acquired by sonar. To facilitate path planning and control, the AUV's continuous control parameters (such as steering angle and propulsion commands) are discretized into a number of fixed actions in this embodiment, forming a discrete action space.
[0092] To ensure that the planned return path is safe, efficient, and smooth, a composite reward function is designed in this embodiment. This reward function comprehensively considers six factors: distance, collision and boundary collision, energy consumption, ocean current influence, time cost, and task completion reward. In this way, during the path planning process, the AUV is encouraged to continuously approach the target, effectively avoid obstacles, save energy, and balance the requirements of various objectives.
[0093] Specifically, the components of the reward function are as follows:
[0094] (1) Distance Reward
[0095] The reward is determined based on the Euclidean distance between the AUV and the target dock to encourage the agent to reduce the distance to the target in each time step and promote the path planning to move in the optimal direction, which is defined as
[0096] R dist =L euc (P t ,P goal )-L euc (P t+1 ,P goal ) (2)
[0097] In formula (2), L euc (P t ,P goal ) represents the distance between the current position and the target point at time t.
[0098] (2) Collision Penalty
[0099] When the distance between the AUV and the obstacle or mission boundary is lower than the set safety threshold d safe , a negative reward is given; if a collision occurs, a larger negative value is directly given, defined as
[0100]
[0101] In formula (3), C is the penalty factor, which determines the intensity of the penalty; d min,i is the distance between the agent and the edge of the closest obstacle; k is the adjustment parameter of exponential decay; d safe is the safety distance threshold.
[0102] (3) Ocean Current Rewards
[0103] According to the angle θ between the AUV heading and the ocean current direction at the current position c Give reward, the expression is
[0104]
[0105] In formula (4), cos(θ c ) is the current reward coefficient. This reward encourages the AUV to travel with the current. When the current direction is favorable for the AUV to move toward the target, the cosine value is large, and a positive reward is obtained; otherwise, a penalty signal is generated.
[0106] (4) Energy consumption penalty
[0107] In order to encourage the selection of paths with lower energy consumption, the moving distance of the AUV in each time step is penalized. This penalty term makes longer paths or actions requiring high power propulsion receive larger negative rewards, thereby guiding the agent to choose a more energy-efficient trajectory. Its form can be expressed as
[0108] R energy =-κΔs t (5)
[0109] In formula (5), Δs t is the moving distance of the AUV in the time step, and κ is the energy consumption coefficient.
[0110] (5) Time step penalty
[0111] In order to promote the rapid completion of the task, a fixed negative reward R is given to each time step step , so that the agent tends to choose the path that can shorten the total task duration.
[0112] (6) Rewards for reaching the target
[0113] When the AUV successfully reaches the preset target dock, a one-time fixed positive reward R is given. goal, this reward serves as the ultimate incentive for task completion.
[0114] Combining the above sub-rewards, the present invention defines the total reward function as
[0115] R(s,a)=k1R dist +k2R obs +k3R cur +k4R step +k5R energy +T done R goal (6)
[0116] In formula (6), T done is the termination flag, which takes 1 when the AUV successfully returns to the dock, and 0 otherwise; coefficients such as k2 and k4 are the weights of reward items such as collision and smoothing, respectively. Each weight parameter is set and optimized according to specific mission requirements and experimental data.
[0117] To ensure the reward function balances various metrics in simulation experiments, this embodiment further normalizes each reward item and adjusts the weights based on different task scenarios. For example, when the task requires high safety, k2 takes a larger value; when minimizing energy consumption, the weights of k3 and k5 are relatively increased. Through comprehensive design, the weighted reward function of the present invention can simultaneously reflect path distance, collision risk, energy consumption level, ocean current utilization, and task completion time, effectively guiding the agent to learn a safe, energy-efficient, smooth, and efficient re-docking path.
[0118] In step S3, an AUV path planning model based on deep reinforcement learning is constructed, including:
[0119] Discretize the steering angle and speed control parameters of the AUV from a continuous space to a finite action space;
[0120] The action space consists of 12 discrete actions, which are mathematically defined as:
[0121] A={±e s |s∈{1,2,...,12}}
[0122] Among them, e s is the 6-DOF standard basis vector, corresponding to the 3 linear velocity degrees of freedom (Δv x ,Δv y ,Δv z ) and three angular velocity degrees of freedom (Δω x ,Δω y ,Δω z ) unit increment adjustment;
[0123] like Figure 3 and Figure 4 As shown, this embodiment uses a deep reinforcement learning-based strategy for path planning. To achieve autonomous docking of the AUV under ocean current disturbances, a neural network model based on the Dueling DQN architecture, namely the path planning network, is constructed. The model includes the following modules:
[0124] a. Input layer: The input layer of the model receives the state vector collected by the AUV state space. This vector has a dimension of 12 and contains various information such as the AUV's position information, attitude information, ocean current disturbance data, and the obstacle distance matrix obtained by sonar detection.
[0125] b. Shared Feature Extraction Layer: This layer consists of two fully connected layers, each of which uses a Leaky ReLU activation function to abstract the input features. Through this shared layer, the model extracts high-dimensional features that are discriminative for path planning tasks. This feature representation helps separate the value of subsequent states from the advantages of actions.
[0126] Value stream branch: The features output by the shared layer are further passed to the value stream branch, which consists of a fully connected layer. Its output is the state value function V(s), which is used to evaluate the basic value in state s without considering the specific action.
[0127] Dominance Stream: Shared features are fed simultaneously into the dominance stream, which also consists of several fully connected layers. Its output is the dominance value A(s, a) for each discrete action. The dominance stream is primarily used to characterize the relative performance of each action relative to the average level within a state.
[0128] c. Output layer: The outputs of the value stream and advantage stream are synthesized into the final Q value in the output layer according to the following formula, which is defined as:
[0129]
[0130] In formula (7), represents the size of the action space, A(s,a) represents the advantage function, V(s) represents the advantage function, and Q(s,a) represents the advantage of each action relative to the average level. This fusion operation ensures that the advantage value has zero mean, so that the state value V(s) can reflect the average level of all actions in that state.
[0131] To reduce the problem of Q-value overestimation, this embodiment introduces the Double DQN strategy. Specifically, during the training process, two neural networks are used: the first neural network (online network) is used for action selection, and the second neural network (target network) is used to evaluate the Q-value of the best action selected in the next state. The target value calculation process is defined as:
[0132]
[0133] In formula (8), θ online is the online network parameter, Q target is the target network parameter, γ is the discount factor, s′ is the next state, and r is the immediate reward.
[0134] The parameters of the target network are synchronized from the online network through the Polyak soft update strategy, and the update formula is:
[0135] θ target ←τθ online +(1-τ)θ target (9)
[0136] In formula (9), τ is the soft update coefficient, and its value is usually small to ensure that the target network parameters smoothly follow the online network changes.
[0137] Furthermore, to improve training efficiency and accelerate convergence, this embodiment adopts the Prioritized Experience Replay (PER) mechanism. For each experience sample, its TD error is calculated:
[0138]
[0139] In formula (10), r is the immediate reward for taking an action in the current state; γ is the discount factor; Q(s′,a′) is the Q value of the next state s′ performing action a′; Q(s,a) is the Q value of the next state s performing action a. The priority of the sample is defined as:
[0140] p i =(|δ i |+∈) α (11)
[0141] In formula (11), ∈ is a small constant to prevent zero priority, and α is a tuning parameter.
[0142] The sampling probability P(i) of each sample is given by p i The formula for decision is:
[0143]
[0144] In formula (12), N is the total number of samples in the experience pool.
[0145] Given, at the same time introduce the importance sampling weight w i , the formula is:
[0146]
[0147] In formula (13), N is the total number of samples in the current experience pool, and β is a parameter that gradually increases to 1 to correct the deviation caused by non-uniform sampling.
[0148] The network training loss function uses the mean square error (MSE), and the loss of batch samples is defined as
[0149]
[0150] In formula (14), N is the number of batch samples, y i is the target Q value, Q(s i ,a i θ online ) is the Q value currently predicted by the network.
[0151] The Adam optimizer is used to adaptively adjust the learning rate, and the online network parameters are updated through the back-propagation algorithm.
[0152] In step S4, path planning is performed based on the reward function and the deep reinforcement learning model. During the training phase, the AUV agent continuously interacts in the simulation environment. In each training round, the AUV is placed at a preset starting point, and its initial state includes position and posture. The agent observes the current state at each time step t and selects an action according to the ε-greedy strategy. Specifically, an action is randomly selected with probability ∈; and the action that maximizes the probability 1-∈ is selected, that is:
[0153]
[0154] After executing the action, the environment updates its state to s according to the AUV motion and ocean current disturbance t+1 , and calculate the immediate reward (according to formula (5)). State transition sample (s t ,a t ,r t ,s t+1 ) are stored in the experience replay pool. The experience pool uses a circular buffer mechanism, and the oldest samples are eliminated when the capacity reaches the upper limit. After every 20 steps, a batch of samples is extracted from the experience pool according to the PER weight for network training. The target value is calculated using formula (8), and the loss is calculated according to formula (14). Then, the Adam optimizer is used to perform gradient descent to update the online network parameters. After the update, the target network is soft-updated according to formula (9). As training continues, the exploration rate ∈ gradually decays until the network loss function converges and the cumulative reward tends to stabilize, thus obtaining a stable strategy.
[0155] After the training is completed, the trained strategy is deployed in the actual AUV. In the actual path planning process, the AUV observes the current state s at each time step, inputs the trained online network, and uses greedy decision-making to select the optimal action, that is,
[0156]
[0157] In formula (16), argmax means selecting the action a that maximizes the Q value. * .
[0158] Based on this, a series of discrete path points are generated to achieve safe re-docking from the starting point to the target docking station. To eliminate potential unevenness in the discrete path, this embodiment further employs Bezier curve post-processing on the discrete path. Specifically, several key path points are selected from the preliminary planned path as control points, and the nth-order Bezier curve formula is used.
[0159]
[0160] In formula (17), B(k) is the n-order Bezier curve, P i is a control point, k∈[0,1], used to interpolate the position of points along the curve (e.g., k=0 corresponds to the starting point, k=1 corresponds to the end point).
[0161] Each discrete path segment is smoothed and fitted. To ensure that the smoothed path meets safety requirements, the generated Bezier curve is discretely sampled and then collision detection is performed on the sampling points using the environment grid map. If necessary, local adjustments are made to the control points until the smoothed path does not collide with obstacles or boundaries globally.
[0162] Example 2: This example constructs an underwater three-dimensional simulation environment with a size of 100m×100m×100m, and comprehensively verifies the path planning scheme of the present invention. In the environment, in addition to randomly arranging 20 obstacles with a radius uniformly distributed between 3 and 6 meters, the system also uses a first-order Gauss-Markov process to generate dynamic ocean current disturbances to simulate the complex and non-stationary flow field characteristics underwater with a flow rate range of approximately 0 to 1m / s. The dynamic ocean current disturbances present different directions and amplitudes at each spatial point and change over time, thereby forcing the AUV to perceive and respond to local water flow information in real time during the path planning process, laying a solid foundation for the application of subsequent algorithms in actual underwater scenes.
[0163] In terms of state and motion design, the AUV's initial and target states are described using 12-dimensional vectors, where the first three dimensions represent spatial position, and the remaining dimensions involve heading, speed, and other control variables. The action space is designed to consist of 12 discrete actions, covering positive and negative motion in the x, y, and z directions, as well as heading adjustments. To ensure timely collision avoidance when approaching obstacles, the system further introduces an obstacle avoidance mechanism after the initial action selection, fine-tuning the initial actions. This allows the AUV to flexibly adjust its motion strategy when dealing with multiple factors such as environmental boundaries, obstacles, and ocean current disturbances.
[0164] In terms of deep reinforcement learning, this embodiment adopts a network architecture based on Dueling DQN, dividing the network into a value branch and an advantage branch, thereby more accurately evaluating the relative merits of different actions in each state. At the same time, combined with the Double DQN concept, it effectively reduces the overestimation problem that may occur during target value estimation. To further improve training effectiveness, the system adopts a Polyak soft update strategy to smoothly update the target network within each training cycle, which not only accelerates the convergence of network parameters but also improves the stability and robustness of the policy network in the later stages of training. During training, the system collects state transition samples generated by the interaction between the AUV and the environment and introduces a PER mechanism for weighted sampling to ensure that samples with large TD errors and rich information are more frequently used for network training. At the same time, the Adam optimizer is used to further accelerate parameter updates. The entire training process is carried out for 10,000 rounds, with a maximum time step of 500 steps per round.
[0165] like Figure 5 As shown in the figure, in order to smooth the reward changes, the reward of each round is taken as the average reward of the 300 groups nearby to smooth the curve. In the early stage of training, since the AUV has not yet fully mastered the information about obstacles, ocean currents and boundaries in the environment, its path planning behavior shows a strong exploratory nature, and therefore obtains higher immediate rewards; as the number of training rounds increases, the AUV gradually masters the characteristics of the environment, its path planning tends to be stable, and the reward curve gradually smoothes, indicating that the network parameters have achieved good convergence in the continuous iteration process, and at the same time reflects that a balance has been reached between exploration and utilization. Figure 6 As shown in the figure, after sufficient training, the AUV can make full use of the ever-changing ocean currents and obstacle distribution to plan a smooth and continuous optimal path. This path not only fully utilizes the positive driving effect of the ocean currents, but also effectively avoids obstacles and environmental boundaries, thereby achieving safe and efficient travel from the starting point to the end point.
[0166] Further analysis of the training process and results revealed that the optimal path occurred in round 9637, with a round duration of 8.01 seconds; an average round duration of 13.08 seconds; a total training time of 135,617.12 seconds; and a total length of the optimal path of 180.87 meters. These indicators fully demonstrate the superior performance of the present invention's solution in achieving efficient and stable path planning in dynamic underwater environments, providing strong support for the practical application of autonomous underwater vehicles.
[0167] Example 3, based on the same inventive concept, this embodiment provides an AUV docking path planning system under ocean current disturbance, comprising:
[0168] Initialization module, used to obtain obstacle and ocean environment data, build underwater 3D environment model and introduce ocean environment disturbance;
[0169] The reward function module is used to determine the initial state and target state of the AUV based on the underwater 3D environment model and the mission requirements, and to construct a path planning reward function;
[0170] The model building module is used to build a path planning model based on deep reinforcement learning using the Dueling DQN architecture. It combines the Double DQN mechanism, the target network, and the Prioritized Experience Replay (PER) mechanism to train the path planning model.
[0171] The planning module is used to perform path planning based on the path planning reward function and the trained path planning model.
[0172] Example 4, based on the same inventive concept, this embodiment provides a computing device, including: one or more processors, one or more memories and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the AUV docking path planning method under ocean current disturbances according to any one of the above items are implemented.
[0173] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is merely a specific embodiment of the present invention and is not intended to limit the present invention. The description may also be a reasonable combination of the features described in the above embodiments. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A method for planning an AUV's return path under ocean current disturbances, characterized in that: include: S1: Obtain obstacle and ocean environment data, build an underwater 3D environment model and introduce ocean environment disturbances; S2: Based on the underwater 3D environment model, determine the initial state and target state of the AUV according to the mission requirements, and construct a path planning reward function; S3: Using the Dueling DQN architecture, we build a path planning model based on deep reinforcement learning. We train the path planning model by combining the Double DQN mechanism, the target network, and the Prioritized Experience Replay (PER) mechanism. S4: Perform path planning based on the path planning reward function and the trained path planning model.
2. The AUV docking path planning method under ocean current disturbance according to claim 1, characterized in that: The step S1 comprises: Obtain the locations of obstacles in the underwater environment and construct an underwater three-dimensional environment model based on the locations of the obstacles; The marine environment data is introduced to improve the underwater three-dimensional environmental model. The marine environment disturbance is simulated and controlled according to the first-order Gauss-Markov process to optimize the underwater three-dimensional environmental model. The first-order Gauss-Markov process is: X t+1 =AX t +W t Among them, X t is the state vector, A is the state transfer matrix, W t is a zero-mean Gaussian white noise model, and t is a discrete time step.
3. The AUV docking path planning method under ocean current disturbance according to claim 1, characterized in that: The reward function R(s,a) in step S2 consists of six terms: R(s,a)=k1R dist +k2R obs +k3R cur +k4R step +k5R energy +T done R goal Among them, R dist is the distance reward; R obs is the collision penalty; R cur Is the ocean current reward; R step is the time step penalty; R energy is the energy consumption penalty; R goal is the target reward; k1, k2, k3, k4 and k5 are weight coefficients; T done is the termination mark. When the round ends and the target is reached, T done =1, otherwise 0.
4. The AUV docking path planning method under ocean current disturbance according to claim 3 is characterized in that: The distance reward R dist The definition is as follows: R dist =L euc (P t ,P goal )-L euc (P t+1 ,P goal ) Among them, L euc (P t ,P goal ) represents the distance between the current position and the target point at time t, P t is the current position at time t, P goal is the target location; The collision penalty R obs The definition is as follows: Among them, C is the penalty factor; d min,i is the distance between the agent and the edge of the closest obstacle i; k is the adjustment parameter of exponential decay; d safe is the safety distance threshold; The current rewards R cur The definition is as follows: Among them, V cur Indicates the speed of ocean current; V AUV represents the navigation speed of AUV, θ c It represents the horizontal angle between the direction of the ocean current and the direction of movement of the AUV.
5. The AUV docking path planning method under ocean current disturbance according to claim 1, characterized in that: The step S3 comprises: Discretize the steering angle and speed control parameters of the AUV from a continuous space to a finite action space; The action space consists of 12 discrete actions, which are mathematically defined as: A={±e s ∣s∈{1,2,...,12}} Among them, e s is the 6-DOF standard basis vector, corresponding to the 3 linear velocity degrees of freedom (Δv x ,Δv y ,Δv z ) and three angular velocity degrees of freedom (Δω x ,Δω y ,Δω z ) unit increment adjustment; using Dueling DQN architecture to build path planning network; The path planning network is trained using the Double DQN mechanism, which uses two neural networks: an online network for action selection and a target network for evaluating the Q value of the best action selected in the next state. The target value calculation process of the target network is: Among them, θ online is the online network parameter, It represents the optimal action selected by the online network in state S, which is the Q value function output by the target network. target is the target network parameter, γ is the discount factor, s′ is the next state, a′ is the executed action, and r is the immediate reward; The parameters of the target network are synchronized from the online network through the Polyak soft update strategy, and the update formula is: i target ←tth online +(1-τ)θ target Where τ is the soft update coefficient, which is used to ensure that the target network parameters smoothly follow the online network changes; The Prioritized Experience Replay (PER) mechanism is used to calculate the TD error and use the target network to estimate the maximum Q value in the next state: Where r is the immediate reward for taking an action in the current state; γ is the discount factor; Q(s′,a′) is the Q value of the next state s′ performing action a′; Q(s,a) is the Q value of the next state s performing action a; The priority experience replay PER mechanism is used to improve sample utilization. The priority calculation formula is: p i =(|δ i |+∈) α Among them, p i Indicates the priority of the i-th sample; δ i represents the TD error of the i-th sample; ∈ is used to smooth and avoid zero division errors or numerical instability when the error is zero; α is the power of the control error, which is used to determine the speed at which the priority increases as the error increases; Calculate sampling probability and introduce importance sampling weights to correct bias; During the training process, the exploration rate ∈ t , the update formula is: ∈ t =∈ min +(∈ max -∈ min )×e -λ×episode Among them, ∈ max is the initial exploration rate, ∈ min is the final exploration rate, λ is the exploration rate attenuation factor, and episode is the current training round.
6. The AUV docking path planning method under ocean current disturbance according to claim 5, characterized in that: The path planning network includes an input layer, a shared feature extraction layer and an output layer; The shared feature extraction layer outputs two branches, one of which is a value flow network for estimating the global value of the state, and the other is an advantage flow network for evaluating the local advantage of each discrete action; The outputs of the value flow network and the advantage flow network are synthesized into the final Q-value output layer through the Q-value addition layer: in, represents the size of the action space, A(s,a) represents the advantage function, V(s) represents the advantage function, and Q(s,a) represents the advantage of each action relative to the average level; Each layer uses the Leaky ReLU activation function, which is defined as: Among them, x is the signal passed through the previous layer in the network; α is a small constant.
7. The AUV docking path planning method under ocean current disturbance according to claim 1, characterized in that: The step S4 comprises: Collect sample data through the interaction between AUV and the environment, and perform training based on the sample data; Dynamically update reward function parameters according to task requirements; After obtaining the preliminary planned path, the Bezier curve formula is used for post-processing to complete the path planning.
8. The AUV docking path planning method under ocean current disturbance according to claim 7, characterized in that: The post-processing using the Bezier curve formula includes: Select the starting point, end point and inflection point from the preliminary planned path as the control points of the Bezier curve; The formula for defining the n-order Bezier curve is as follows: Among them, B(k) is the n-order Bezier curve, P i is the control point; During the smoothing process, collision detection is performed on the generated curve to ensure that the smoothed path does not collide with obstacles or task boundaries. If necessary, local adjustments are made to the control points or curves.
9. A system for planning the path of an AUV returning to dock under ocean current disturbance, characterized in that: include: Initialization module, used to obtain obstacle and ocean environment data, build underwater 3D environment model and introduce ocean environment disturbance; The reward function module is used to determine the initial state and target state of the AUV based on the underwater 3D environment model and the mission requirements, and to construct a path planning reward function; The model building module is used to build a path planning model based on deep reinforcement learning using the Dueling DQN architecture. It combines the Double DQN mechanism, the target network, and the Prioritized Experience Replay (PER) mechanism to train the path planning model. The planning module is used to perform path planning based on the path planning reward function and the trained path planning model.
10. A computing device, characterized in that include: One or more processors, one or more memories, and one or more programs, wherein the programs are stored in the memories and configured to be executed by the processors, and when the programs are loaded into the processors, the steps of the AUV docking path planning method under ocean current disturbances are implemented according to any one of claims 1 to 8.