Vector propulsion AUV path planning method based on deep reinforcement learning

Through the vector-promoting AUV path planning method of deep reinforcement learning, combined with sonar and inertial navigation information, the SAC algorithm is optimized, which solves the problems of large amount of computing and insufficient maneuverability in AUV path planning, and achieves fast and energy-saving path planning.

CN120370940AActive Publication Date: 2025-07-25INST OF ACOUSTICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510491385.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing AUV path planning algorithms have large computational volume and poor generalization performance, and most of them ignore the influence of dynamics. Traditional single-control volume planning leads to insufficient maneuverability, and the joint control method has not been fully studied, resulting in suboptimal paths.

Method used

A vector propulsion AUV path planning method based on deep reinforcement learning is adopted to collect states in real time through sonar and inertial navigation information, and the action is output using the strategy network, and combined with reward function and dynamic equations, the state space, action space and reward function are designed, the SAC algorithm is optimized to accelerate convergence, and adaptive temperature parameters and batch normalization technology are introduced.

Benefits of technology

It realizes path planning with strong environmental adaptability and rapid convergence, optimizes energy consumption, and improves the maneuverability of AUV and the efficiency of reaching the target.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120370940A_ABST
    Figure CN120370940A_ABST
Patent Text Reader

Abstract

The invention discloses a vector propulsion AUV path planning method based on deep reinforcement learning, and the method comprises the steps: 1), collecting the current state st and action at of an AUV in real time based on sonar and inertial navigation information of the AUV, inputting a strategy network pi theta (atst), outputting a strategy action at, achieving the path planning of the AUV, calculating a reward value rt through a reward function, (2) when the number of tuples in the experience pool # imgabs 1 # is larger than a set value N, M tuples are sampled from # imgabs 2 #, and when the number of tuples in the experience pool # imgabs 1 # is larger than a set value N, M tuples are sampled from # imgabs 2 #; turning to step 3); otherwise, turning to the step 1); 3) under a maximum entropy reinforcement learning framework, updating a strategy network parameter theta, an evaluation network parameter # imgabs3 # and a temperature parameter alpha respectively; and 4) when the AUV meets the termination condition, ending, otherwise, turning to the step 1).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of AUV path planning, and particularly relates to a vector propulsion AUV path planning method based on deep reinforcement learning. Background Technique

[0002] Autonomous underwater vehicle (AUV) is an important tool for humans to explore the ocean. As the difficulty of AUV's mission execution increases day by day, higher requirements are put forward for its mobility and autonomy. Path planning technology, as one of the core technologies to achieve AUV autonomy, has received the attention of many researchers in recent years.

[0003] Path planning refers to planning a path from the starting point to the target point under certain constraints (such as dynamic constraints, environmental constraints, etc.) and meeting certain performance requirements (such as the shortest planned distance, the shortest planned time, etc.). There are already relevant patents studying the path planning problem of AUV. For example, the Chinese patent application with the publication number CN118857320A proposes a graph-based path planning algorithm, which determines local path points through the bidirectional A* * algorithm. This algorithm conducts bidirectional searches on the starting point and the target point, and improves the search efficiency compared with the traditional A* * algorithm. Subsequently, the DWA (Dynamic Window Algorithm) is used to connect local path points to avoid unknown obstacles in the environmental map. The Chinese patent application with the publication number CN118838357A proposes a sampling-based AUV path planning algorithm to solve the two-dimensional path planning problem in the ocean current environment. It uses the rapidly-exploring random tree algorithm to plan a low-energy consumption path, and by integrating ocean current information during random sampling, the algorithm can quickly reduce the energy consumption of the AUV with the help of ocean currents. The Chinese patent application with the publication number CN115202356B proposes an AUV path planning algorithm based on a heuristic optimization method. Based on the kinematic model of an underactuated AUV in a three-dimensional environment, it gives the hard constraints and soft constraints of the path planning problem, and searches for the optimal path through an improved ant colony algorithm. The improved algorithm uses adaptive parameters to prevent the algorithm from falling into local optima.

[0004] With the rapid development of deep reinforcement learning (DRL) technology, researchers are increasingly using DRL technology to solve the AUV path planning problem. The advantage of using DRL technology is that such methods learn by interacting with the environment, have strong environmental adaptability. At the same time, compared with heuristic-based optimization strategies, DRL pays more attention to the internal structure of the strategy. Coupled with the powerful representation ability of neural networks, it can solve relatively complex planning and control problems. For example, the Chinese patent application with the publication number CN118394089A uses the DQN algorithm in DRL technology to solve the high-dimensional AUV path planning problem. It integrates terrain information, inertial navigation error information, and target point information into the reward function, enabling the AUV to reach the target point faster. The Chinese patent application with the publication number CN113064422A also uses the DQN algorithm to solve the path planning problem, but it uses the priority experience replay technique when sampling from the experience pool to accelerate the convergence speed of the algorithm. The Chinese patent application with the publication number CN115790608A uses the SAC algorithm in DRL technology to solve the AUV path planning problem in a three-dimensional environment with ocean currents. In order to make the training model more suitable for the actual environment, it uses actual ocean current data, models the ocean current based on the Gaussian Markov model, and trains the algorithm in the modeled environment to improve the robustness and environmental adaptability of the algorithm.

[0005] Through research on the above existing technologies, it is found that traditional algorithms (graph-based algorithms, sampling-based algorithms, and heuristic optimization-based algorithms) more or less have problems such as large computational complexity and poor generalization performance. Moreover, most of them model problems at the kinematic level, ignoring the impact of the dynamics of AUVs on the algorithm. In addition, the planning environment of some algorithms is a two-dimensional grid, which requires using a certain method (such as cubic spline method) to smooth the planned path after the planning is completed. Although the algorithms based on deep reinforcement learning have a certain degree of environmental self-adaptability, there is still room for optimization in the convergence speed of the algorithm. In addition, the control input of most of the AUV path planning methods designed in patents is relatively single, only the rudder angle, which will result in poor maneuverability. There is little research on the AUV path planning problem using the combined control method of vector thrusters and rudders, and such a control method can effectively improve the maneuverability of AUVs.

[0006] Different from the conventional single-control variable path planning problem, the combined control method needs to balance the usage ratio of the two to avoid sub-optimal paths caused by overusing a certain control variable, which brings new challenges to the path planning problem. Summary of the Invention

[0007] In view of this, the present invention proposes a vector propulsion AUV path planning method based on deep reinforcement learning, including:

[0008] Step 1) Based on sonar and its own inertial navigation information, the current state s of the AUV is collected in real time t , according to the parameterized policy network π θ (a t |s t ) the action a is output t , the reward value r is calculated using the reward function t , the state s at the next moment is calculated using the AUV dynamics equation t+1 , the tuple [s t ,a t ,r t ,s t+1 is stored in the experience pool

[0009] Step 2) When the number of tuples in the experience pool is greater than the set value N, M tuples are sampled from ; go to Step 3); otherwise, go to Step 1);

[0010] Step 3) Under the framework of maximum entropy reinforcement learning, the policy network parameters θ and the evaluation network parameters and the temperature parameter α are updated respectively;

[0011] Step 4) When the AUV meets the termination condition, end; otherwise, go to Step 1).

[0012] Preferably, before the said Step 1), it further includes:

[0013] Based on sonar and its own inertial navigation information, the current state s of the AUV is collected in real time t , the action a t , and the state s at the next moment t+1 ;

[0014] According to the state s t , the action a t , the current reward r is calculated by the reward function t ;

[0015] The tuple [s t ,a t ,r t ,s t+1 is stored in the experience pool

[0016] The above steps are continuously carried out until the number of tuples in the experience pool is greater than the set value N.

[0017] Preferably, the policy network and the evaluation network are two independent neural networks; batch renormalization and bounded activation functions are used to connect between layers of each neural network; the action distribution is modeled using a Gaussian distribution, and the mean and standard deviation of the Gaussian distribution are finally output.

[0018] Preferably, the reward value r in step 1) t is:

[0019] r t = r z + r s + r d

[0020] where r z represents the terminal reward, a large positive feedback is given when reaching the target point, and a small negative feedback is given when colliding with an obstacle; r s represents the immediate reward, and r d represents the reward related to the rudder angle and geometric deflection angle.

[0021] Preferably, step 3) includes:

[0022] Step 3-1) Update the parameters of the evaluation network using as the loss function;

[0023] Step 3-2) Update the parameters θ of the policy network π π (a θ |s t ) using J t (θ) as the loss function;

[0024] Step 3-3) Update the temperature parameter α using J(α) as the loss function.

[0025] Preferably, the loss function in step 3-1) is:

[0026]

[0027] In the formula, |·| sg represents that its gradient is not required, represents the parameterized evaluation network, a t+1 represents the action at the next moment, s t+1 represents the state at the next moment, q t , q t+1 respectively represent that when the input of the evaluation network is [s t , a t , [s t+1 , a t+1The output value, γ represents the discount factor, γ ∈ [0, 1].

[0028] Preferably, the loss function J π (θ) in step 3-2) is:

[0029]

[0030] In the formula, D KL (·||·) is used to measure the difference degree between two strategies, π θ (·||s t ) represents the strategy distribution in state s t . represents the unnormalized probability distribution of the action-state value function in state s t . represents the normalization constant.

[0031] Preferably, the loss function J(α) in step 3-3) is:

[0032]

[0033] In the formula, represents a constant, π θ (a t |s t ) is the policy network. represents an expectation symbol, and the action a inside the expectation symbol t uses the parameterized policy π θ for sampling.

[0034] Preferably, the termination condition in step 4) includes: the actual number of running episodes is greater than the maximum number of running episodes hyperparameter T me .

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] 1. The present invention uses deep reinforcement learning technology to solve the path planning problem of vector propulsion AUV, models based on the dynamics of vector propulsion AUV, adopts an "end-to-end" control method, and makes full use of the characteristics of deep reinforcement learning technology to make the planned strategy have strong environmental adaptability.

[0037] 2. The present invention designs a state space, an action space, and a reward function that fit the path planning problem under joint control according to the characteristics of the problem, enabling the algorithm to converge quickly.

[0038] 3. For the SAC algorithm, an adaptive temperature parameter is introduced in this paper, the target evaluation network is eliminated, the batchnorm architecture is added, the sample efficiency of the algorithm is increased, and the convergence time of the algorithm is reduced. Brief Description of the Drawings

[0039] Figure 1 It is a schematic diagram of the path planning of the vector propulsion AUV;

[0040] Figure 2(a) is a schematic diagram of the path under different propulsion modes, and Figure 2(b) is the relative power required by the three when reaching the same speed;

[0041] Figure 3(a) is a schematic diagram of the AUV sonar model, and Figure 3(b) is an example of sonar detection data for a certain time;

[0042] Figure 4 It is a flowchart of the method of the present invention;

[0043] Figure 5 It is a graph showing the changing trend of the total reward of the ISAC algorithm of the present invention;

[0044] Figure 6 It is the path planned by the algorithm under Environment 1;

[0045] Figure 7 It is the rotation ratio of the rudder angle and the vector deflection angle under Environment 1;

[0046] Figure 8 It is the path planned by the algorithm under Environment 2;

[0047] Figure 9 It is the rotation ratio of the rudder angle and the vector deflection angle under Environment 2;

[0048] Figure 10 It is the path planned by the algorithm under Environment 3;

[0049] Figure 11 It is the rotation ratio of the rudder angle and the vector deflection angle under Environment 3. Detailed Implementation Manner

[0050] To solve the above problems, the present invention proposes a path planning method for a vector propulsion AUV based on deep reinforcement learning. First, it directly maps the state space of the AUV to the rudder plate deflection angle and the vector thruster deflection angle to control the AUV to reach the specified target point, and models the action space as a continuous space to avoid subsequent path smoothing steps; second, according to the turning characteristics of the rudder plate deflection angle and the vector deflection angle, the penalty for using the vector thruster is increased in the reward function. Finally, it optimizes the SAC algorithm in the DRL technology to accelerate the convergence speed of the algorithm.

[0051] The present invention aims to solve the path planning problem of the vector propulsion AUV in the marine environment. The schematic diagram of its path planning is as Figure 1 shown.

[0052] Figure 1 In the XOZ, Xu O u Z u represent the world coordinate system and the hull coordinate system respectively. The origin of the hull coordinate system is the center of buoyancy. The blue sector represents the detectable range of the sonar; and it is stipulated that the rotation angle from the X-axis to the Z-axis is a positive angle.

[0053] The AUV with vector propulsion in a two-dimensional environment can be modeled in the following form:

[0054]

[0055] In the above formula, x c represents the coordinate component of the center of gravity on the X u axis in the hull coordinate system. x0, z0, ψ represent the azimuth and attitude of the AUV in the world coordinate system, v x , v z , ω y represent the velocity and angular velocity of the AUV in the hull coordinate system, ρ represents the density of water, S represents the maximum cross-sectional area of the AUV, m represents the mass of the AUV, I yy represents the moment of inertia of the AUV, L represents the length of the AUV, λ 11 , λ 33 , λ 35 , λ 55 , C xS , represent the parameters related to hydrodynamic forces, represent the parameters related to hydrodynamic moments, δ r represents the vertical rudder angle, A(ξ) represents the thrust correction coefficient. ξ characterizes the geometric deflection angle of the vector thruster. When the nozzle of the vector thruster deflects, there will be a thrust loss, so thrust correction is required:

[0056]

[0057] After introducing the vector thruster, the turning radius of the AUV will be reduced, and the maneuverability of the AUV will be improved. However, due to the thrust loss, the speed during turning is reduced. As shown in Fig. 2(a), the sailing routes and distances traveled in three cases when the rudder angle and the vector geometric deflection angle take the maximum values at the same running time are respectively shown. Fig. 2(b) shows the relative power required to reach the same speed in three propulsion modes (the expressions are respectively characterize different propulsion modes). It is observed that although the use of the vector thruster reduces the turning radius of the AUV, the power required to reach the same speed is larger, resulting in more energy consumption for the same distance traveled. Therefore, when using both the vector nozzle and the rudder plate for dual control, it is necessary to balance the usage ratio of the two, so that the AUV can reach the target point as soon as possible while optimizing the energy consumption.

[0058] The sonar model in this paper adopts the ray model. As shown in Fig. 3(a), it is a schematic diagram of the AUV sonar model, and Fig. 3(b) is an example of sonar detection data. Among them, R represents the maximum detection radius of the sonar, Θ represents the maximum detection angle of the sonar, and d = [d1, d2, …, d i , …], 0 < i ≤ i max represents the distance between the sonar and the obstacle in the i-th direction.

[0059] Based on the above model and the characteristics of vector geometric angle turning, the following path planning problem can be established in this paper: that is, by adjusting the rudder angle and the vector geometric angle, and combining the sonar perception information, the AUV is guided to avoid obstacles and reach the target with the minimum energy consumption E. Therefore, it can be written in the following form:

[0060]

[0061] subject to:

[0062]

[0063] In the above formula, π θ (δ r , ξ|s) represents the probability that the AUV executes δ r , ξ in the state s, that is, the policy; x0(t), z0(t) represent the coordinates of the AUV in the world coordinate system at time t; x g , z g represent the target coordinates; x ob,i , z ob,i , r ob,i represent the coordinates and radius of the i-th obstacle; T max , δ max , ξ max , ζ represent the time limit, the rudder angle limit, the geometric angle limit, and the target threshold.

[0064] The above path planning problem can be modeled as a Markov decision process (MDP), which is represented by the tuple <S, A, P, R, P0, γ, H>. Among them, S represents the state space, A represents the action space, P represents the state transition function P(s t+1 |s t , a t ), R represents the reward function, P0 represents the initial state distribution, γ represents the discount factor, γ ∈ [0, 1], and H represents the length of each episode. For the MDP, the objective function of reinforcement learning can be expressed as:

[0065]

[0066] Among them, π θ (a t |s t) represents the policy parameterized by θ. In this paper, the maximum entropy reinforcement learning framework is adopted to solve the path planning problem of the vector propulsion AUV, enabling the AUV to have a certain exploration ability. The objective function under this framework can be written in the following form:

[0067]

[0068] In the above formula, α represents a hyperparameter, and H(·|s t ) represents the entropy of the action distribution in state s t . In this paper, the SAC (soft-actor-critic) algorithm will be used to optimize the policy parameter θ and improve it to accelerate the convergence speed of the algorithm.

[0069] The invention first establishes the dynamic equation of the vector propulsion AUV in the marine environment. The simulation verifies that the vector propeller can improve the maneuverability of the AUV, but at the same time, it will also increase the energy loss of the AUV. Secondly, based on the above dynamic model, the invention proposes a state space and an action space adapted to this problem, and increases the penalty for using the vector propeller in the reward function according to the characteristics of the vector propeller to control the unnecessary energy consumption caused by excessive use of the vector propeller. Finally, the invention updates the parameters of the policy network with the "soft actor-critic" algorithm (SAC) as the framework and improves it to accelerate the convergence speed of the algorithm. The simulation at the end of the paper verifies that the improved SAC algorithm (ISAC) of the invention can better balance the use ratio of the rudder angle and the vector deflection angle in various environments and complete the path planning task.

[0070] The vector propulsion AUV path planning method based on deep reinforcement learning of the present invention adopts the vector propulsion AUV path planning algorithm based on SAC, and the specific introduction is as follows:

[0071] The state space, action space, reward function, and parameter update method required by the SAC algorithm will be introduced below.

[0072] In this paper, the state space is designed into two parts, namely the state information of the vehicle itself and the information sensed by the sonar from the outside world

[0073]

[0074] Among them, δ r (t), ξ(t) represent the rudder angle and the vector deflection angle at time t; d(t) represents the sensing data of the sonar at time t; D represents the distance between the initial point and the target in the current episode; the motivation for dividing each state component by a constant is to map all state components to the interval [-1, 1], which helps with training generalization. d Δ (t), ψΔ (t) represents the distance and line-of-sight angle between the AUV and the target, which are expressed as follows:

[0075]

[0076]

[0077] Therefore, the final state vector can be expressed as follows:

[0078]

[0079] In this paper, an "end-to-end" control method is adopted to directly map the state space into the action space. Therefore, the action space is modeled in the following form:

[0080]

[0081] where The proportionality coefficients characterizing the rudder angle and the geometric deflection angle of the vector thruster can be used to obtain the actual output action after the following transformation:

[0082] a t ←a t ·[δ max ,ξ max

[0083] To accelerate the algorithm convergence, the reward function is designed in the following form:

[0084] r t =r z +r s +r d

[0085] where r z represents the terminal reward:

[0086]

[0087] That is, a large positive feedback is given when reaching the target point, and a small negative feedback is given when colliding with an obstacle. r s represents an immediate reward:

[0088] r s =a(d Δ (t - 1)-d Δ (t))-b

[0089] d Δ (t - 1),d Δ ​(t) represents the distances of the AUV from the target point at the previous moment and the current moment. The motivation for expressing it in the above form is to reward actions that bring the AUV closer to the target point, and a negative number (i.e., the hyperparameter b in the above formula) is added at each time step to control the AUV to reach the target point as soon as possible. r d Represents a reward related to the rudder angle and geometric deviation angle:

[0090] r d =-c(|δ r (t - 1)-δ r (t)|+|δ r (t)|)-d(|ξ(t - 1)-ξ(t)|+|ξ(t)|)

[0091] In the above formula, δ r (t - 1) and ξ(t - 1) represent the rudder angle and geometric deviation angle at the previous moment respectively. The motivation for adopting the above reward function is to hope that the control quantity is as smooth as possible. In addition, in the setting of hyperparameters, c < d needs to be taken because, as mentioned before, using vector propulsion will reduce the speed of the AUV. Therefore, the reward function needs to increase the penalty for using vector propulsion.

[0092] The SAC algorithm needs to parameterize the action-value function network as well as the policy network π θ (a t |s t ) and guide the algorithm to converge to the optimal policy by updating the parameters of the above two networks. The specific update strategy is as follows: The action-value function network is updated using the following loss function:

[0093]

[0094] where represents the target value function network, which is used to alleviate the unstable performance during the training process. The update method is as follows:

[0095]

[0096] In the above formula, τ represents a hyperparameter used to control the update speed of the target value network. In order to reduce the overestimation bias of the Q value during the update, a double Q-network architecture is used. The policy network π θ (a t |s t ) is updated using the following loss function:

[0097]

[0098] In the above formula Represents a normalization constant. To reduce the variance during training, the SAC algorithm utilizes the reparameterization trick:

[0099] a t = f θ (ε t ; s t )

[0100] In the above formula, ε t represents an input noise vector, which is sampled from a fixed distribution (such as a Gaussian distribution with mean 0 and variance 1), and substituting it into the loss function of the policy network π θ (a t | s t ), we can obtain:

[0101]

[0102] Different from traditional reinforcement learning algorithms, a temperature parameter α is introduced in the maximum entropy reinforcement learning framework. Generally, α can be treated as a fixed value, but this treatment method will make the policy maintain the same exploration in all state spaces. Therefore, the SAC algorithm improves the above problem by adaptively adjusting the temperature parameter:

[0103]

[0104] where, represents a constant, which is usually initialized as the opposite of the action dimension in practical applications.

[0105] However, there is still room for improvement in the convergence speed of the above SAC algorithm. Based on this, this paper introduces the CrossQ algorithm, which improves the algorithm from the following two parts:

[0106] First, the algorithm adopts a more "wide" network architecture, that is, the number of neurons in the hidden layer is increased; second, it deletes the target network architecture and uses batch normalization technology to stabilize the training of the algorithm and achieve the above effects. The improved surrogate loss function is as follows:

[0107]

[0108] where, |·| sg represents that its gradient is not required. In summary, the pseudo-code of the vector propulsion AUV path planning based on the ISAC algorithm of the present invention can be given as shown in Table 1:

[0109] Table 1 Vector Propulsion AUV Path Planning Algorithm Based on ISAC

[0110]

[0111] In the above table, N represents that when the number of tuples in reaches N, the update of subsequent parameters starts, M represents the number of samples to be taken from the experience pool each time, and T me , T mt respectively represent the maximum number of episodes for the algorithm to run and the maximum number of steps for each episode.

[0112] The method flow chart of the present invention is as shown in Figure 4 . Figure 4 The two networks in the lower middle of Figure 4 represent the evaluation network and the action network. Batch renormalization (BRN) and a bounded activation function (sigmoid) are used to connect between layers. The action distribution is modeled using a Gaussian distribution, and finally the mean and standard deviation of the Gaussian distribution are output.

[0113] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0114] Embodiment

[0115] The embodiment of the present invention provides a vector propulsion AUV path planning method based on deep reinforcement learning. The vector propulsion AUV path planning algorithm is simulated using python3 + pytorch. The hyperparameters of the algorithm are as shown in the following table, and the specific simulation results are shown in Table 2:

[0116] Table 2: Algorithm hyperparameter table

[0117] Figure 5 shows the change curve of the total reward for each episode during the training using the improved SAC algorithm (hereinafter referred to as the ISAC algorithm). It is obtained by changing the random number seed and performing 4 Monte Carlo simulations. The solid line represents the mean, and the colored area represents the variance. It is not difficult to find that the ISAC algorithm has a good convergence speed and stability.

[0118] Figure 6 is the path planned by the algorithm in Environment 1; Figure 7 is the rotation ratio of the rudder angle and the vector deflection angle in Environment 1; Figure 8 is the path planned by the algorithm in Environment 2; Figure 9 is the rotation ratio of the rudder angle and the vector deflection angle in Environment 2; Figure 10 is the path planned by the algorithm in Environment 3; Figure 11 is the rotation ratio of the rudder angle and the vector deflection angle in Environment 3. Figures 6 to 11The figure respectively shows the paths planned by using the ISAC algorithm in different environments and the vector thruster deflection angles and rudder plate deflection angles required to generate the above paths. Among them, the blue dots represent the starting point of the AUV, the orange pentagrams represent the target points that the AUV needs to reach, the yellow circles represent the obstacles in the environment, the blue solid lines represent the paths traveled by the AUV, and the physical meaning of the deflection ratio represents the ratio of the actual rotation angle of the deflection angle to the maximum deflection angle. It is not difficult to observe that the algorithm can reach the target point in all three environments. From the perspective of the rotation ratio of the rudder angle and the vector deflection angle, the usage ratio of the vector deflection angle is less than that of the rudder angle in all three environments, thus avoiding the reduction of the AUV's speed caused by excessive use of the vector deflection angle and improving the maneuverability of the AUV at the same time, enabling the AUV to reach the target point faster.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A path planning method for a vector propulsion AUV based on deep reinforcement learning, comprising: Step 1) Based on sonar and its own inertial navigation information, the current state s of the AUV is collected in real time t , according to the parameterized policy network π θ (a t |s t ) outputs the policy action a t , calculates the reward value rt using the reward function, calculates the state s at the next moment using the AUV dynamics equation t+1 , and stores the tuple [s t , a t , r t , s t+1 in the experience pool Step 2) When the number of tuples in the experience pool is greater than the set value N, sample M tuples from ; go to Step 3); Otherwise, go to step 1); Step 3) Under the maximum entropy reinforcement learning framework, update the policy network parameter θ, the evaluation network parameter and the temperature parameter α respectively; Step 4) When the AUV meets the termination condition, end; otherwise, go to step 1).

2. The vector propulsion AUV path planning method based on deep reinforcement learning according to claim 1, wherein, Before the said step 1), it further includes: Collect the current state s of the AUV in real time based on sonar and its own inertial navigation information t , action a t , and the state s at the next moment t+1 ; According to the state s t and the action a t , the current reward r is calculated by the reward function t ; Store the tuple [s t , a t , r t , s t+1 into the experience pool Continuously perform the above steps until the number of tuples in the experience pool is greater than the set value N.

3. The vector propulsion AUV path planning method based on deep reinforcement learning according to claim 1, wherein, The policy network and the evaluation network are two independent neural networks; between each layer of each neural network, batch renormalization and bounded activation functions are used for connection; The action distribution is modeled using a Gaussian distribution, and finally the mean and standard deviation of the Gaussian distribution are output.

4. The vector propulsion AUV path planning method based on deep reinforcement learning according to claim 1, characterized in that The reward value r in step 1) t is as follows: r t =r z +r s +r d Among them, r z represents the terminal reward, giving a large positive feedback when reaching the target point and a small negative feedback when colliding with an obstacle; r s represents the immediate reward, r d represents the reward related to the rudder angle and the vector geometric deflection angle.

5. The vector propulsion AUV path planning method based on deep reinforcement learning according to claim 1, characterized in that The said step 3) includes: Step 3-1) Using as the loss function, update the parameters of the evaluation network ; ​ Step 3-2) Using J π (θ) as the loss function, update the parameters θ of the policy network π θ (a t |s t ); Step 3-3) Update the temperature parameter α with J(α) as the loss function.

6. The method for path planning of a vector propulsion AUV based on deep reinforcement learning according to claim 5, wherein The loss function in step 3-1) is as follows: where, |·| sg represents that its gradient is not required represents the parameterized evaluation network, a t+1 represents the action at the next moment, s t+1 represents the state at the next moment, q t , q t+1 respectively represent that when the inputs of the evaluation network are [s t , a t , [s t+1 , a t+1 the output values, γ represents the discount factor, γ ∈ [0, 1].

7. The method for path planning of a vector propulsion AUV based on deep reinforcement learning according to claim 5, wherein The loss function J π (θ) in step 3-2) is as follows: where D KL (·||·) is used to measure the difference between two policies, π θ (·||s t ) represents the policy distribution in state s t , represents the unnormalized probability distribution of the action-state value function in state s t , represents the normalization constant.

8. The vector propulsion AUV path planning method based on deep reinforcement learning according to claim 5, wherein, The loss function J(α) in the said step 3-3) is: In the formula, represents a constant, represents an expected symbol, and the action a inside the expected symbol t uses the parameterized policy π θ to perform sampling.

9. The vector propulsion AUV path planning method based on deep reinforcement learning according to claim 5, characterized in that, The termination condition of the step 4) includes that the actual number of running curtains is greater than the maximum number of running curtains hyperparameter T me .

Citation Information

Patent Citations

  • Underwater vehicle docking control method and device based on imaging sonar

    CN118244755A

  • Generating and classifying training data for machine learning functions

    WO2019126755A1