Autonomous underwater vehicle path following control method based on value distribution reinforcement learning
By constructing a risk-sensitive policy exploration method based on value distribution reinforcement learning, and optimizing the policy network parameters, the problem of insufficient policy exploration and unstable updates of AUVs in complex underwater environments is solved, and high-precision path following control is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-05
AI Technical Summary
Existing AUV path-following control methods suffer from insufficient strategy exploration capabilities and unstable strategy updates in complex underwater environments, resulting in reduced tracking accuracy and stability, making it difficult to achieve high-precision path following.
We employ a value distribution-based reinforcement learning approach, constructing a risk-sensitive policy exploration method. By introducing a policy network and a value network, and combining maximum entropy reinforcement learning with risk-sensitive state sequences, we optimize policy parameters to enhance the exploration capability and learning stability of the policy network.
It improves the path-following accuracy and learning efficiency of AUVs in complex underwater environments, achieves faster algorithm convergence speed and a more stable training process, and enhances the autonomous path-following capability of AUVs.
Smart Images

Figure CN121979252A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of deep reinforcement learning and intelligent control, and relates to a path following control method for autonomous underwater vehicles based on value distribution reinforcement learning. Background Technology
[0002] With the development of marine exploration, there is an increasing need for unmanned underwater vehicles (UUVs) for marine scientific research, marine resource exploration, and marine engineering construction. Autonomous Underwater Vehicles (AUVs) are now widely used in marine exploration. Underwater path following is crucial for achieving these complex tasks and has become a key research focus in the field of intelligent control in recent years. However, current AUV motion control still faces many challenges, mainly including the following aspects: First, AUVs have highly nonlinear dynamic models and time-varying hydrodynamic coefficients, and are strongly coupled nonlinear multi-input multi-output systems, making it difficult to obtain accurate AUV models. When facing complex path following situations, the tracking accuracy and stability of AUVs decrease significantly, making stable high-precision control a challenge. The quality of automatic control methods for underwater vehicles directly affects whether they can successfully complete their missions and influences their safety performance. Therefore, improving the accuracy and stability of AUV path following technology is of great significance for its application and exploration in the marine field.
[0003] In recent years, scholars both domestically and internationally have conducted extensive research in the field of AUV path tracking control, with the main methods falling into two categories: linear control and intelligent control. Linear control methods rely on precise mathematical or physical models. While they can provide theoretically accurate control, they are difficult to apply to complex, nonlinear systems in practice. Furthermore, they lack self-learning and adaptive capabilities; when underwater robots operate over large areas, the control performance of linear controllers deteriorates significantly, limiting their widespread application. With the advent of the intelligent information age, advancements in artificial intelligence technology and data processing capabilities have driven the rapid development of intelligent control methods. These methods reduce reliance on precise mathematical models by training and learning the agent in a simulation system, resulting in controllers with superior adaptability and transferability. In particular, the application of deep learning technology, by constructing deep neural networks as nonlinear function approximators, enables the learning of the mapping relationship between complex environmental states and values. For vast and complex environmental state spaces, deep neural networks can automatically extract state features. This technology significantly improves the decision-making ability of the agent, reduces learning and training time, and promotes the development of AUV path control towards intelligence and efficiency. However, in practical engineering applications, the large and complex state space in the underwater environment can lead to insufficient policy exploration capabilities and unstable policy updates, thereby reducing the tracking accuracy of path-following control for autonomous underwater vehicles based on deep reinforcement learning. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention proposes a path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning. This method constructs risk-sensitive state sequences and introduces a policy exploration approach, enabling the agent to further adapt to complex underwater environments, thereby improving the stability of the model's global learning and training.
[0005] The technical solution adopted by the present invention to achieve the above objectives is as follows: A path-following control method for an autonomous underwater vehicle based on value distribution reinforcement learning includes the following steps: S1. Define the path following problem of an autonomous underwater vehicle (AUV). The definition of the AUV path following control problem includes three parts: determining the AUV system input, determining the AUV system output, and defining the path following control error. S2. Establish a Markov decision model for the AUV path following problem. Model the Markov decision process for the AUV path following problem in step S1. The decision model for the AUV path following problem consists of a state space... Action space Reward function and the next state composition; S3. Construct a policy network and a value network based on value distribution to solve the Markov decision process in step S2. This includes two parts: constructing a value distribution soft policy iterative framework and constructing a risk-sensitive policy function.
[0006] In maximum entropy reinforcement learning, a reward distribution function is embedded to learn the continuous distribution of state-action reward (also known as the value distribution), which is derived from the value distribution function. and random policy function A value distribution soft policy iterative framework was constructed, embedding a reward distribution function into maximum entropy reinforcement learning to learn a continuous distribution of state-action reward, where... and These are the weight parameters of its neural network; The value network, which includes the value distribution function, is implemented using a fully connected deep neural network. The input to the value network is... The distribution of Q-values output by the value distribution network is the mean of the returns. and standard deviation The value distribution soft policy iterative framework evaluates policies by characterizing the distribution of random cumulative returns. Compared to expected future returns, the complete Gaussian distribution of returns contains more information and can provide a faster update speed for policy learning. The aforementioned risk-sensitive policy function, for risk-sensitive scenarios, constructs a risk-sensitive state sequence based on the standard deviation of the value distribution and the average reward: in The preset standard deviation threshold of the value distribution. The average reward value for all trajectories; Let these represent the risk-sensitive state space and action space, respectively. Simultaneously, a risk-sensitive policy function is obtained by combining a stochastic policy function with uniform sampling of the risk-sensitive action space. This will enable the improved strategy exploration method.
[0007] Construct the corresponding policy optimization objective function Optimize policy parameters by maximizing the value of the policy objective function. The policy network, which includes a risk-sensitive policy function, is implemented using a fully connected deep neural network, consisting of an input layer, two hidden layers, and an output layer; the input to the policy network is a state vector. The output of the policy network is an action vector. .
[0008] S4. Solve the Markov decision model through the policy network and value network, and train the policy network and value network to obtain the optimal path following strategy for the autonomous underwater vehicle.
[0009] The beneficial effects of this invention are as follows: The AUV path following control method proposed in this invention does not rely on a model. It obtains a better target strategy by sampling data during the AUV's driving process. This process does not require any assumptions about the AUV model. Furthermore, it makes full use of the standard deviation of the value distribution and trajectory rewards to enable the AUV to better understand environmental information and improve training efficiency.
[0010] This invention is based on a value distribution soft policy iteration framework that embeds a value distribution function. By introducing a policy balancing method, it encourages more stable exploration in the later stages of training, thereby reducing the imbalance between learning new knowledge and experience, accelerating the algorithm's convergence speed, and making the global learning and training process more stable, resulting in higher AUV path following accuracy. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the structural process of the present invention; Figure 2 This is a schematic diagram of the algorithm flow of the present invention; Figure 3 This is a comparison chart of the learning reward curves of the method proposed in this invention and existing reinforcement learning methods; Figure 4 This is a comparison chart of the AUV path following performance between the method proposed in this invention and existing reinforcement learning methods. Detailed Implementation
[0012] To better understand the above-described objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Many specific details are set forth in the following description to provide a thorough understanding of the present invention; however, the present invention may also be implemented in other ways different from those described herein, and therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0013] Example 1: like Figure 1 and Figure 2 As shown, this invention provides an AUV path-following control method based on value distribution reinforcement learning. First, a path-following controller based on reinforcement learning is employed. By taking the current position, velocity, direction, and state value of the AUV as inputs, the velocity error between the desired velocity and the current velocity in the X and Y axes is calculated. After processing using the line-of-sight (LOS) navigation method, the lateral tracking error between the AUV's current position and the projection point is calculated. Calculate the yaw angle error between the current yaw angle and the target yaw angle of the AUV. The final outputs are the AUV's propulsion speed error, lateral tracking error, and yaw angle error. The reinforcement learning controller uses the path tracking error to control the AUV's thrusters and rudders, adjusting the AUV's motion to keep it moving along the desired path.
[0014] A path-following control method for an autonomous underwater vehicle based on value distribution reinforcement learning includes the following steps: S1. Define the path following problem of an autonomous underwater vehicle (AUV). The definition of the AUV path following control problem includes three parts: determining the AUV system input, determining the AUV system output, and defining the path following control error.
[0015] S2. Establish a Markov decision model for the AUV path following problem. Model the Markov decision process for the AUV path following problem in step S1. The decision model for the AUV path following problem consists of a state space... Action space Reward function and the next state composition; S3. Construct a policy network and a value network based on value distribution to solve the Markov decision process in step S2. This includes two parts: constructing a value distribution soft policy iterative framework and constructing a risk-sensitive policy function.
[0016] S4. In maximum entropy reinforcement learning, embed the reward distribution function to learn the continuous distribution of state-action reward (also known as the value distribution), which is derived from the value distribution function. and random policy function A value distribution soft policy iterative framework was constructed, embedding a reward distribution function into maximum entropy reinforcement learning to learn a continuous distribution of state-action reward, where... and It is the weight parameter of its neural network.
[0017] The value network, which includes the value distribution function, is implemented using a fully connected deep neural network. The input to the value network is... The distribution of Q-values output by the value distribution network is the mean of the returns. and standard deviation The value distribution soft policy iteration framework evaluates policies by characterizing the distribution of random cumulative returns. Compared to expected future returns, the complete Gaussian distribution of returns contains more information and can provide faster update speeds for policy learning.
[0018] S5. The risk-sensitive policy function is constructed by building a risk-sensitive state sequence based on the standard deviation of the value distribution and the average reward for risk-sensitive scenarios: in The preset standard deviation threshold of the value distribution. The average reward value for all trajectories; Let these represent the risk-sensitive state space and action space, respectively. Simultaneously, a risk-sensitive policy function is obtained by combining a stochastic policy function with uniform sampling of the risk-sensitive action space. This will enable the improved strategy exploration method.
[0019] Risk-sensitive state sequences correspond to states and actions with lower standard deviations in value distribution and higher rewards. Policies derived from these sequences tend to explore less biased and higher-value actions. This means that the overall action space can be better explored while avoiding negative rewards, and the policy update direction can converge more stably toward the optimal policy, thereby improving the exploration capability of the policy network.
[0020] Construct the corresponding policy optimization objective function Optimize policy parameters by maximizing the value of the policy objective function. The policy network, which includes a risk-sensitive policy function, is implemented using a fully connected deep neural network, consisting of an input layer, two hidden layers, and an output layer; the input to the policy network is a state vector. The output of the policy network is an action vector. .
[0021] S6. By learning and training the model constructed in the above steps on the reference trajectory, solve the Markov decision model constructed in the above steps to obtain the optimal path following strategy for the autonomous underwater vehicle.
[0022] Further, in step S1, the AUV system input is determined, and the AUV system input vector is set to... ,in , These represent propeller thrust and rudder angle, respectively, with the subscript t indicating the t-th time step; , The range of values for are respectively and ,in and These represent the maximum propeller thrust and the maximum rudder angle, respectively.
[0023] Further, in step S1, the AUV system output is determined, and the AUV system output vector is... ,in Let X and Y be the coordinates of the AUV along the X and Y axes in the inertial coordinate system at time step t. Let be the angle between the forward direction of the AUV at time step t and the Y-axis of the fixed coordinate system.
[0024] Furthermore, in step S1, a path following control error is defined, and a trajectory reference point at the t-th time step is selected based on the target path of the AUV. yaw angle is If the only objective is to approximate the desired distance from the actual flight path, then the optimal maneuver for an AUV is to sail directly from the starting point to the destination at maximum cruising speed.
[0025] Therefore, in order to meet the requirements of trajectory tracking accuracy, we need to keep the AUV's yaw angle and the desired yaw angle (the angle between the reference trajectory tangent) as close as possible, and minimize the shortest distance between the current AUV's horizontal position and the reference trajectory; based on the trajectory reference point and the yaw angle, we introduce a lateral trajectory error. and yaw angle error (heading error).
[0026] To ensure that the AUV maintains a desired cruise speed, we introduce a speed error. .
[0027] in, This is the current speed. It is the expected speed. That is the maximum speed.
[0028] Furthermore, in step S2, a state space is defined. In the two-dimensional underwater operating environment, the AUV's position information includes its lateral position x, longitudinal position y, and yaw angle ψ. Therefore, the AUV's operating environment needs to consider surge velocity u, roll velocity v, and yaw rate r to accurately describe its operating state. To ensure the AUV completes the path-following task, a lateral trajectory error is introduced. and yaw angle error Based on the path tracing MDP framework, the state space of the AUV path following problem is defined as follows: Further, in step S2, an action space is defined, and the action vector at time step t is defined as the AUV system input vector at that time step, i.e.: Furthermore, in step S2, a reward function is defined, and the reward function at time step t is used to characterize the state. Take action The execution effect is determined by multiplying the defined trajectory tracking control error by a negative coefficient. , and By reducing these two errors, the state of the controlled object is adjusted to continuously bring it closer to the target trajectory. To ensure that the AUV's running direction points from the starting position to the target position, we map the AUV's current position to the arc length of the reference trajectory. By using a time step Mapped arc length within and the distance traveled at the desired cruising speed Perform division, then give Multiply by a positive coefficient Received a reward.
[0029] Finally, the AUV reward function at time step t is obtained: Furthermore, in step 3, corresponding parameter settings are made, including setting the maximum number of iteration rounds M and the maximum time step per iteration. Size of the training set extracted from experience replay Learning rate of the value distribution network Learning rate of the policy network Policy entropy coefficient updates the network learning rate Initialize target network update parameters and discount factor .
[0030] Furthermore, steps S4 and S5 require initializing the diversity value distribution value network and policy network. The weight parameters of the value distribution network and policy network are randomly initialized. and Initialize the policy entropy coefficient Initialize the preset value distribution standard deviation threshold. ; Constructing an experience set Let this set of experiences be... The maximum capacity is It is stored in the experience cache pool and initialized to empty.
[0031] The iteration of the policy network and value network begins, training is performed on the value network and policy network, and the number of training iterations is initialized. =1; Sets the current time step. Randomly initialize the state variables of the AUV Let the state variable at the current time step be... Initialize the initial time step size. Determine the action at the current time step. .
[0032] AUV in its current state Next action According to the reward function Calculate the current reward value And a new state was observed. , denoted as a trajectory Let it be an empirical sample. If the empirical set... The number of samples has reached the maximum capacity. If so, first delete the first sample added, then add the new empirical sample. Store in experience collection In the middle; otherwise, directly use the empirical sample. Store in experience collection middle.
[0033] Select N empirical samples from the empirical set R, specifically as follows: when the empirical set R... When the number of samples does not exceed N, all empirical samples in the empirical set R are selected; when the empirical set R is less than N, all empirical samples in the empirical set R are selected. When the number of samples exceeds N, N samples are drawn from the experience set according to priority sampling.
[0034] Through iterative learning, the standard deviation of each trajectory and the average reward of all trajectories are calculated to obtain the risk-sensitive action space. Construct risk-sensitive policy functions: in, This represents an annealing hyperparameter that increases with the number of training iterations. This represents the time step of the current iteration round, and mod represents the modulo operation, which is used in the early stages of training. When the value is 0, the original policy function outputs the policy; in the later stages of training, actions are obtained from the risk-sensitive action space through uniform sampling. This indicates uniform sampling. Indicates the state The risk-sensitive action space is defined below; the selected N experience samples are used to calculate the output action through the policy network in step S5. .
[0035] Furthermore, in step S4, the state-action value function of the value network... Output the current policy Random cumulative returns generated In step S5, the policy network output is a policy. The policy network optimizes the objective function (policy network loss function) through policies. Parameter updates are performed, and the value network updates the objective function (value network loss function) through value distribution. Update the parameters. Define random cumulative reward. : Its probability density function is defined as follows Also known as the value distribution function, the corresponding state-action value function is: Let the current strategy be State-action value function of value network The goal of policy update is to obtain a new policy that maximizes the state-action value function and the policy entropy. , is defined as: Finally, the new strategy The actions performed in the process complete the path-following control of the autonomous underwater vehicle.
[0036] Furthermore, in steps S4, S5, and S6, the weight parameters of the value network and the policy network are updated by calculating the gradient of the loss function with respect to the weight parameters. and Corresponding target network weight parameters and Updates are performed using an extremely low synchronization rate, and the policy entropy coefficient is... The update is performed through a dynamic adjustment mechanism, as shown in the following formula: in The target entropy is the minimum expected entropy. To update parameters; Furthermore, let And on Make a judgment, such as If the policy network outputs the next action after selecting N experience samples, it will be used as the input AUV to continue following the reference trajectory; otherwise, the number of training iterations will be used to determine whether the constraints are met. If the number of training iterations is less than M, the AUV proceeds to the next iteration; otherwise, the iteration ends, the training process of the value network and policy network is terminated, the parameter values of the value network and policy network at the time of iteration termination are saved, and the policy output by the policy network implements path following control for the AUV.
[0037] To facilitate understanding of the above technical solutions of the present invention, the following detailed description of the above technical solutions of the present invention will be provided through AUV path following examples.
[0038] Example 2: A path-following control method for AUVs based on value distribution reinforcement learning includes the following steps: S1. Simulation experiments were used to verify the performance of the AUV underwater horizontal path following. The AUV underwater environment and AUV motion model were developed using Python and the Gym library provided by OpenAI, while the value network and policy network were constructed using the machine language library (PyTorch).
[0039] Example 1 is experimentally verified on path RT1, and comparative experiments are conducted under different algorithm controls. The final reward value and AUV path tracking effect are summarized. The method in this paper is referred to as MS-DSAC.
[0040] The reference trajectory RT1 is shown below: in and This represents the horizontal and vertical coordinates of the trajectory in the coordinate system. This represents a time step, and also represents a curved trajectory.
[0041] The maximum number of iterations in the simulation task is set to 300, and the initial position of the AUV is randomly located in each round. The constraints are: in The reference trajectory starting point is; the initial yaw angle is... ,in The yaw angle at the starting point; the initial velocity of the AUV is... Specifically initialized as The input to the AUV system is the initial state of the AUV mentioned above, and the corresponding initial path-following error is calculated accordingly. All comparative experiments are based on the widely used REMUS autonomous unmanned aerial vehicle, with its maximum propeller thrust... and rudder angle They are respectively and This is the output constraint of the AUV system.
[0042] S2. The time step of reinforcement learning is set to 0.1s, the expected tracking time of the reference trajectory is 100s, and each round is divided into 1000 time steps. The AUV system output is randomly initialized according to the constraints to obtain the new state after the AUV moves.
[0043] S3. Set the parameters of the MS-DSAC algorithm as follows: batch sampling size is 64, experience replay pool size is 10000, policy network learning rate is 0.0003, evaluation network learning rate is 0.003, policy entropy target value is -2, hidden layers are 2, hidden units are 128, discount factor is 0.99, soft update parameter is 0.005, Gaussian noise standard deviation is 0.01. The policy entropy target value is set to -2 (negative value of action space dimension), hidden layers are 2, hidden units are 128, discount factor is 0.99, soft update parameter is 0.005, and the optimization method for all algorithms is Adam.
[0044] The basic parameter settings for S4, DDPG, TD3, SAC, and DSAC algorithms are the same as above. The TD3 algorithm has the following specific parameter settings: delay update frequency is 5, exploration noise variance is 0.1, and policy noise variance is 0.2.
[0045] S5. Taking the reference trajectory RT1 as an example, in the same underwater simulation environment, each algorithm is applied to the reference trajectory RT1 for comparison.
[0046] Figure 3 The graph shows the average total reward learning curves for MS-DSAC and other current mainstream algorithms under the RT1 curve environment, based on 20 sets of trials. The results show that MS-DSAC and other reinforcement learning algorithms can converge to satisfactory values in the current environment. Furthermore, the fluctuations in the reward function during the convergence phase indicate that the MS-DSAC algorithm has a faster convergence speed, higher reward value, and a more stable learning process than existing mainstream algorithms. These results demonstrate that this method helps improve the algorithm's policy learning ability in the vast state space of the underwater environment, and validates the effectiveness of this method in AUV path following.
[0047] Figure 4 The simulation results show the AUV tracking trajectories obtained by MS-DSAC and existing main algorithms under the RT1 reference trajectory. The simulation results demonstrate that, under the same environment, the MS-DSAC algorithm exhibits superior tracking control accuracy and smaller tracking error compared to existing main algorithms on trajectories RT1 and RT2. The following trajectory obtained by this method almost coincides with the reference trajectory, indicating that the MS-DSAC method can achieve high-precision AUV following control. The results show that the MS-DSAC algorithm can effectively complete path tracking tasks in underwater environments, demonstrating that this method enables AUVs to learn effectively in complex underwater environments.
[0048] Table 1 shows the average path following error and path following standard deviation of each algorithm. The data in the table can be seen that the MS-DSAC algorithm of this invention has a lower average tracking error under the reference path than the other algorithms, and has better tracking accuracy; in particular, it has a lower standard deviation and more stable learning compared with other algorithms.
[0049] Table 1 Path following error statistics for each algorithm on RT1
[0050] Table 2 summarizes the key performance indicators (KPIs) of the MS-DSAC, DSAC, SAC, and TD3 algorithms. R represents 20 independent learning and training processes, and each process records the rewards for all rounds. , as well as These are the optimal value, mean, and standard deviation of R in the interval [100, 300] rounds, respectively. The improvement rate is defined as: The improvement of the MS-DSAC algorithm over existing major algorithms is measured by comparing the average reward value of the MS-DSAC algorithm with the average reward value of the existing major algorithms. The larger the value, the greater the improvement of the MS-DSAC algorithm compared to existing algorithms.
[0051] Table 2. Reward statistics for each algorithm on RT1
[0052] Analyzing Table 2, we can draw the following conclusions: Compared to other algorithms, MS-DSAC has the highest average reward, proving its superior performance in control accuracy. This is because the policy exploration method introduced in MS-DSAC's policy network guides the network to select actions that are more suitable for the current AUV environment, thereby improving the accuracy of the policy network's output.
[0053] Compared to other algorithms, MS-DSAC has the smallest standard deviation of rewards, proving that it performs well in terms of stability during training. This is because the MS-DSAC algorithm improves the stability of policy updates in the policy network by uniformly sampling actions from the risk-sensitive action space.
Claims
1. A path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning, characterized in that, It includes the following steps: S1. Define the path following problem of the autonomous underwater vehicle (AUV), including determining the AUV system input, determining the AUV system output, and defining the path following control error. S2. Establish a Markov decision model for the AUV path following problem and model the Markov decision process for the AUV path following problem. S3. Construct a policy network and a value network based on value distribution; S4. Solve the Markov decision model through the policy network and value network, and train the policy network and value network to obtain the optimal path following strategy for the autonomous underwater vehicle.
2. The path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning according to claim 1, characterized in that, In step S1, determining the AUV system input specifically involves: letting the AUV system input vector be... ,in , These represent propeller thrust and rudder angle, respectively, with the subscript t indicating the t-th time step; , The range of values for are respectively and ,in and These are the maximum propeller thrust and the maximum rudder angle, respectively. The specific steps for determining the AUV system output are: Let the AUV system output vector be... ,in Let X and Y be the coordinates of the AUV along the X and Y axes in the inertial coordinate system at time step t. Let be the angle between the forward direction of the AUV at time step t and the Y-axis of the fixed coordinate system; The defined path following control error is specifically defined as: selecting the trajectory reference point at the t-th time step based on the target path of the AUV. yaw angle is Based on the trajectory reference point and yaw angle, define the lateral trajectory error. Yaw angle error and the introduction of speed error .
3. The path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning according to claim 2, characterized in that, The Markov decision model for the AUV path-following problem is based on the state space. Action space Reward function and the next state composition; Define the state space. In a two-dimensional underwater operating environment, the position information of an AUV includes its lateral position (x), longitudinal position (y), and yaw angle (ψ). The AUV's operating environment needs to consider surge velocity (u), roll velocity (v), and yaw rate (r), introducing lateral trajectory error. and yaw angle error Based on the path tracing MDP framework, the state space of the AUV path following problem is defined as follows: ; Define the action space, and define the action vector at time step t as the AUV system input vector at that time step, i.e. ; Define the reward function; multiply it by negative coefficients based on the defined trajectory tracking control error. , and The arc length that maps the current position of the AUV to the reference trajectory. By using a time step Mapped arc length within and the distance traveled at the desired cruising speed Divide, then give Multiply by a positive coefficient The reward is obtained, and finally the AUV reward function at time step t is obtained: 。 4. The path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning according to claim 3, characterized in that, The construction of the value distribution-based policy network and value network includes building a value distribution-based soft policy iterative framework and constructing a risk-sensitive policy function; and setting parameters, such as the maximum number of iteration rounds. Maximum time steps per iteration Size of the training set extracted from experience replay Learning rate of the value distribution network Learning rate of the policy network Policy entropy coefficient updates the network learning rate Initialize target network update parameters and discount factor .
5. The path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning according to claim 4, characterized in that, The constructed value distribution soft strategy iterative framework is based on the value distribution function. and random policy function A value distribution soft policy iterative framework is constructed, embedding a reward distribution function into maximum entropy reinforcement learning to learn a continuous distribution of state-action reward, where... and These are the weight parameters of the neural network; the value network, which includes the value distribution function, is implemented using a fully connected deep neural network, and the input to the value network is... The distribution of Q-values output by the value distribution network is the mean and standard deviation of the returns; the value distribution soft policy iterative framework evaluates the policy by characterizing the distribution of random cumulative returns. The aforementioned risk-sensitive policy function, for risk-sensitive scenarios, constructs a risk-sensitive state sequence based on the standard deviation of the value distribution and the average reward: in The preset standard deviation threshold of the value distribution. This represents the average reward value for all current trajectories. Let the risk-sensitive state space and action space be represented respectively. Furthermore, a risk-sensitive policy function is obtained by combining a stochastic policy function with uniform sampling of the risk-sensitive action space. Obtain the action; construct the corresponding policy optimization objective function. Optimize policy parameters by maximizing the policy objective function value. The policy network, which includes a risk-sensitive policy function, is implemented using a fully connected deep neural network. The input to the policy network is a state vector. The output of the policy network is an action vector. .
6. The path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning according to claim 5, characterized in that, The implementation process of the policy network and value network is as follows: Weight parameters of the random initialization value distribution network and the policy network and Initialize the policy entropy coefficient ; Initialize the preset value distribution standard deviation threshold ; Construct an experience queue set Let this set of experiences be... The maximum capacity is Stored in the experience cache pool and initialized to empty; The policy network and value network iterations begin; training is performed on the value network and policy network, and the number of iterations is initialized; the current time step is set, and the state variables of the AUV are randomly initialized. Let the state variable at the current time step be... Initialize the initial time step size. Determine the action at the current time step. ; AUV in its current state Next action According to the reward function Calculate the current reward value And a new state was observed. , denoted as a trajectory For an empirical sample; if the empirical set The number of samples has reached the maximum capacity. If so, first delete the first sample added, then add the new empirical sample. Store in experience collection In the middle; otherwise, directly use the empirical sample. Store in experience collection middle; From experience set Select N empirical samples from the set of experiences: When the number of samples does not exceed N, then the empirical set is selected. All empirical samples in the empirical set; when the empirical set When the number of samples exceeds N, N samples are drawn from the experience set according to priority sampling. Through iterative learning, the standard deviation of each trajectory and the average reward of all trajectories are calculated to obtain the risk-sensitive action space. Construct risk-sensitive policy functions: in, This represents an annealing hyperparameter that increases with the number of training iterations. This represents the time step of the current iteration round, and mod represents the modulo operation, which is used in the early stages of training. When the value is 0, the original policy function outputs the policy; in the later stages of training, actions are obtained from the risk-sensitive action space through uniform sampling. This indicates uniform sampling. Indicates the state In the risk-sensitive action space, the selected N experience samples are used to calculate the output action through a policy network. .
7. The path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning according to claim 6, characterized in that, The implementation process of the value distribution soft policy iteration framework on the policy network and value network is as follows: State-action value function of the value network Output the current policy Random cumulative returns generated The policy network output is the policy. ; Policy networks optimize objective functions through policies. The value network updates the objective function by updating the parameters and the value distribution. Update parameters; define random cumulative reward. : Define the probability density function of random cumulative return as follows: Also known as the value distribution function, the corresponding state-action value function is: ; Let the current strategy be State-action value function of value network The goal of policy update is to obtain a new policy that maximizes the state-action value function and the policy entropy. , is defined as: in, The strategy entropy coefficient; Finally, the new strategy The actions performed in the process complete the path-following control of the autonomous underwater vehicle.
8. The path-following control method for autonomous underwater vehicles based on value distribution reinforcement learning according to claim 7, characterized in that, In step S4, the weight parameters of the value network and the policy network are updated by calculating the gradient of the loss function with respect to the weight parameters. and Corresponding target network weight parameters and Updates are performed using the synchronization rate, and the policy entropy coefficient is... The update is performed through a dynamic adjustment mechanism, as shown in the following formula: Among them The target entropy is the minimum expected entropy; To update parameters; make And on Make a judgment, such as If the policy network outputs the next action after selecting N experience samples, it will be used as the input AUV to continue following the reference trajectory; otherwise, the number of training iterations will be used to determine whether the constraints are met. If the number of training iterations is less than M, AUV proceeds to the next iteration. Otherwise, the iteration ends, terminating the training process of the value network and policy network. The parameter values of the value network and policy network at the time of iteration termination are saved, and the policy output by the policy network implements path following control for the AUV.