A path tracking control method for unmanned vessels based on deep reinforcement learning

By constructing a time-varying marine environment model and optimizing the control strategy using the improved TD3 algorithm, the problem of insufficient environmental realism and adaptability in unmanned surface vessel (USV) path tracking control was solved, enabling USVs to achieve high-precision path tracking and rapid response in complex marine environments.

CN119882740BActive Publication Date: 2025-10-31TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510044485.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-11
Publication Date
2025-10-31
Estimated Expiration
2045-01-11

AI Technical Summary

Technical Problem

Existing unmanned vessel path tracking control methods suffer from insufficient environmental realism and control strategy effectiveness in complex marine environments. Traditional methods also have high computational complexity and poor adaptability to dynamic environments.

Method used

A time-varying marine environment model integrating multiple environmental factors is constructed using a deep reinforcement learning approach. The model is modeled as a Markov decision process, and the improved TD3 algorithm is used to optimize the control strategy. The path tracking control of the unmanned vessel is trained using real marine data.

Benefits of technology

It enables high-precision path tracking of unmanned vessels in complex marine environments, improves response speed and system stability, and enhances the realism, reliability, adaptability, and real-time performance of the environmental model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119882740B_ABST
    Figure CN119882740B_ABST
Patent Text Reader

Abstract

This invention relates to a path tracking control method for unmanned surface vessels (USVs) based on deep reinforcement learning, comprising the following steps: Step 1, constructing a time-varying marine environment model integrating multiple environmental factors; Step 2, modeling the USV path tracking control problem into a solution framework based on a Markov decision process based on the marine environment model of Step 1; Step 3, optimizing the USV path tracking control using an improved Twin Delayed Deep Deterministic policy gradient (TD3) algorithm based on the solution framework of Step 2; Step 4, using the policy neural network trained in Step 3 to control the USV to perform the path tracking task. The deep reinforcement learning-based USV path tracking control method proposed in this invention employs the improved TD3 algorithm for policy training and optimization, significantly improving the convergence speed of the control policy and achieving accurate and reliable path tracking performance in dynamic and complex marine environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous motion control of unmanned vessels, and particularly relates to a path tracking control method for unmanned vessels based on deep reinforcement learning. Background Technology

[0002] With the continuous advancement of the national maritime power development plan, marine resources have been increasingly developed and utilized, leading to a greater demand for intelligent and autonomous marine equipment. Among these, unmanned surface vessels (USVs), due to their high maneuverability, long range, and autonomy, are widely used in resource exploration, maritime search and rescue, and intelligence reconnaissance. In the research of motion control for USVs, high-precision path tracking technology in complex marine environments remains an important and challenging problem. The main function of path tracking control is to generate motion control commands, enabling the USV to track a reference path with minimal dynamic error.

[0003] The marine environment for unmanned surface vessel (USV) path tracking is dynamic and complex, influenced by environmental factors such as ocean currents and wind. In actual navigation, these environmental factors can cause USVs to deviate from their reference paths, leading to significant tracking errors. However, early research neglected the impact of environmental factors on USV motion, resulting in low practical application value for their marine environment models. Simulating the marine environment using simple mathematical function models or uniform vector fields lacks environmental realism. Furthermore, current marine environment models ignore the dynamic changes of environmental factors. With the development of marine meteorological data acquisition and forecasting technologies, real ocean observation and forecast data are widely used in other fields. Therefore, establishing a realistic and accurate dynamic marine environment model based on real ocean data is crucial for USV path tracking control.

[0004] Over the past few decades, various research efforts have focused on solving the path tracking problem for unmanned surface vessels (USVs). Traditional control methods, such as sliding mode control (SMC), model predictive control (MPC), and linear active disturbance suppression controllers (LADRC), have been widely applied to USV path tracking control. However, these model-based controllers are inherently based on mathematical estimation and analysis, leading to high computational complexity and poor portability. To overcome these problems of traditional controllers, intelligent path tracking controllers have been proposed, including fuzzy logic control and neural network-based control. However, these intelligent controllers all face challenges such as high requirements for expert experience, poor adaptability to dynamic environments, and slow inference speed. Furthermore, path tracking control of USVs in complex marine environments places even higher demands on the model adaptability and real-time model inference capabilities of the methods.

[0005] In recent years, with the development of deep learning and reinforcement learning technologies, deep reinforcement learning has emerged as a possibility to meet the aforementioned requirements. Deep reinforcement learning methods can control extremely complex systems without requiring precise system dynamics models and prior knowledge, and are widely used in fields such as intelligent robot control and recommender systems. For unmanned surface vessel (USV) path tracking control, the USV agent learns the optimal control strategy in dynamic and complex marine environments through trial-and-error interaction with the environment. Specifically, deep reinforcement learning methods control the USV's action output, allowing the USV to continuously interact with the simulation environment in real time, learning and training a control strategy that minimizes path tracking error. Classical deep reinforcement learning methods such as DQN and DDPG suffer from the problem of overestimating the Q-value, leading to model instability and poor control strategy performance.

[0006] In summary, significant progress has been made in the theoretical research of control methods for unmanned surface vessels (USVs) path tracking. However, many researchers have not carefully considered the impact of environmental factors on the motion of USVs, resulting in less than ideal results when applied in practice. Furthermore, traditional control methods lack robustness and adaptability.

[0007] A search revealed no prior art that is identical or similar to this invention. Summary of the Invention

[0008] To overcome the shortcomings in environmental realism and control strategy effectiveness in unmanned surface vessel (USV) path tracking control, this invention proposes a deep reinforcement learning-based USV path tracking control method. By introducing real ocean data, a time-varying ocean environment model integrating multiple environmental factors is constructed. Then, the USV path tracking control problem is modeled as a solution framework based on Markov decision processes, designing a reasonable state-action space and reward function. An improved TD3 algorithm is used to optimize the USV path tracking control solution, and the trained policy neural network controls the USV to perform the path tracking task. This invention enables rapid and stable convergence of the control strategy, allowing the USV to accurately track a reference path in complex ocean environments.

[0009] The present invention solves its practical problem by adopting the following technical solution:

[0010] A path tracking control method for unmanned surface vessels based on deep reinforcement learning includes the following steps:

[0011] Step 1: Construct a time-varying marine environment model that integrates multiple environmental factors;

[0012] Step 2: Based on the marine environment model in Step 1, the unmanned vessel path tracking control problem is modeled as a solution framework based on Markov decision process;

[0013] Step 3: Based on the solution framework of Step 2, the improved Twin Delayed Deep Deterministicpolicy gradient (TD3) algorithm is used to optimize the path tracking control of the unmanned vessel.

[0014] Step 4: Use the policy neural network trained in Step 3 to control the unmanned vessel to perform path tracking tasks.

[0015] Furthermore, the specific steps for constructing the marine environment model in step 1 include:

[0016] 1.1 Ocean current and wind field environmental data for a portion of the South China Sea (113.5°E-115.5°E, 17.25°N to 19.25°N) were obtained from the National Marine Science Data Center;

[0017] 1.2 The region was divided into a 120×120 grid using a grid method, and ocean current and wind field environmental data were interpolated into the grid environment using inverse distance weighted interpolation.

[0018] 1.3 When the USV is at position P, the environmental data at a certain time t is represented as follows:

[0019]

[0020] In the formula I P (t) represents the environmental information received in real time by the USV at position P and time t via sensors such as a flow profiler and anemometer. and These represent the lateral and longitudinal velocities of the ocean current, respectively. and These represent the lateral and longitudinal velocities of the wind, respectively.

[0021] The environmental data I(t) at a non-grid point location P can be obtained from the four surrounding adjacent grid points P. i Environmental data I at (i = 1, ..., 4) i (t) is obtained by distance weighting and is calculated by the following formula:

[0022]

[0023] Furthermore, step 2 specifically includes:

[0024] The unmanned surface vessel path tracking control problem is modeled as a solution framework based on Markov decision processes, specifically including the state space S, action space A, reward function R, state transition relation P, and discount factor γ of the Markov decision process:

[0025] Here, the state space S represents the set of state information s perceived by the unmanned surface vessel, including motion parameter information sm Environmental information e LOS guidance information LOS The state s during the t-th control time. t Represented as:

[0026] s t ={s m ,s e ,s LOS}

[0027] Where: motion parameter information s m = {(x,y),U,ψ}, where (x,y) is the current position of the unmanned surface vessel (USV), and U and ψ are the current speed and heading angle of the USV, respectively; environmental data as described in step 1, environmental information s e =I P (t); LOS guidance information s LOS =(ψ d ,x e ,y e ), where ψ d For the ideal heading angle of the unmanned surface vessel, x e and y e These are the longitudinal tracking error and the lateral tracking error, respectively. The LOS guidance method is used to calculate the desired heading angle and tracking error to improve tracking accuracy. The desired heading angle can be expressed as:

[0028]

[0029] In the formula, ψ k The slope of the tracking path is Δ, where Δ is the forward-looking distance of the unmanned vessel and β is the sideslip angle.

[0030] For the unmanned surface vessel at position (x, y), the parameterized path P is tracked. k The tracking error can be expressed as: (t),

[0031]

[0032] Wherein, the action space A represents the set of actions a of the unmanned vessel, and the action a at the t-th control time is... t Defined as the output speed and heading angle of the unmanned surface vessel (USV), the heading angle is controlled by controlling the direction of the rudder, while the speed of the USV is adjusted by controlling the thrust of the propeller, as shown in the following formula:

[0033]

[0034] In the formula, U(t) represents the velocity of the unmanned vessel. Indicates the heading angle of the unmanned vessel;

[0035] The reward function R defines the unmanned surface vessel in different states s.t and action a t The reward signal r obtained below t It is used to guide the learning and decision-making process of control strategies, and is defined as follows:

[0036]

[0037] In the formula, K1 and K2 are proportional parameters assigned to different error penalty terms, and t step The total control time is [amount missing]. Furthermore, the unmanned surface vessel (USV) receives a substantial reward when it successfully completes the tracking task. However, if the USV fails to track within the specified number of steps or deviates from the tracking area, it will be subject to a severe penalty.

[0038] The state transition relationship is defined as the probability that the unmanned vessel will update from the current environmental state to the next state after making an action decision, as shown in the following formula:

[0039] P(S t+1 |S t ×[A] t )=ρ

[0040] In the formula, ρ is the environmental transition rate, which takes the value of 1 in the unmanned vessel path tracking environment;

[0041] The discount factor is defined as γ, which is a value between 0 and 1, representing the importance of future rewards. The larger the discount factor γ, the more attention is paid to the impact of long-term rewards on the control strategy.

[0042] Furthermore, the specific steps of step 3 include:

[0043] 3.1 Initialize the marine environment, perform network parameter initialization processing on the actor policy neural network and critic value neural network of the improved TD3 algorithm, and set the experience replay buffer.

[0044] 3.2 The unmanned surface vessel (USV) agent interacts with the marine environment, following the same interaction process as in step 2. The USV's state information is input into the actor policy network. The output action of the actor policy network is calculated through forward propagation. After the USV executes the action, it undergoes a state transition and calculates the reward. During this interaction, experience data is collected and stored in the experience replay buffer. This step is repeated until the experience replay buffer is full.

[0045] 3.3 The Critic value network uses the Q-value function to evaluate the performance of the actor policy network and guide its next stage update; the actor policy network determines to take a better action by changing policy π to obtain a higher reward.

[0046] 3.4 Using a two-state value network to address the overestimation of the Q-value function, the next state s t+1 The Q value can be estimated by two target value networks:

[0047]

[0048] In the formula, Q′1 and Q′2 are the target Q values ​​of the two value networks, π′ is the policy of the target policy network, and θ μ′ For the parameters of the target policy network, and These are the parameters for the two target value networks, respectively.

[0049] Choosing the smaller target Q value, the target value of the value network is calculated according to the Bellman equation as follows:

[0050]

[0051] 3.5 Based on the priority experience sampling mechanism, N experience samples are sampled from the experience replay buffer for network updates. The state set S and action set A are input into the value neural network, and the parameters of the neural network are updated through backpropagation. The loss function of the value network is constructed using the following formula to guide policy learning and optimization:

[0052]

[0053] 3.6 Update the parameters of the critic value network by minimizing the loss function and gradient descent:

[0054]

[0055] In the formula It is the gradient of the loss function. α1 is the gradient of the Q value, and α1 is the learning rate of the value network.

[0056] 3.7 A policy delay update mechanism is used to improve the stability of the policy network update. The policy network updates at a lower frequency than the value network. Gradient ascent is used to guide the policy network to update parameters θ in the direction of improving the Q-value. μ :

[0057]

[0058] In the formula, The objective function is J(θ) μ The gradient of ) It is strategy π(s) t |θ μ The gradient of ), where α2 is the learning rate of the policy network;

[0059] 3.8 Using a soft update mechanism to update the parameters of the target network and θ μ′ To perform an update, the current network gradually updates the target network at a smaller update rate, as expressed by the following formula:

[0060]

[0061] θ μ′ ←εθ μ +(1-ε)θ μ′

[0062] In the formula, ε is the soft update rate;

[0063] 3.9 Repeat steps 3.2-3.8 for several training iterations to finally obtain a policy network with excellent control strategy.

[0064] Furthermore, the specific steps of step 4 include:

[0065] 4.1 Initialize the marine environment and load the trained policy neural network parameters;

[0066] 4.2 Input the initial environmental state of the unmanned vessel into the policy network trained in step 3, and obtain the speed and heading control actions of the unmanned vessel through forward propagation inference;

[0067] 4.3 The unmanned surface vessel performs output actions in the marine environment, tracks the reference path, and transitions to a new state;

[0068] 4.4 Input the newly acquired state information into the policy network to generate new control actions for the unmanned vessel;

[0069] 4.5 Repeat steps 4.3 and 4.4 until the unmanned vessel completes the path tracking task of the reference path.

[0070] Advantages and beneficial effects of the present invention:

[0071] 1. This invention addresses the problems of low control accuracy and poor adaptability to dynamic environments in traditional control methods by proposing a path tracking control algorithm for unmanned vessels based on deep reinforcement learning. It models a Markov decision process and trains and optimizes the control strategy based on the improved TD3 algorithm to improve the path tracking control accuracy of unmanned vessels in complex marine environments, thereby improving the response speed and system stability of unmanned vessels.

[0072] 2. This invention addresses the issue of poor realism in environmental models by introducing real marine environmental data to construct a time-varying marine environmental model that integrates multiple environmental factors, thereby improving its realism and reliability. Furthermore, the support of real data compensates for the shortcomings of existing work and enhances its practical application value.

[0073] 3. When modeling the Markov decision process, the reward function proposed in this invention comprehensively considers the tracking error, control cycle and task reward of the unmanned vessel, effectively alleviating the problem of sparse rewards in complex marine environments and helping the unmanned vessel to quickly learn an effective and superior control strategy.

[0074] 4. This invention is easy to deploy on embedded terminal devices. After the trained control strategy is deployed on the terminal device, it can realize rapid inference and meet the needs of real-time decision-making. It has great application value in real ship control. Attached Figure Description

[0075] Figure 1 This is a schematic diagram of the path tracking control task of the present invention;

[0076] Figure 2 This is a schematic diagram of the marine environment of the present invention;

[0077] Figure 3 This is the Markov decision process diagram of the present invention;

[0078] Figure 4 This is a schematic diagram of the deep reinforcement learning-based unmanned vessel path tracking control optimization algorithm based on the TD3 algorithm of the present invention;

[0079] Figure 5 This is a flowchart illustrating the solution process of a path tracking control method for unmanned vessels based on deep reinforcement learning, as described in this invention. Detailed Implementation

[0080] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings:

[0081] A path tracking control method for unmanned surface vessels based on deep reinforcement learning includes the following steps:

[0082] Step 1: Construct a time-varying marine environment model that integrates multiple environmental factors;

[0083] The specific steps of step 1 include:

[0084] 1.3 Ocean current and wind field environmental data for a portion of the South China Sea (113.5°E-115.5°E, 17.25°N to 19.25°N) were obtained from the National Marine Science Data Center;

[0085] 1.4 The region was divided into a 120×120 grid using a grid method, and ocean current and wind field environmental data were interpolated into the grid environment using inverse distance weighted interpolation.

[0086] 1.3 When the USV is at position P, the environmental data at a certain time t is represented as follows:

[0087]

[0088] In the formula I P (t) represents the environmental information received in real time by the USV at position P and time t via sensors such as a flow profiler and anemometer. and These represent the lateral and longitudinal velocities of the ocean current, respectively. and These represent the lateral and longitudinal velocities of the wind, respectively.

[0089] The environmental data I(t) at a non-grid point location P can be obtained from the four surrounding adjacent grid points P. i Environmental data I at (i = 1, ..., 4) i (t) is obtained by distance weighting and is calculated by the following formula:

[0090]

[0091] Step 2: Based on the marine environment model in Step 1, the unmanned vessel path tracking control problem is modeled as a solution framework based on Markov decision process;

[0092] The specific steps of step 2 include:

[0093] The unmanned surface vessel path tracking control problem is modeled as a solution framework based on Markov decision processes, specifically including the state space S, action space A, reward function R, state transition relation P, and discount factor γ of the Markov decision process:

[0094] Here, the state space S represents the set of state information s perceived by the unmanned surface vessel, including motion parameter information s m Environmental information e LOS guidance information LOS The state s during the t-th control time. t Represented as:

[0095] s t ={s m ,s e ,s LOS}

[0096] Where: motion parameter information s m = {(x,y),U,ψ}, where (x,y) is the current position of the unmanned surface vessel (USV), and U and ψ are the current speed and heading angle of the USV, respectively; environmental data as described in step 1, environmental information s e =I P (t); LOS guidance information s LOS =(ψ d ,x e ,y e), where ψ d For the ideal heading angle of the unmanned surface vessel, x e and y e These are the longitudinal tracking error and the lateral tracking error, respectively. The LOS guidance method is used to calculate the desired heading angle and tracking error to improve tracking accuracy. The desired heading angle can be expressed as:

[0097]

[0098] In the formula, ψ k The slope of the tracking path is Δ, where Δ is the forward-looking distance of the unmanned vessel and β is the sideslip angle.

[0099] For the unmanned surface vessel at position (x, y), the parameterized path P is tracked. k The tracking error can be expressed as: (t),

[0100]

[0101] Wherein, the action space A represents the set of actions a of the unmanned vessel, and the action a at the t-th control time is... t Defined as the output speed and heading angle of the unmanned surface vessel (USV), the heading angle is controlled by controlling the direction of the rudder, while the speed of the USV is adjusted by controlling the thrust of the propeller, as shown in the following formula:

[0102]

[0103] In the formula, U(t) represents the velocity of the unmanned vessel. Indicates the heading angle of the unmanned vessel;

[0104] The reward function R defines the unmanned surface vessel in different states s. t and action a t The reward signal r obtained below t It is used to guide the learning and decision-making process of control strategies, and is defined as follows:

[0105]

[0106] In the formula, K1 and K2 are proportional parameters assigned to different error penalty terms, and t step The total control time is [amount missing]. Furthermore, the unmanned surface vessel (USV) receives a substantial reward when it successfully completes the tracking task. However, if the USV fails to track within the specified number of steps or deviates from the tracking area, it will be subject to a severe penalty.

[0107] The state transition relationship is defined as the probability that the unmanned vessel will update from the current environmental state to the next state after making an action decision, as shown in the following formula:

[0108] P(S t+1 |S t ×[A]t )=ρ

[0109] In the formula, ρ is the environmental transition rate, which takes the value of 1 in the unmanned vessel path tracking environment;

[0110] The discount factor is defined as γ, which is a value between 0 and 1, representing the importance of future rewards. The larger the discount factor γ, the more attention is paid to the impact of long-term rewards on the control strategy.

[0111] Step 3: Based on the solution framework of Step 2, the improved TD3 algorithm is used to optimize the solution of the unmanned vessel path tracking control.

[0112] The specific steps of step 3 include:

[0113] 3.1 Initialize the marine environment, perform network parameter initialization processing on the actor policy neural network and critic value neural network of the improved TD3 algorithm, and set the experience replay buffer.

[0114] 3.2 The unmanned surface vessel (USV) agent interacts with the marine environment, following the same interaction process as in step 2. The USV's state information is input into the actor policy network. The output action of the actor policy network is calculated through forward propagation. After the USV executes the action, it undergoes a state transition and calculates the reward. During this interaction, experience data is collected and stored in the experience replay buffer. This step is repeated until the experience replay buffer is full.

[0115] 3.3 The Critic value network uses the Q-value function to evaluate the performance of the actor policy network and guide its next stage update; the actor policy network determines to take a better action by changing policy π to obtain a higher reward.

[0116] 3.4 Using a two-state value network to address the overestimation of the Q-value function, the next state s t+1 The Q value can be estimated by two target value networks:

[0117]

[0118] In the formula, Q′1 and Q′2 are the target Q values ​​of the two value networks, π′ is the policy of the target policy network, and θ μ′ For the parameters of the target policy network, and These are the parameters for the two target value networks, respectively.

[0119] Choosing the smaller target Q value, the target value of the value network is calculated according to the Bellman equation as follows:

[0120]

[0121] 3.5 Based on the priority experience sampling mechanism, N experience samples are sampled from the experience replay buffer for network updates. The state set S and action set A are input into the value neural network, and the parameters of the neural network are updated through backpropagation. The loss function of the value network is constructed using the following formula to guide policy learning and optimization:

[0122]

[0123] 3.6 Update the parameters of the critic value network by minimizing the loss function and gradient descent:

[0124]

[0125] In the formula It is the gradient of the loss function. α1 is the gradient of the Q value, and α1 is the learning rate of the value network.

[0126] 3.7 A policy delay update mechanism is used to improve the stability of the policy network update. The policy network updates at a lower frequency than the value network. Gradient ascent is used to guide the policy network to update parameters θ in the direction of improving the Q-value. μ :

[0127]

[0128] In the formula, The objective function is J(θ) μ The gradient of ) It is strategy π(s) t |θ μ The gradient of ), where α2 is the learning rate of the policy network;

[0129] 3.8 Using a soft update mechanism to update the parameters of the target network and θ μ′ To perform an update, the current network gradually updates the target network at a smaller update rate, as expressed by the following formula:

[0130]

[0131] θ μ′ ←εθ μ +(1-ε)θ μ′

[0132] In the formula, ε is the soft update rate;

[0133] 3.9 Repeat steps 3.2-3.8 for several training iterations to finally obtain a policy network with excellent control strategy.

[0134] Step 4: Use the policy neural network trained in Step 3 to control the unmanned vessel to perform the path tracking task. The specific steps of Step 4 include:

[0135] 4.1 Initialize the marine environment and load the trained policy neural network parameters;

[0136] 4.2 Input the initial environmental state of the unmanned vessel into the policy network trained in step 3, and obtain the speed and heading control actions of the unmanned vessel through forward propagation inference;

[0137] 4.3 The unmanned surface vessel performs output actions in the marine environment, tracks the reference path, and transitions to a new state;

[0138] 4.4 Input the newly acquired state information into the policy network to generate new control actions for the unmanned vessel;

[0139] 4.5 Repeat steps 4.3 and 4.4 until the unmanned vessel completes the path tracking task of the reference path.

[0140] In this embodiment, unmanned surface vessel path tracking is as follows: Figure 1 As shown, the unmanned vessel operates in an ocean environment with ocean currents and wind. It tracks the reference path with minimal error and reaches the target point. Our constructed time-varying marine environment model, which integrates multiple environmental factors, is as follows: Figure 2 As shown, arrows indicate the direction of ocean currents and wind fields, and the background color represents the intensity of the ocean current and wind field environment. The environmental factors at different locations and times are dynamically changing.

[0141] In addition, the flowchart of the deep reinforcement learning process is as follows: Figure 3 As shown, the environment is the marine environment in which the unmanned vessel moves, and the intelligence... The agent is an unmanned vessel, and the interaction between the agent and the environment is achieved through a Markov decision process. This invention provides a method for sequentially defining the process... Righteousness Figure 3 Status, action, and reward in It maximizes rewards through interaction with the environment and training.

[0142] Figure 4 The diagram shows the model optimization process, specifically the training process of the improved TD3 algorithm. The agent interacts with the marine environment through a Markov decision process, storing collected experience samples in an experience replay buffer. During the network parameter update phase, mini-batch samples are sampled from the experience replay buffer for network updates. The actor policy network learns superior control policies, and the critic value network guides the policy updates. The policy learning process is accelerated and stabilized through mechanisms such as policy delayed updates and soft updates.

[0143] Figure 5 The diagram shows the steps of the unmanned vessel path tracking control method based on deep reinforcement learning according to the present invention.

[0144] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0145] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0146] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A path tracking control method for unmanned surface vessels based on deep reinforcement learning, characterized in that: The implementation method specifically includes the following steps: Step 1: Construct a time-varying marine environment model that integrates multiple environmental factors; A time-varying marine environment model integrating multiple environmental factors was constructed, including: acquiring ocean current and wind field environmental data for a portion of the South China Sea (113.5°E-115.5°E, 17.25°N to 19.25°N) from the National Marine Science Data Center; dividing the region into a 120×120 grid using a grid method, and interpolating the ocean current and wind field environmental data into the grid environment using inverse distance weighted interpolation; and setting time windows for environmental data changes to describe the dynamic changes in the marine environment. When the USV is at position P, the environmental data at a certain time t is represented as follows: In the formula I P (t) represents the environmental information received in real time by the USV at position P and time t via sensors such as a flow profiler and anemometer. and These represent the lateral and longitudinal velocities of the ocean current, respectively. and These represent the lateral and longitudinal velocities of the wind, respectively. The environmental data I(t) at a non-grid point location P can be obtained from the four surrounding adjacent grid points P. i Environmental data I at (i = 1, ..., 4) i (t) is obtained by distance weighting and is calculated by the following formula: Step 2: Based on the marine environment model from Step 1, the unmanned vessel path tracking control problem is modeled as a solution framework based on a Markov decision process. The modeling content in Step 2 specifically includes the state space S, action space A, reward function R, state transition relation P, and discount factor γ of the Markov decision process. Here, the state space S represents the set of state information s perceived by the unmanned surface vessel, including motion parameter information s m Environmental information e LOS guidance information LOS The state s during the t-th control time. t Represented as: s t ={s m ,s e ,s LOS } Where: motion parameter information s m = {(x,y),U,ψ}, where (x,y) is the current position of the unmanned surface vessel (USV), and U and ψ are the current speed and heading angle of the USV, respectively; environmental data as described in step 1, environmental information s e =I P (t); LOS guidance information s LOS =(ψ d ,x e ,y e ), where ψ d For the ideal heading angle of the unmanned surface vessel, x e and y e These are the longitudinal tracking error and the lateral tracking error, respectively. The LOS guidance method is used to calculate the desired heading angle and tracking error to improve tracking accuracy. The desired heading angle can be expressed as: In the formula, ψ k The slope of the tracking path is Δ, where Δ is the forward-looking distance of the unmanned vessel and β is the sideslip angle. For the unmanned surface vessel at position (x, y), the parameterized path P is tracked. k The tracking error can be expressed as: (t), Wherein, the action space A represents the set of actions a of the unmanned vessel, and the action a at the t-th control time is... t Defined as the output speed and heading angle of the unmanned surface vessel (USV), the heading angle is controlled by controlling the direction of the rudder, while the speed of the USV is adjusted by controlling the thrust of the propeller, as shown in the following formula: In the formula, U(t) represents the velocity of the unmanned vessel. Indicates the heading angle of the unmanned vessel; The reward function R defines the unmanned surface vessel in different states s. t and action a t The reward signal r obtained below t It is used to guide the learning and decision-making process of control strategies, and is defined as follows: In the formula, K1 and K2 are proportional parameters assigned to different error penalty terms, and t step The total control time is calculated; in addition, when the unmanned vessel successfully completes the tracking task, it will receive a huge reward incentive; however, if the unmanned vessel fails to track within the specified number of steps or deviates from the tracking area, it will be subject to a huge penalty. The state transition relationship is defined as the probability that the unmanned vessel will update from the current environmental state to the next state after making an action decision, as shown in the following formula: P(S t+1 |S t ×[A] t )=ρ In the formula, ρ is the environmental transition rate, which takes the value of 1 in the unmanned vessel path tracking environment; The discount factor is defined as γ, which is a value between 0 and 1, representing the importance of future rewards. The larger the discount factor γ, the more attention is paid to the impact of long-term rewards on the control strategy. Step 3: Based on the solution framework of Step 2, the improved Twin Delayed Deep Deterministicpolicy gradient (TD3) algorithm is used to optimize the path tracking control of the unmanned vessel. Specific steps include: 3.1 Initialize the marine environment, perform network parameter initialization on the actor policy neural network and critic value neural network of the improved TD3 algorithm, and set up an experience replay buffer; 3.2 The unmanned vessel intelligent agent interacts with the marine environment, and the interaction process is the same as in step 2. The state information of the unmanned vessel is input into the actor policy network, and the output action of the actor policy network is calculated through forward propagation. After the unmanned vessel executes the action, it performs a state transition and calculates the reward. During this interaction, experience data is collected and stored in the experience replay buffer. This step is repeated until the experience replay buffer is full. 3.3 The Critic value network uses the Q-value function to evaluate the performance of the actor policy network and guide its next stage update; the actor policy network determines to take a better action by changing policy π to obtain a higher reward. 3.4 Using a two-state value network to address the overestimation of the Q-value function, the next state s t+1 The Q value can be estimated by two target value networks: In the formula, Q′1 and Q′2 are the target Q values ​​of the two value networks, π′ is the policy of the target policy network, and θ μ′ For the parameters of the target policy network, and These are the parameters for the two target value networks, respectively. Choosing the smaller target Q value, the target value of the value network is calculated according to the Bellman equation as follows: 3.5 Based on the priority experience sampling mechanism, N experience samples are sampled from the experience replay buffer for network updates. The state set S and action set A are input into the value neural network, and the parameters of the neural network are updated through backpropagation. The loss function of the value network is constructed using the following formula to guide policy learning and optimization: 3.6 Update the parameters of the critic value network by minimizing the loss function and gradient descent: In the formula It is the gradient of the loss function. α1 is the gradient of the Q value, and α1 is the learning rate of the value network. 3.7 A policy delay update mechanism is used to improve the stability of the policy network update. The policy network updates at a lower frequency than the value network. Gradient ascent is used to guide the policy network to update parameters θ in the direction of improving the Q-value. μ : In the formula, The objective function is J(θ) μ The gradient of ) It is strategy π(s) t |θ μ The gradient of ), where α2 is the learning rate of the policy network; 3.8 Using a soft update mechanism to update the parameters of the target network and θ μ′ To perform an update, the current network gradually updates the target network at a smaller update rate, as expressed by the following formula: i μ′ ←th μ +(1−e)θ μ′ In the formula, ε is the soft update rate; 3.9 Repeat steps 3.2-3.8 for a number of training iterations to finally obtain a policy network with excellent control strategy; Step 4: Use the policy neural network trained in Step 3 to control the unmanned vessel to perform path tracking tasks. Specific steps include: 4.1 Initialize the marine environment and load the trained policy neural network parameters; 4.2 Input the initial environmental state of the unmanned vessel into the policy network trained in step 3, and obtain the speed and heading control actions of the unmanned vessel through forward propagation inference; 4.3 The unmanned surface vessel performs output actions in the marine environment, tracks the reference path, and transitions to a new state; 4.4 Input the newly acquired state information into the policy network to generate new control actions for the unmanned vessel; 4.5 Repeat steps 4.3 and 4.4 until the unmanned vessel completes the path tracking task of the reference path.

Citation Information

Patent Citations

  • Unmanned ship path tracking method based on reinforcement learning

    CN112947431A

  • Unmanned ship path tracking method based on reinforcement learning and line-of-sight method

    CN113110504A