Unmanned ship path planning system based on under-actuated navigation deep reinforcement learning

By using under-driven navigation deep reinforcement learning algorithm in the unmanned ship path planning system to generate and track local paths, the problem that traditional systems cannot consider the dynamic constraints of unmanned ships is solved, and the safe and stable path planning and tracking of unmanned ships in a dynamic environment is realized.

CN119987351AActive Publication Date: 2025-05-13DALIAN MARITIME UNIVERSITY
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202411954823.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing traditional path planning system does not consider the dynamic constraints of under-driven unmanned ships, and cannot guarantee that the path is accurately tracked by unmanned ships, or even that the path is feasible.

Method used

Adopt an unmanned ship path planning system based on underdrive navigation deep reinforcement learning, including an environment perception module, a path planner and a guidance controller. The status information of the unmanned ship is collected through multiple sensors, and local paths are generated using the trained Actor-Critic deep reinforcement learning algorithm model, and the unmanned ship is driven to track these paths through the guidance control algorithm.

Benefits of technology

It realizes the unmanned ship directing to the target area without collision in a dynamic environment, ensuring path feasibility and traceability, reducing dependence on high-precision mathematical models, and improving immunity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987351A_ABST
    Figure CN119987351A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned ship path planning system based on under-actuated navigation deep reinforcement learning. The unmanned ship path planning system comprises an environment sensing module used for collecting pose and speed information of an unmanned ship and distance information of the unmanned ship and an obstacle through various sensors; the path planner is used for gradually generating each section of local path of the unmanned ship in the process from the starting point to the terminal point through a trained under-actuated navigation Actor-Critic deep reinforcement learning algorithm model based on a state vector generated by the pose and speed information of the unmanned ship and the obstacle distance information collected by the environment sensing module; and the guidance controller is used for tracking the local path by adopting a guidance control algorithm on the basis of each section of local path generated by the path planner until the unmanned ship reaches the final target area from the starting point of the local path to the end point of the local path. According to the method, the control process is considered in the training and execution process, it is ensured that the planned path can be accurately tracked, and the underactuated USV can reach the target position without collision in the dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of fully automated products and relates to an unmanned ship path planning system based on under-actuated navigation deep reinforcement learning. Background Art

[0002] Compared with traditional manual methods, the use of surface unmanned vessels for autonomous water quality monitoring has significant advantages in reducing labor costs and improving operational efficiency. However, the near-shallow sea environment is complex and there are various static and dynamic obstacles. Whether the unmanned vessel can operate normally depends on its autonomous navigation ability in complex environments. For under-actuated unmanned vessels, their motion performance is limited and they are usually unable to flexibly perform obstacle avoidance operations such as steering and speed control. If the special under-actuated dynamic characteristics of the unmanned vessel are not considered, it cannot be guaranteed that the planned path can be accurately tracked, resulting in mission failure. Therefore, when performing path planning, the dynamic characteristics of the under-actuated unmanned vessel must be considered to ensure that the unmanned vessel completes the task safely and stably.

[0003] Traditional path planning methods often plan paths based on static maps or assume that the speed of unmanned ships can be adjusted in time and directly output the expected speed. Both planning schemes fail to consider the dynamic constraints of underactuated unmanned ships and make it difficult to accurately guide unmanned ships to avoid obstacles and reach the target area. Smoothing the path through the path smoothing method can make the path easier to track, but this method is suitable for global path planning and requires frequent replanning in a dynamic environment, which consumes a lot of computing resources. With the development of artificial intelligence and machine learning technologies, deep reinforcement learning has shown great potential in the field of path planning. Deep reinforcement learning simulates the learning process of humans, allowing unmanned ships to continuously learn and optimize their motion strategies in the environment. However, there are still some difficulties in applying deep reinforcement learning to path planning for surface unmanned ships, such as high model complexity, long training time, and poor strategy generalization ability. Some reinforcement learning-based methods train control strategies based on the mathematical model of the unmanned ship, and output force and torque to control the motion of the underactuated unmanned ship. This scheme relies on a high-precision mathematical model of the unmanned ship, and the obtained strategy is usually difficult to adapt to environmental disturbances, and the actual applicability of the method is poor. Therefore, when navigating an unmanned ship, in addition to considering the dynamic characteristics of the unmanned ship, it is also very important to ensure that the strategy does not rely on high-precision mathematical models and has a certain degree of anti-interference.

[0004] The existing traditional path planning system does not take into account the dynamic constraints of the under-actuated unmanned ship, and cannot ensure that the path is accurately tracked by the under-actuated unmanned ship or even that the path is feasible. Summary of the invention

[0005] In order to solve the shortcomings of the existing traditional path planning system that does not consider the dynamic constraints of the under-actuated unmanned ship, cannot ensure that the path is accurately tracked by the under-actuated unmanned ship, and cannot even ensure that the path is feasible, the technical solution adopted by the present invention is: an unmanned ship path planning system based on under-actuated navigation deep reinforcement learning, characterized in that it includes: an environmental perception module: used to collect the position and speed information of the unmanned ship and the distance information from obstacles through a variety of sensors;

[0006] Path planner: used to generate a state vector based on the position, speed and distance information of the unmanned ship collected by the environment perception module, and gradually generate each local path of the unmanned ship from the starting point to the end point through the trained underactuated navigation Actor-Critic deep reinforcement learning algorithm model;

[0007] Guidance controller: used to drive the unmanned ship to track each local path generated by the path planner, and use the guidance control algorithm to track the local path from the starting point of the local path to the end point of the local path until the unmanned ship reaches the final target area, thus completing the path planning task.

[0008] Further: the environment perception module includes:

[0009] IMU: collects the position information of the unmanned ship.

[0010] GPS: collects the absolute position information of the unmanned ship;

[0011] LiDAR: Obtain the relative distance information between obstacles at a specific angle to the bow of the unmanned ship and the unmanned ship.

[0012] Further: the process of gradually generating each local path of the unmanned ship from the starting point to the end point based on the underactuated navigation Actor-Critic deep reinforcement learning algorithm model is as follows:

[0013] The state vector of the unmanned ship is input into the Actor network of the trained underactuated navigation Actor-Critic, and the Actor network of the trained underactuated navigation Actor-Critic outputs the action vector a k , and the position of the unmanned ship p: = [x, y] is added to get the waypoint position p of the unmanned ship w :=[x w ,y w ] T .

[0014] Further: the underactuated navigation Actor-Critic deep reinforcement learning algorithm model includes:

[0015] Actor network: parameter is θ μ, the input is the state vector s of the unmanned ship k , the output is the action vector a k and the probability of choosing this action

[0016] Two Critic networks: the parameters are Input is a multi-step state vector and the action vector a k , the output is the state action value Used to guide the Actor network to obtain the optimal strategy during training;

[0017] Two Critic target networks: the parameters are The input is a multi-step state vector and action vector, and the output is a state action value Used to update the Critic network and its own parameters.

[0018] Further: the Actor network includes two first fully connected layers with ReLU activation function and one first linear layer; the first fully connected layer and the first linear layer are sequentially cascaded;

[0019] The Critic network includes a first LSTM layer, a second fully connected layer with an activation function of ReLU, a third connected layer and a linear layer; the first LSTM layer, the second fully connected layer, the third fully connected layer and the second linear layer are sequentially cascaded;

[0020] The structure of the Critic target network is the same as that of the Critic network, including a second LSTM layer, a fourth fully connected layer with ReLU activation function, a fifth fully connected layer, and a second linear layer;

[0021] The second LSTM layer, the fourth fully connected layer, the fifth fully connected layer and the third linear layer are cascaded sequentially.

[0022] Further: The parameter update process of the underactuated navigation Actor-Critic deep reinforcement learning algorithm model is as follows:

[0023] Optimize the Actor network parameters θ by gradient descent μ , the loss function used is as follows:

[0024]

[0025] Among them, s k From the experience replay buffer Sampling acquisition, is the minimum state action value output by the two Critic networks, and α is the entropy temperature coefficient;

[0026] The loss function of the Critic network is defined as follows:

[0027]

[0028] Among them, y is obtained from The reward for sampling r k and the state-action value from the target critic network The target values ​​are as follows:

[0029]

[0030] Among them, γ is the discount factor;

[0031] In addition, the loss function of the entropy temperature coefficient α is determined by the following formula:

[0032]

[0033] in: represents the target entropy;

[0034] According to the loss function of the Actor network, the loss function of the Critic network, and the loss function of the entropy temperature coefficient α, the parameters θ μ and α are adaptively updated via gradient descent as follows:

[0035]

[0036] Among them, λ Q ,λ μ ,λ α >0 is the learning rate.

[0037] In addition, the parameters of the target network are updated by soft updating:

[0038]

[0039] Among them, ∈ is a sufficiently small constant.

[0040] Further: the process for driving the unmanned ship to track each local path generated by the path planner, using a guidance control algorithm to track the local path from the starting point of the local path to the end point of the local path until the unmanned ship reaches the final target area is as follows:

[0041] From the initial state of the unmanned ship's starting point to the final state of the end point, the guidance algorithm is used to calculate the expected heading and expected speed according to the unmanned ship's posture, speed and path: the tracking guidance algorithm is used to obtain the expected heading, and the expected speed is preset as U d =4m / s, expected heading ψ d Dynamically determined by:

[0042]

[0043] in, Indicates the angle between the line connecting the position of the unmanned ship and the position of the waypoint and the x-axis, is the sideslip angle, u and v represent the longitudinal and transverse speeds of the unmanned ship respectively;

[0044] Use the control algorithm to calculate the force and torque of the under-actuated unmanned ship according to the expected speed and heading as well as the current speed and heading of the unmanned ship to control the movement of the unmanned ship:

[0045] The incremental PID controller is used to obtain the longitudinal force and torque of the unmanned ship. The incremental longitudinal force and torque output by the controller at time h are determined by the following formula:

[0046]

[0047] in, is the controller parameter, t c is the sampling period, U e =U d -U and ψ e =ψ d -ψ are the speed and heading errors, U is the speed of the unmanned ship, ψ is the heading of the unmanned ship, and the longitudinal force and torque acting on the unmanned ship are:

[0048] τ u (h) = τ u (ht c )+Δτ u (h) (21)

[0049] τ r (h) = τ r (ht c )+Δτ r (h) (22)

[0050] The saturation constraints for force and torque are This can control the movement of the unmanned ship;

[0051] Generate the next local path one by one and track it. When the unmanned ship has reached the final target area, the path planning task is completed.

[0052] The present invention provides an unmanned ship path planning system based on deep reinforcement learning of underactuated navigation, which is an unmanned ship path planning system based on deep reinforcement learning of underactuated navigation, realizing an unmanned ship path planning scheme integrating functions such as environment perception, path planning, and path tracking control, and has the following advantages: Compared with the prior art, the present invention has the following advantages: The beneficial effects of the present invention are: high-precision dynamic obstacle perception and acquisition of accurate posture information of the unmanned ship are achieved through multiple sensors, the dimension of state data is reduced by embedding the LSTM layer in the Critic network, the algorithm learning efficiency is improved, the control process is considered during training and execution, and it is ensured that the planned path can be accurately tracked. The underactuated USV can reach the target position without collision in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0054] Figure 1 It is a schematic diagram of an unmanned ship path planning system based on underactuated navigation Actor-Critic deep reinforcement learning provided by the present invention;

[0055] Figure 2 It is a schematic diagram of the Critic network structure provided by the present invention;

[0056] Figure 3 It is a schematic diagram of the Actor network structure provided by the present invention;

[0057] Figure 4 It is a schematic diagram of the underactuated navigation deep reinforcement learning framework composed of an Actor network and a Critic network provided by the present invention;

[0058] Figure 5 It is a schematic diagram of the training process of the Actor-Critic neural network of the unmanned ship path planning system based on underactuated navigation deep reinforcement learning provided by the present invention;

[0059] Figure 6 is a schematic diagram of a prototype of an unmanned ship according to a specific implementation of the present invention;

[0060] Figure 7Schematic diagrams of the performance of the unmanned ship path planning system based on under-actuated navigation deep reinforcement learning provided by the present invention in a harbor environment, (a) is the performance schematic diagram I, (b) is the performance schematic diagram II, (c) is the performance schematic diagram III, and (d) is the performance schematic diagram IV. DETAILED DESCRIPTION

[0061] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0062] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0063] Figure 1 It is a schematic diagram of an unmanned ship path planning system based on underactuated navigation Actor-Critic deep reinforcement learning provided by the present invention;

[0064] An unmanned ship path planning system based on underactuated navigation deep reinforcement learning, comprising:

[0065] Environmental perception module: used to collect the position, speed and distance information of the unmanned ship from obstacles through a variety of sensors;

[0066] Path planner: used to generate a state vector based on the position, speed and distance information of the unmanned ship collected by the environment perception module, and gradually generate each local path of the unmanned ship from the starting point to the end point through the trained underactuated navigation Actor-Critic deep reinforcement learning algorithm model;

[0067] Guidance controller: used to drive the unmanned ship to track each local path generated by the path planner, and use the guidance control algorithm to track the local path from the starting point of the local path to the end point of the local path until the unmanned ship reaches the final target area, thereby completing the path planning task.

[0068] Furthermore: the multi-sensors carried by the unmanned ship include RTK, IMU, Doppler odometer, and laser radar to obtain the position, speed, and distance information of the unmanned ship from obstacles;

[0069] The environment perception module comprises:

[0070] IMU: Obtain the relative position change of the carrier in a short period of time by integration to obtain the position and posture information of the unmanned ship;

[0071] GPS: Collect the absolute position information of the unmanned ship; use the multi-sensor fusion algorithm to solve the absolute position information provided by GPS to adjust the state estimation of IMU, eliminate the accumulated error of IMU, and obtain the accurate position information of the unmanned ship;

[0072] LiDAR: Obtain the relative distance information between obstacles at a specific angle to the bow of the unmanned ship and the unmanned ship.

[0073] The above information is combined into a column vector to form a state vector s that can be used by the underactuated navigation Actor-Critic network. k .

[0074] Further: the process of gradually generating each local path of the unmanned ship from the starting point to the end point based on the underactuated navigation Actor-Critic deep reinforcement learning algorithm model is as follows:

[0075] The state vector of the unmanned ship is input into the Actor network of the trained underactuated navigation Actor-Critic, and the Actor network of the trained underactuated navigation Actor-Critic outputs the action vector a k , and the position of the unmanned ship p: = [x, y] is added to get the waypoint position p of the unmanned ship w :=[x w ,y w ] T ;

[0076] A local path is formed by connecting the UAV position with the waypoint positions.

[0077] Further: the underactuated navigation Actor-Critic deep reinforcement learning algorithm model includes:

[0078] Actor network: parameter is θ μ , the input is the state vector s of the unmanned ship k , the output is the action vector a k and the probability of choosing this action

[0079] Two Critic networks: the parameters are Input is a multi-step state vector and the action vector a k , the output is the state action value Used to guide the Actor network to obtain the optimal strategy during training;

[0080] Two Critic target networks: the parameters are The input is a multi-step state vector and action vector, and the output is a state action value Used to update the Critic network and its own parameters;

[0081] Figure 2 It is a schematic diagram of the Critic network structure provided by the present invention;

[0082] Figure 3 It is a schematic diagram of the Actor network structure provided by the present invention;

[0083] Figure 4 It is a schematic diagram of the under-actuated navigation deep reinforcement learning framework composed of an Actor network and a Critic network provided by the present invention.

[0084] The Actor network includes two first fully connected layers with ReLU activation functions and one first linear layer; the first fully connected layer and the first linear layer are sequentially cascaded;

[0085] The Critic network includes a first LSTM layer, a second fully connected layer with an activation function of ReLU, a third connected layer and a linear layer; the first LSTM layer, the second fully connected layer, the third fully connected layer and the second linear layer are sequentially cascaded;

[0086] The structure of the Critic target network is the same as that of the Critic network, including a second LSTM layer, a fourth fully connected layer with ReLU activation function, a fifth fully connected layer, and a second linear layer;

[0087] The parameter update process of the underactuated navigation Actor-Critic deep reinforcement learning algorithm model is as follows:

[0088] Optimize the Actor network parameters θ by gradient descent μ , the loss function used As shown below:

[0089]

[0090] Among them, s k From the experience replay buffer Sampling acquisition, is the minimum state action value output by the two Critic networks, and α is the entropy temperature coefficient;

[0091] Similarly, the loss function of the Critic network is The definition is as follows:

[0092]

[0093] Among them, y is obtained from The reward for sampling r k and the state-action value from the target critic network The target values ​​are as follows:

[0094]

[0095] Among them, γ is the discount factor.

[0096] In addition, the loss function of the entropy temperature coefficient α is determined by the following formula:

[0097]

[0098] in Represents the target entropy, which is an empirical value.

[0099] According to the loss function of the Critic network, the loss function of the Actor network, and the loss function of the entropy temperature coefficient α, the parameters θ μ and α can be adaptively updated via gradient descent as follows:

[0100]

[0101] Among them, λ Q ,λ μ ,λ α >0 is the learning rate.

[0102] In addition, the parameters of the target network are updated via soft updating;

[0103]

[0104] Among them, ∈ is a sufficiently small constant.

[0105] Figure 5 It is a schematic diagram of the training process of the Actor-Critic neural network of the unmanned ship path planning system based on under-actuated navigation deep reinforcement learning provided by the present invention.

[0106] Furthermore, the specific steps of the training process of the Actor-Critic neural network of the unmanned ship path planning system based on underactuated navigation deep reinforcement learning are as follows:

[0107] S1: Reset all network parameters: and θ μ and clear

[0108] S2: Determine whether the current round is less than the maximum number of rounds;

[0109] S21: If the current round is less than the maximum number of rounds, go to step S3;

[0110] S22: If the current round is greater than or equal to the maximum number of rounds, end the training;

[0111] S3: Reset the environment state at the beginning of each round of training, including the position, speed, force and torque of the unmanned ship, the position of obstacles, and the position of the target area;

[0112] S4: Obtain the state vector of the unmanned ship through multiple sensors;

[0113] S5: Determine whether the current number of rounds is greater than the initial number of rounds;

[0114] S51: If the number of rounds currently performed is less than or equal to the initial number of rounds, a waypoint position is randomly generated;

[0115] S52: If the current number of rounds is greater than the initial number of rounds, the state vector is input into the Actor network to generate a waypoint position;

[0116] S6: Determine whether the distance between the unmanned ship and the waypoint is greater than a threshold;

[0117] S61: If it is greater than the threshold, the unmanned ship has not reached the waypoint, and the guidance algorithm is used to calculate the expected speed and heading of the unmanned ship at this time;

[0118] S62: If it is less than the threshold, the unmanned ship has reached the waypoint, and go to step (12);

[0119] S7: using a control algorithm to calculate the force and torque for controlling the unmanned ship;

[0120] S8: Control the movement of the unmanned ship;

[0121] S9: Determine whether the unmanned ship is in a feasible area;

[0122] S10: If the unmanned ship is in the feasible area, return to step (6);

[0123] S11: If the unmanned ship is not in the feasible area, the environment is reset to the state before the unmanned ship obtains the waypoint;

[0124] S12: Get the new state vector of the unmanned ship and calculate the reward. The reward function used is:

[0125]

[0126] Among them, Δd is the Euclidean distance from the unmanned ship to the center point of the target area;

[0127] S13: Store (state, action, new state, reward) into the experience replay buffer pool;

[0128] S14: Determine whether the number of storage times is greater than the sampling batch;

[0129] S141: If the number of storage times is greater than the sampling batch, sampling is performed and the network parameters are updated using the gradient descent method, and then the process goes to step S15;

[0130] S142: If the number of storage times is less than the sampling batch, no operation is performed and the process goes to step S15;

[0131] S15: Determine whether the unmanned ship is in the target area;

[0132] S151: If the unmanned ship is in the target area, go to step S2;

[0133] S152: If the unmanned ship is not in the target area, go to step S5;

[0134] Through training, a trained Actor-Critic network model that drives the path planning of the unmanned ship can be obtained.

[0135] Further: the process for driving the unmanned ship to track each local path generated by the path planner, from the starting point of the local path to the end point of the local path, until the unmanned ship reaches the final target area is as follows:

[0136] From the initial state of the unmanned ship's starting point to the final state of the end point, the guidance algorithm is used to calculate the desired heading and desired speed based on the unmanned ship's position, speed and path;

[0137] In the specific implementation, the tracking guidance algorithm is used to obtain the desired heading, and the desired speed is preset as U d =4m / s, expected heading ψ d Dynamically determined by:

[0138]

[0139] in, Indicates the angle between the line connecting the position of the unmanned ship and the position of the waypoint and the x-axis, is the sideslip angle, u and v represent the longitudinal and transverse velocities of the unmanned ship respectively.

[0140] Use the control algorithm to calculate the force and torque of the underactuated unmanned ship according to the expected speed and expected heading as well as the current speed and heading of the unmanned ship, and control the movement of the unmanned ship;

[0141] In the specific implementation, an incremental PID controller is used to obtain the longitudinal force and torque of the unmanned ship. The incremental longitudinal force and torque output by the controller at time h are determined by the following formula:

[0142]

[0143] in, is the controller parameter, t c is the sampling period, U e =U d -U and ψ e =ψ d -ψ are the speed and heading errors, U is the speed of the unmanned ship, and ψ is the heading of the unmanned ship. The longitudinal force and torque acting on the unmanned ship are:

[0144] τ u (h) = τ u (ht c )+Δτ u (h) (33)

[0145] τ r (h) = τ r (ht c )+Δτ r (h) (34)

[0146] The saturation constraints for force and torque are This can control the movement of the unmanned ship.

[0147] Generate local paths one by one and track them. When the unmanned ship has reached the final target area, the path planning task is completed.

[0148] Figure 6 is a schematic diagram of a prototype of an unmanned ship according to a specific implementation of the present invention;

[0149] Figure 7 Schematic diagrams of the performance of the unmanned ship path planning system based on underactuated navigation deep reinforcement learning provided by the present invention in a harbor environment; (a) is a schematic diagram of performance I, (b) is a schematic diagram of performance II, (c) is a schematic diagram of performance III, and (d) is a schematic diagram of performance IV;

[0150] In Lingshui Port waters, Figure 6 The prototype of the "Zhi Hai No. 2" unmanned ship shown in the figure verifies the effectiveness of the proposed path planning system based on deep reinforcement learning for underactuated navigation. The initial position of the unmanned ship is [150, -285] T ,The final target area is randomly selected in the harbor, t c =0.1s,

[0151] Using the Actor network trained for 100,000 rounds to obtain the path, the path planning effect in the four final target area scenarios is as follows: Figure 7 shown.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An unmanned ship path planning system based on underactuated navigation deep reinforcement learning, characterized by: include: Environmental perception module: used to collect the position, speed and distance information of the unmanned ship from obstacles through a variety of sensors; Path planner: used to generate a state vector based on the position, speed and distance information of the unmanned ship collected by the environment perception module, and gradually generate each local path of the unmanned ship from the starting point to the end point through the trained underactuated navigation Actor-Critic deep reinforcement learning algorithm model; Guidance controller: used to drive the unmanned ship to track each local path generated by the path planner, and use the guidance control algorithm to track the local path from the starting point of the local path to the end point of the local path until the unmanned ship reaches the final target area, thereby completing the path planning task.

2. The unmanned ship path planning system based on underactuated navigation deep reinforcement learning according to claim 1, characterized in that: The environment perception module comprises: IMU: collects the position information of the unmanned ship. GPS: collects the absolute position information of the unmanned ship; LiDAR: Obtain the relative distance information between obstacles at a specific angle to the bow of the unmanned ship and the unmanned ship.

3. The unmanned ship path planning system based on underactuated navigation deep reinforcement learning according to claim 1, characterized in that: The process of gradually generating each local path of the unmanned ship from the starting point to the end point based on the underactuated navigation Actor-Critic deep reinforcement learning algorithm model is as follows: The state vector of the unmanned ship is input into the Actor network of the trained underactuated navigation Actor-Critic, and the Actor network of the trained underactuated navigation Actor-Critic outputs the action vector a k , and the position of the unmanned ship p: = [x, y] is added to get the waypoint position p of the unmanned ship w :=[x w ,y w ] T .

4. The unmanned ship path planning system based on underactuated navigation deep reinforcement learning according to claim 1, characterized in that: The underactuated navigation Actor-Critic deep reinforcement learning algorithm model includes: Actor network: parameter is θ μ , the input is the state vector s of the unmanned ship k , the output is the action vector a k and the probability of choosing this action Two Critic networks: the parameters are Input is a multi-step state vector and the action vector a k , the output is the state action value Used to guide the Actor network to obtain the optimal strategy during training; Two Critic target networks: the parameters are The input is a multi-step state vector and action vector, and the output is a state action value Used to update the Critic network and its own parameters.

5. The unmanned ship path planning system based on underactuated navigation deep reinforcement learning according to claim 1, characterized in that: The Actor network includes two first fully connected layers with ReLU activation functions and one first linear layer; the first fully connected layer and the first linear layer are sequentially cascaded; The Critic network includes a first LSTM layer, a second fully connected layer with an activation function of ReLU, a third connected layer and a linear layer; the first LSTM layer, the second fully connected layer, the third fully connected layer and the second linear layer are sequentially cascaded; The structure of the Critic target network is the same as that of the Critic network, including a second LSTM layer, a fourth fully connected layer with ReLU activation function, a fifth fully connected layer, and a second linear layer; The second LSTM layer, the fourth fully connected layer, the fifth fully connected layer and the third linear layer are cascaded sequentially.

6. The unmanned ship path planning system based on underactuated navigation deep reinforcement learning according to claim 1, characterized in that: The parameter update process of the underactuated navigation Actor-Critic deep reinforcement learning algorithm model is as follows: Optimize the Actor network parameters θ by gradient descent μ , the loss function used is as follows: Among them, s k From the experience replay buffer Sampling acquisition, is the minimum state action value output by the two Critic networks, and α is the entropy temperature coefficient; The loss function of the Critic network is defined as follows: Among them, y is obtained from The reward for sampling r k and the state-action value from the target critic network The target values ​​are as follows: Among them, γ is the discount factor; In addition, the loss function of the entropy temperature coefficient α is determined by the following formula: in: represents the target entropy; According to the loss function of the Actor network, the loss function of the Critic network, and the loss function of the entropy temperature coefficient α, the parameters θ μ and α are adaptively updated via gradient descent as follows: Among them, λ Q ,λ μ ,λ α >0 is the learning rate. In addition, the parameters of the target network are updated by soft updating: Among them, ∈ is a sufficiently small constant.

7. The unmanned ship path planning system based on underactuated navigation deep reinforcement learning according to claim 1, characterized in that: The process for driving the unmanned ship to track each local path generated by the path planner, using a guidance control algorithm to track the local path from the starting point of the local path to the end point of the local path until the unmanned ship reaches the final target area is as follows: From the initial state of the unmanned ship's starting point to the final state of the end point, the guidance algorithm is used to calculate the expected heading and expected speed according to the unmanned ship's posture, speed and path: the tracking guidance algorithm is used to obtain the expected heading, and the expected speed is preset as U d =4m / s, expected heading ψ d Dynamically determined by: in, Indicates the angle between the line connecting the position of the unmanned ship and the position of the waypoint and the x-axis, is the sideslip angle, u and v represent the longitudinal and transverse speeds of the unmanned ship respectively; Use the control algorithm to calculate the force and torque of the under-actuated unmanned ship according to the expected speed and heading as well as the current speed and heading of the unmanned ship to control the movement of the unmanned ship: The incremental PID controller is used to obtain the longitudinal force and torque of the unmanned ship. The incremental longitudinal force and torque output by the controller at time h are determined by the following formula: in, is the controller parameter, t c is the sampling period, U e =U d -U and ψ e =ψ d -ψ are the speed and heading errors, U is the speed of the unmanned ship, ψ is the heading of the unmanned ship, and the longitudinal force and torque acting on the unmanned ship are: t u (h)=τ u (ht c )+Δτ u (h) (10) t r (h)=τ r (ht c )+Δτ r (h) (11) The saturation constraints for force and torque are This can control the movement of the unmanned ship; Generate the next local path one by one and track it. When the unmanned ship has reached the final target area, the path planning task is completed.

Citation Information

Patent Citations

  • Unmanned ship autonomous navigation method based on deep reinforcement learning and genetic algorithm

    CN110362089A

  • Collision avoidance planning method for mobile robots based on deep reinforcement learning in dynamic environment

    CN110632931A

  • Predictive active-disturbance-rejection stabilization control method of rudder fin system

    CN114815626A

  • Unmanned ship cluster task scheduling and collaborative confrontation method based on MADDPG

    CN116050795A

  • Unmanned ship path planning method and system considering environmental disturbance and multi-target constraint

    CN116700269A