An obstacle avoidance method for an unmanned ship
By adopting a deep Q-learning algorithm based on the Markov process framework in the unmanned ship obstacle avoidance algorithm, combining the value of decomposition state and adding noise, the obstacle avoidance decision and motion control of unmanned ships is optimized, and the problem of rapid changes in angle and speed in traditional algorithms is solved, and the obstacle avoidance efficiency and safety of unmanned ships are improved.
Patent Information
- Application Number
- CN202211497994.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-27
AI Technical Summary
The traditional unmanned ship obstacle avoidance algorithm has problems such as the angle and speed change too quickly and the number of changes is too much, resulting in unnecessary mechanical losses.
The deep Q learning algorithm based on the Markov process framework is adopted, combined with the decomposition state value and the method of adding noise, and the obstacle avoidance decision and motion control of the unmanned ship is optimized through priority sampling and n-step playback based on the ensemble tree.
It improves the safety and efficiency of the unmanned ship's obstacle avoidance and reaches the target point, and reduces mechanical losses.
Smart Images

Figure CN115752475B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of obstacle avoidance for unmanned vessels, and particularly relates to an obstacle avoidance method for unmanned vessels. Background Art
[0002] With the development of technology and the continuous introduction of navigation algorithms for unmanned vessels, the research on unmanned vessels has entered the fast lane. In recent years, unmanned vessels with low manpower requirements and high tracking accuracy have become active in fields such as water quality detection, underwater terrain surveying, and water body cleaning. An important symbol of the intelligence of unmanned vessels is autonomous navigation, and obstacle avoidance is one of the basic requirements for realizing the autonomous navigation of unmanned vessels. Obstacle avoidance means that when an unmanned vessel collects the state information of surrounding obstacles through sensors and perceives that it hinders its normal passage, it performs effective obstacle avoidance according to a certain method and finally reaches the target point. In the actual application of unmanned vessels, the environment in which they are located is dynamically variable and unknown. To solve the above problems, some algorithms in the fields of computers and artificial intelligence are introduced in the decision-making of unmanned vessels. Thanks to the improvement of the computing power of processors and the development of sensor technology, it is easier to perform complex intelligent algorithm operations on the unmanned vessel platform. However, in traditional intelligent methods, the genetic algorithm has a slow convergence speed, poor local search ability, and requires encoding and decoding, while the neural network has a long development time, requires a large amount of data, and the process is opaque, and the fuzzy algorithm cannot make full use of the obtained information. Therefore, the deep Q-learning algorithm with high algorithm efficiency and few sampled data has emerged, and it has good performance in aspects such as unmanned vessel navigation, obstacle avoidance decision-making, and obstacle avoidance control. However, most algorithms have problems such as complex network structures and low accuracy, and only consider the performance during decision-making during application, ignoring the kinematic performance under the corresponding decision-making, that is, there are problems of too fast and too many changes in the angle and speed when the unmanned vessel moves, resulting in unnecessary mechanical losses. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide an obstacle avoidance method for unmanned vessels, which reduces the too fast and too many changes in the angle and speed during the obstacle avoidance of unmanned vessels and avoids unnecessary mechanical losses.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] An obstacle avoidance method for unmanned vessels, comprising the following steps:
[0006] Step S1: Construct a two-dimensional environment, and generate static and dynamic obstacles, a target point, and the starting point of the unmanned vessel;
[0007] Step S2: Establish an environment model based on the Markov process framework, including the state space s t , action space a, reward r, and termination flag d;
[0008] Step S3: The unmanned ship starts to move after selecting an action based on the set random seed. After executing the action, it interacts with the environment to generate a new state and calculates the reward, and stores the state s t , action a, reward r, next state s t+1 , and termination flag d as a five-tuple in the database for learning for sampling training;
[0009] Step S4: Sample the environment model for network update and assign values to the sample state and action;
[0010] Step S5: Select the optimal action based on the value of each action in the sample state, and then obtain the optimal policy; if the maximum number of steps is not reached or the target point is not reached, return to Step S4.
[0011] Furthermore, the two-dimensional environment includes an unmanned ship with a warning range and collision monitoring points set, the unmanned ship trajectory, static obstacles at different positions, a dynamic obstacle moving from the upper right to the lower left along the unmanned ship route, and a circular target point in the upper left corner.
[0012] Furthermore, the reward r is expressed by the formula:
[0013]
[0014] In the formula, ship_length is the length of the unmanned ship, r is the total reward during the movement of the unmanned ship, r 1 is the reward generated by the change in the distance between the unmanned ship and the target point before reaching the target point, dist goal is the distance between the unmanned ship and the target point, dist prev is the distance between the unmanned ship and the target point in the previous state; r 2 is the reward for the unmanned ship moving towards the target point; r 3 is the reward obtained when the unmanned ship reaches the target point; r 4 is the penalty when the unmanned ship collides with an obstacle, DCPA is the minimum distance between the unmanned ship and the obstacle, dist warn is the size of the warning range of the unmanned ship; r 5 is the penalty when an obstacle appears within the warning range of the unmanned ship.
[0015] Furthermore, when storing in Step S3, calculate the priority of the current five-tuple according to the temporal difference algorithm, that is:
[0016] Priority = TD error (Q target - Q value )
[0017] In the formula, TD error is the temporal difference algorithm, Q target is the target value network, Qvalue is the current value network;
[0018] Store the five - tuple expressing the state of the unmanned ship in the database, and store the priority in the corresponding set tree of the database.
[0019] Furthermore, the network update is specifically as follows:
[0020] (1), Initialize the current network and the target network by adding noise ε to the nodes;
[0021] (2), Select an action a, and the action selection formula is as follows:
[0022]
[0023] In the formula, a best is the best action, represents the most valuable action a selected by the value maximization function under the state s and the policy π with network parameters θ ;
[0024] (3), Update the state s of the unmanned ship according to the action a in (2) t+1 , separately calculate the state value V(s) and the action value A(s,a), and then sum them to obtain Q(s,a). The formula is expressed as:
[0025]
[0026] In the formula, q π (s,a) is the value under the policy π(a|s), and a′ is other actions under the state s.
[0027] Calculate the reward return r, and the formula is as follows:
[0028]
[0029] In the formula, r j is the value obtained after the current value network takes the action a, γ is the discount rate of the target value network, d = False and d = True are respectively the judgments of whether the termination condition is met or not;
[0030] (4), After n - step replay, store (s t ,a,r,s t+1 ,d) in the database. If the database is not full, store it directly; otherwise, delete the oldest data and then store it; update the priority value in the corresponding set tree of the database;
[0031]
[0032] In the formula, t is the current step, n is the n - step replay, k is the intermediate step between t and n, G t:t+nis the value after playing back n steps starting from t;
[0033] (5), if the minimum learning batch is reached, sample and learn from the database according to the n-step replay priority and update the current network. The update formula is as follows:
[0034] Q new (s t ,a t ) = Q old (s t ,a t ) + α(G t:t+n - Q old (s t ,a t ))
[0035] In the formula, Q new is the new network, Q old is the old network, s t is the state at the t-th step, a t is the action at the t-th step.
[0036] (6), every time n steps are reached, update the target network. If not reached, return to (2). The update formula is as follows:
[0037]
[0038] Furthermore, it further includes step S6. When the maximum number of steps is reached or the target point is reached, substitute the change in the state velocity direction of this round into the kinematic formula:
[0039]
[0040] In the formula, is the positive definite symmetric inertia matrix of the added mass, is the Coriolis force and centripetal force matrix, is the linear damping matrix, is the kinematic force matrix, τ 1 is the longitudinal force, τ 2 is the lateral force, τ 3 is the yaw moment;
[0041] Calculate T to obtain τ in this round 1 τ 2 τ 3 Analyze the curve and end this round. If the maximum total number of steps limit is reached, end the training. Otherwise, return to step S3.
[0042] The present invention has the following beneficial effects compared with the prior art:
[0043] The present invention combines the method of decomposing state value and adding noise, and in terms of sampling learning, it uses the combination of priority sampling based on set tree and n-step replay, which improves the sampling effectiveness during the training process and effectively enhances the safety and efficiency of the unmanned ship during obstacle avoidance and reaching the target point; Description of the Drawings
[0044] Figure 1 is the schematic diagram of the method flow of the present invention;
[0045] Figure 2 is the schematic diagram of the two-dimensional environment in an embodiment of the present invention;
[0046] Figure 3 is the reward setting diagram in an embodiment of the present invention
[0047] Figure 4 is the network structure diagram in an embodiment of the present invention;
[0048] Figure 5 is the two obstacle avoidance processes when the unmanned ship encounters obstacles in an embodiment of the present invention. Detailed Embodiment
[0049] The present invention will be further described below in conjunction with the drawings and embodiments.
[0050] Aiming at the situation that traditional algorithms overestimate the action value, have high network complexity and large fluctuations during training, the present invention combines the advantages of several algorithms. Based on a dual network in the network structure, it combines the method of decomposing state value and adding noise.
[0051] The update formula of the traditional deep Q-learning obstacle avoidance algorithm is as follows:
[0052]
[0053] In the formula, Q Target is the value of the target network, r is the reward, γ is the discount rate, θ is the network parameter, and Q(s′, a′; θ) is the value of the unmanned ship taking the a′ action in the s′ state of the current network.
[0054] The loss function L(θ) of the traditional deep Q-learning obstacle avoidance algorithm is as follows:
[0055] L(θ) = (Q Target - Q(s, a; θ)) 2 (7)
[0056] The update formula of the obstacle avoidance method in the present invention is as follows:
[0057] Q π (s, a) = A π (s, a) + V π(s) (8)
[0058]
[0059] The loss function L(θ) of the obstacle avoidance method in the present invention is as follows:
[0060] L(θ) = (Q Target (s,a;θ t ,ε)-Q(s,a;θ,ε)) 2 (10)
[0061] In view of the fact that the traditional algorithm has too strong randomness and the training effect is sometimes good and sometimes bad, the database that directly stores random sampling is changed to use n-step replay when storing, and its formula is as follows:
[0062]
[0063] In the formula, V(s) represents the current information, R t+1 is the information of the next step, and γ is the information discount rate. After storage, update the priority value in the corresponding set tree of the database, and select according to this priority during sampling. The simplified structural diagram of the improved algorithm process is as Figure 4 shown.
[0064] Please refer to Figure 1 , the present invention provides an obstacle avoidance method for an unmanned ship, including the following steps:
[0065] Step S1: Establish a two-dimensional environment as Figure 2 shown. The entire environment includes an unmanned ship with a blue circular warning range and red collision monitoring points set, the blue trajectory of the unmanned ship, three static obstacles at different positions, a dynamic obstacle moving from the upper right to the lower left along the route of the unmanned ship, and a green circular target point in the upper left corner.
[0066] Step S2: Establish an environment model based on the Markov process framework, including the state space s t , action space a, reward r, and termination flag d. Establish a state space including the speed and angle of the unmanned ship, and an action space including forward movement, left and right changes in the sailing angle, and speed changes, etc.
[0067] The reward setting is as Figure 3 shown, and the formula expression is:
[0068]
[0069] In the formula, ship_length is the length of the unmanned ship, r is the total reward during the movement of the unmanned ship, r 1 is the reward generated by the change in the distance between the unmanned ship and the target point before reaching the target point, dist goal is the distance between the unmanned ship and the target point, distprev is the distance between the unmanned ship and the target point in the previous state; r 2 is the reward for the unmanned ship moving towards the target point; r 3 is the reward obtained by the unmanned ship reaching the target point; r 4 is the penalty when the unmanned ship collides with an obstacle. DCPA is the minimum distance between the unmanned ship and the obstacle, dist warn is the size of the warning range of the unmanned ship; r 5 is the penalty when an obstacle appears within the warning range of the unmanned ship.
[0070] Step S3: After the unmanned ship selects an action based on the set random seed and starts moving, it interacts with the environment after executing the action to generate a new state and calculates the obtained reward. The state s t , action a, reward r, next state s t+1 and termination flag d are composed of a five-tuple and stored in the database for sampling training. The specific storage method is to calculate the priority of the current five-tuple according to the temporal difference algorithm, that is:
[0071] Priority = TD error (Q target -Q value ) (13)
[0072] In the formula, TD error is the temporal difference algorithm, Q target is the target value network, Q value is the current value network.
[0073] The five-tuple expressing the state of the unmanned ship is stored in the database, and the priority is stored in the corresponding set tree of the database.
[0074] Step S4: Sample the environment model and update the network, and assign values to the sample state and action. When the number of movements of the unmanned ship reaches the minimum learning batch of the algorithm database, sample the environment model and update the network. The update process is divided into several steps:
[0075] Step 4.1, Initialize the current network and the target network with noise ε added to the nodes;
[0076] Step 4.2, Select action a. The action selection formula is as follows:
[0077]
[0078] In the formula, a best is the best action, represents the most valuable action a selected by the value maximum function under the state s and the policy π with network parameters θ;
[0079] Step 4.3, update the state s of the unmanned ship according to the action a in Step 4.2 t+1 , calculate the state value V(s) and the action value A(s,a) separately and then sum them to obtain Q(s,a). The formula is expressed as:
[0080]
[0081] In the formula, q π (s,a) is the value under the policy π(a|s), and a′ is other actions in the state s.
[0082] Calculate the reward return r. The formula is as follows:
[0083]
[0084] In the formula, r j is the value obtained after taking the action a by the current value network, γ is the discount rate of the target value network, d = False and d = True are the determinations of whether the termination condition is met or not;
[0085] Step 4.4, store (s t ,a,r,s t+1 ,d) in the database after n-step replay. The n-step replay formula is shown in formula (7). If the database is not full, store it directly; otherwise, delete the oldest data and then store it. Update the priority value in the corresponding set tree of the database;
[0086]
[0087] In the formula, t is the current step, n is the n-step replay, k is the intermediate step between t and n, and G t:t+n is the value after n-step replay starting from t.
[0088] Step 4.5, if the minimum learning batch is reached, sample and learn from the database according to the n-step replay priority and update the current network. The update formula is as follows:
[0089] Q new (s t ,a t ) = Q old (s t ,a t ) + α(G t:t+n - Q old (s t ,a t )) (19)
[0090] In the formula, Q new is the new network, Q old is the old network, s t is the state at the t-th step, a tis the action at the t-th step.
[0091] Step 4.6: Every time n steps are reached, update the target network; if not reached, return to Step 4.2. The update formula is as follows:
[0092]
[0093] Step S5: Select the optimal action based on the values of each action in the state of the sample, and then obtain the optimal strategy; if the maximum number of steps is not reached or the target point is not reached, return to Step S4.
[0094] Preferably, in this embodiment, it further includes Step S6. When the maximum number of steps is reached or the target point is reached, substitute the change in the state speed direction of this round into the kinematic formula:
[0095]
[0096] In the formula, is the positive definite symmetric inertia matrix of the added mass, is the Coriolis force and centripetal force matrix, is the linear damping matrix, is the kinematic force matrix, τ 1 is the longitudinal force, τ 2 is the lateral force, τ 3 is the yaw moment.
[0097] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention.
Claims
1. An obstacle avoidance method for an unmanned ship, characterized in that, it includes the following steps: Step S1: Construct a two-dimensional environment, generate static and dynamic obstacles, target points, and the starting point of the unmanned ship; Step S2: Establish an environmental model based on the Markov process framework, including the state space s t , the action space a, the reward r, and the termination flag d; Step S3: The unmanned ship starts to move after selecting an action based on the set random seed. After executing the action, it interacts with the environment to generate a new state and calculates the reward. The state s t , action a, reward r, next state s t+1 and termination flag d are composed into a five-tuple and stored in the database for learning for sampling training; Step S4: Sample the environmental model for network update, and assign values to the sample states and actions; Step S5: Select the optimal action according to the value of each action in the sample state, and then obtain the optimal strategy; if the maximum number of steps is not reached or the target point is not reached, return to Step S4; The formula expression of the reward r is: In the formula, ship_length is the length of the unmanned ship, r is the total reward during the movement of the unmanned ship, r 1 is the reward generated by the change in the distance between the unmanned ship and the target point before the unmanned ship reaches the target point, dist goal is the distance between the unmanned ship and the target point, dist prev is the distance between the unmanned ship and the target point in the previous state; r 2 is the reward for the unmanned ship moving towards the target point; r 3 is the reward obtained by the unmanned ship reaching the target point; r 4 is the penalty when the unmanned ship collides with an obstacle, DCPA is the minimum distance between the unmanned ship and the obstacle, dist warn is the size of the warning range of the unmanned ship; r 5 is the penalty when an obstacle appears within the warning range of the unmanned ship.
2. The obstacle avoidance method for an unmanned ship according to claim 1, characterized in that, the two-dimensional environment includes an unmanned ship with a warning range and collision monitoring points set, the trajectory of the unmanned ship, static obstacles at different positions, a dynamic obstacle moving from the upper right to the lower left along the unmanned ship route, and a circular target point in the upper left corner.
3. The obstacle avoidance method for an unmanned ship according to claim 1, characterized in that, when storing in Step S3, calculate the priority of the current five-tuple according to the temporal difference algorithm, that is: Priority = TD error (Q target -Q value ) Wherein, TD error is the temporal difference algorithm, Q target is the target value network, and Q value is the current value network; Store the five-tuple expressing the state of the unmanned ship in the database, and store the priority in the corresponding set tree of the database.
4. The obstacle avoidance method for an unmanned ship according to claim 1, characterized in that, the network update is specifically: (1), Initialize the current network and the target network with nodes added with noise ε; (2), Select an action a, and the action selection formula is as follows: where a best is the optimal action, represents the most valuable action a selected by the value maximization function under the policy π with state s and network parameters θ; (3) Update the state s of the unmanned ship according to the action a in (2). t+1 Separate the calculation of the state value V(s) and the action value A(s,a), and then sum them to obtain Q(s,a). The formula is expressed as: where q π (s,a) is the value under policy π(a|s), and a′ is other actions in state s; Calculate the reward return r, and the formula is as follows: where r j is the value obtained after the current value network takes action a, γ is the discount rate of the target value network, and d = False and d = True are the determinations of whether the termination condition is met or not, respectively; (4), store (s t , a, r, s t+1 , d) into the database after n-step replay. If the database is not full, store it directly; otherwise, delete the oldest data and then store it; update the priority value in the corresponding set tree of the database; Where \(t\) is the current step, \(n\) is the replay of \(n\) steps, \(k\) is the intermediate step between \(t\) and \(n\), and \(G\) t:t+n is the value after replaying \(n\) steps starting from \(t\); (5), If the minimum learning batch is reached, sample and learn from the database according to the priority of the n-step replay and update the current network, and the update formula is as follows: Q new (s t ,a t ) = Q old (s t ,a t ) + α(G t:t+n - Q old (s t ,a t )) Where, Q new is the new network, Q old is the old network, s t is the state at the t-th step, a t is the action at the t-th step; (6), Every time n steps are reached, update the target network, if not reached, return to (2), and the update formula is as follows:
5. The obstacle avoidance method for an unmanned ship according to claim 1, characterized in that, it further includes Step S6. When the maximum number of steps is reached or the target point is reached, substitute the change in the state speed direction of this round into the kinematic formula: wherein, is a positive definite symmetric inertia matrix of the added mass, is the Coriolis force and centripetal force matrix, is the linear damping matrix, is the kinematic force matrix, τ 1 is the longitudinal force, τ 2 is the lateral force, τ 3 is the yaw moment; Calculate T to obtain τ in the round 1 τ 2 τ 3 Analyze the curve and end this round; if the maximum total number of steps limit is reached, end the training, otherwise return to step S3.