An informative path planning method for AUV based on learning and sampling
By combining the Q-learning algorithm and probability roadmap technology, the reward matrix is designed to optimize AUV path planning, which solves the problem of difficulty in comprehensively considering the impact of ocean currents, information collection and static obstacle avoidance in the existing technology, and maximizes the safety, energy efficiency and information value of AUV path planning, and has the automatic return function.
Patent Information
- Application Number
- CN202211381884.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-11-02
AI Technical Summary
The existing AUV path planning methods are difficult to comprehensively consider the problems of current impact, information collection and static obstacle avoidance, making it difficult to plan a path to maximize safety, energy efficiency and information value in complex marine environments.
A method of informational AUV path planning based on learning and sampling is proposed. Combined with Q-learning algorithm and probability roadmap technology, the reward matrix is designed to optimize path planning, ensure the safety, energy efficiency and information value of the path, and at the same time realize the automatic return function of AUV.
This method can effectively reduce the energy consumption of AUV, maximize the value of sampling information, realize safe path planning, and have automatic return function, which is suitable for multi-objective optimization problems in complex marine environments.
Smart Images

Figure CN115686031B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of path planning for autonomous underwater vehicles, and in particular relates to an informative path planning method for autonomous underwater vehicles based on learning and sampling taking into account the influence of ocean currents. Background Art
[0002] Informative path planning for autonomous underwater vehicles (AUVs) considering the influence of ocean currents,Informative path planning (IPP) refers to planning a path for AUV that maximizes the value of sampling information under the condition of satisfying budget constraints.At the same time, in the marine environment, the influence of ocean currents on the movement of AUVs should be considered.
[0003] Many path planning methods that consider the influence of ocean currents have been proposed to determine the optimal path with the lowest energy consumption or the shortest navigation time. Graph-based search methods are widely used in path planning problems in marine environments, such as the Dijkstra algorithm and the A* algorithm. These algorithms use discrete representations of the environment and have good consistency and convergence, but are difficult to solve in a time-varying ocean current field environment. The level set method is an effective method for underwater vehicle path planning in a flow field that changes dynamically over time, but there may also be computational problems. The potential field method is a method that artificially generates repulsive and gravitational fields on obstacles and targets to calculate the safe path of underwater vehicles. This method can effectively solve the obstacle avoidance problem, but it often produces local optimal solutions and cannot consider all ocean current variables. The particle swarm optimization method has also been applied to path planning research to obtain the global optimal path, but it has problems such as premature convergence. In complex marine environments, path planning sometimes needs to consider multi-objective optimization problems. For example, in the optimization problem of mixed integer nonlinear programming, a spatiotemporal network model is proposed, which can reduce energy consumption while minimizing the cycle time and achieve collision avoidance.
[0004] The most general form of the IPP problem is to find a path that maximizes information gain under a set of constraints. A variety of solutions have been proposed. For example, an intelligent generation of rapid exploration random loops is used to plan informative paths for sensor robots to achieve periodic continuous environmental monitoring and generate the best estimate of the Gaussian random field distribution. Submodular reward functions and partially observable Markov decision processes have been proposed for the IPP problem, but these methods are usually difficult to apply to large-scale problem instances. The IPP problem that considers both budget constraints and information gain can be formulated using a combinatorial optimization problem. Specific solutions include recursive greedy algorithms and branch-and-bound techniques, but the size of the IPP problem usually grows exponentially with the increase of the budget and the expansion of the search space. The IPP problem is closely related to adaptive sampling, in which the goal is to visit the observation location with the minimum prediction uncertainty or the maximum information gain. A rapid exploration information acquisition algorithm based on iterative adaptive sampling has been proposed to solve the IPP problem. In order to calculate the optimal solution to the IPP problem, one method is to select a set of "informative" sensing locations and then traverse these locations at the lowest cost; another method is to search and plan within a limited range, which can be combined with sampling-based methods to achieve asymptotic optimality.
[0005] Sampling-based methods are also widely used in AUV path planning due to their advantages in dealing with high-dimensional problems. For example, probabilistic roadmaps and rapidly explored random trees (RRT) have been proposed. These algorithms can asymptotically approach the optimal solution within a reasonable computing time. By using the branch-and-bound technique to prune the search tree, a variant of the RRT algorithm, the RRT* algorithm, is proposed, which has real-time capabilities. Sampling-based methods are one of the best ways to generate optimal sampling paths for AUV information collection tasks. However, most of the existing sampling-based path planning methods minimize the path cost without considering the budget constraint, so they cannot be directly applied to the IPP problem. Summary of the invention
[0006] In order to solve the problem that the existing AUV path planning research does not comprehensively consider the influence of ocean currents, information collection and static obstacle avoidance, the present invention proposes an AUV information path planning method based on learning and sampling. The method proposed in the present invention expands the sampling-based path planning technology to improve the solution to the IPP problem, and uses the probabilistic roadmap to improve the efficiency of the IPP algorithm, which not only minimizes the energy consumption of the AUV, but also considers the maximization of the value of the AUV sampling information, while achieving obstacle avoidance and planning a safe and optimal path for the AUV. This method can not only solve the multi-objective optimization problem, but also has the advantages of high computational efficiency and can realize the function of automatic return of the AUV.
[0007] The technical solution adopted by the present invention to solve its technical problem is:
[0008] The present invention comprises the following specific steps:
[0009] Step 1: Use Q-learning for AUV path planning:
[0010] Step 1.1, AUV is in state s t Execute action a t+1 , and receive real-time reward value r t+1 =R(s t ,a t+1 ), where R is the reward matrix. The reward matrix R has states S as rows and actions A as columns. R(s i ,a j )(i,j1,2,…,N) represents the transition from the current state s i Execute action a j Reach the next state s j The reward value obtained after . The reward matrix R is as follows:
[0011]
[0012] When two states cannot be transferred, the corresponding matrix elements are set to -1. When two states can be transferred, if state s j If it is the target state, set the matrix element to 10, otherwise it is set to 0.
[0013] Step 1.2, through the process of learning and updating the Q-table to store Q values, the AUV can learn a target strategy π:S→A, which maps the state set S to the action set A, and the AUV will choose a series of actions from the current state to the target state. The optimal target strategy π * It can guide the AUV to choose the action that maximizes the expected Q value of the cumulative reward, so that the AUV can safely reach the target state in the most energy-efficient way.
[0014] For the AUV path planning problem, the state space S is the set of all possible positions of the AUV, and the action space A is the set of all possible movements of the AUV. The Q value is the AUV at a certain time t, at position s t (s t ∈S) takes an action a t (a t ∈A) The expectation of the future cumulative reward of moving to another location is defined as:
[0015]
[0016] Where π is the target policy, represents the expected operation, r i(i=t+1, t+2,…, t+m) represents the reward value obtained by the AUV at the future time i. t =r t+1 +γr t+2 +γ 2 r t+3 +…+γ m-1 r t+m represents the cumulative discounted reward value at the current time t in the future m moments. The reward at the future moment is multiplied by the discount coefficient γ,γ 2 ,…,γ m-1 Reflected at the current moment. The discount coefficient γ∈[0,1) indicates the degree of foresight of the AUV. The closer γ is to 1, the more foresighted the AUV is, that is, the more the AUV will consider the impact of its action choices on the future.
[0017] Step 1.3, use the temporal difference method to learn the target policy π. The learning and updating process of the cumulative reward expected Q value in the Q-table is:
[0018]
[0019] Where α is the learning rate and s′ is the next state reached after executing action a in state s. The value of the value function Q(s,a) represents the quality of the target strategy π for selecting action a in state s.
[0020] For the entire learning process, first initialize the Q-table to an all-zero matrix of the same size as the reward matrix R, and then use formula (2) to iteratively update the Q-table. In each iteration, an action a is selected from the randomly selected initial state according to the behavior strategy. After executing action a, the next state s' is obtained. When the value of R(s,a) is -1, a new iteration is performed. Otherwise, use formula (2) to update the value of Q(s,a) and reach state s'. Repeat the above process until the target state is reached and this iteration terminates. If the convergence condition of the Q-table is met, the entire learning process ends.
[0021] Through the learning and updating process of the cumulative reward expected Q value in the Q-table, a converged Q * , and learn the optimal target strategy π for the AUV * From formula (2), we can see that the convergence condition of Q-table is that for each state s and action a:
[0022] R(s,a)+γmax a ,Q(s',a')=Q(s,a) (3)
[0023] Or it can be expressed as:
[0024] |R(s,a)+γmax a ,Q(s',a')-Q(s,a)|<δ (4)
[0025] Where δ is a very small positive constant. When the conditions in formula (4) are met, the Q-table can be considered to be convergent. At this time, based on the convergent Q-table, that is, Q * , the optimal target strategy π * It is expressed as:
[0026]
[0027] Using the optimal target strategy π * Select actions in sequence to realize the path planning of the AUV from the starting state to the target state. The obtained state sequence corresponds to the position of the AUV in space. Since the reward matrix is specially designed for the AUV to reach the target position, the AUV is based on π * The selected action will eventually achieve the planning goal of the shortest path. The optimal path P composed of the obtained state sequence * It is expressed as:
[0028]
[0029] in, represents the optimal path P * The path points on the , n is the number of path points, Indicates that from the path point To waypoint subpath segment of .
[0030] Step 2: AUV informative path planning method based on learning and sampling
[0031] Step 2.1, Q-learning hybrid path planning method based on probabilistic roadmap
[0032] The probabilistic roadmap method mainly includes two stages: the graph construction stage and the graph search stage.
[0033] In the graph construction phase, a roadmap is constructed to represent the working environment around the AUV. First, the environment is initialized as an empty undirected graph G(S,A), where the vertex set S represents a set of collision-free AUV position nodes, that is, the state space in Q-learning. The edge set A represents the set of collision-free paths, that is, the action space in Q-learning. Secondly, the roadmap is constructed using the uniform random sampling (URS) method and the K nearest neighbor (KNN) algorithm. Using the URS method, collision-free nodes s are sampled in free space. i(i1,2,…,N) and add it to the vertex set S. Then, use the KNN algorithm to search for s i k neighbor nodes of node s i Each node is connected to its k neighbor nodes to generate lines to build a roadmap. At the same time, check whether the line collides with any obstacle, add the non-collision line to the edge set A, otherwise delete the line. Finally, a low-dimensional collision-free probability roadmap is constructed.
[0034] In the graph search phase, the Q-learning algorithm is integrated with the generated probabilistic roadmap. The probabilistic roadmap is used as the input of the Q-learning algorithm to construct the reward matrix R and Q-table. The set of randomly sampled nodes in the probabilistic roadmap is set as the state space in Q-learning.
[0035] Step 2.2: Hybrid path planning method to solve the IPP problem in ocean current field
[0036] The AUV needs to follow the forward path P f Sampling environmental information, while considering the impact of ocean currents on its energy consumption. For the constructed probabilistic roadmap, first get an initial reward matrix at this time, There are only three element values, namely -1, 0 and 10. Then, the reward matrix is adjusted using the known flow field and environmental information data. Consider the state s h The sampling information value at And by the state s l Transfer to state s h Energy consumption The non-negative values in the initial reward matrix Redesigned to:
[0037]
[0038] Where ρ and ω are positive constant weight coefficients. Energy consumption Calculate by the following formula (8):
[0039]
[0040] Among them, P v is the propulsion power of the AUV, which is related to the propulsion speed of the AUV is proportional to the cube of i is the AUV along the subpath segment The time it takes to travel, k is the drag coefficient of the AUV, which is determined by the design of the AUV itself, and the path point p i Corresponding to state s l, path point p i+1 Corresponding to state s h , is the AUV in the sub-path segment The speed relative to the seabed when traveling on the surface can be determined by and ocean current speed The vector synthesis is obtained, that is:
[0041]
[0042] In formula (7), there is a combined reward r ie =ρr i -ωr e . Due to r i and r e The units and magnitudes are different. i and r e Perform dimensionless processing, that is, normalization. Use the Min-Max normalization method to make r i and r e The value of is in the range [0,1]. At the same time, in order to avoid r ie Negative reward values appear in r ie It is also normalized to be in the range [0,1].
[0043] Due to r ie The range is [0,1], which is smaller than the initial reward matrix The value of 10 at the target position in the middle does not affect the purpose of the AUV approaching the target area. By properly designing the values of ρ and ω, a reasonable trade-off between information collection and energy consumption can be achieved. The redesigned reward matrix as follows:
[0044]
[0045] The AUV moves along the forward path P f During the sampling process, if the energy reserve is insufficient, it is necessary to return along the return path P in the most energy-efficient way. r Return to the starting point. Since on the return path P r The AUV does not perform sampling, and only considers the impact of ocean currents on the energy consumption of the AUV. Therefore, the design of the reward matrix is different. From formula (8), it can be seen that when the propulsion speed of the AUV is constant, the energy consumption of the AUV is proportional to the navigation time. Therefore, the value of the reward matrix is redesigned using the inverse of the AUV navigation time. The shorter the navigation time, the less energy consumption, and the higher the reward value obtained by the AUV. Establish the initial reward matrix as follows:
[0046]
[0047] The AUV needs to return to the starting point, so the original starting state s1 is set as the target state, and the reward matrix The value of the corresponding position in is 10. Then, based on the known ocean current field data, as well as the propulsion speed and direction of the AUV, the navigation time of the AUV is calculated as:
[0048]
[0049] Among them, Δt i,j is the AUV position from state s i (The two-dimensional coordinates of space are [x i ,y i ]) to state position s j (The two-dimensional coordinates of space are [x j ,y j ]) The time it takes to sail, l cell is the length of the unit grid in the environment space, is the speed of the AUV relative to the seabed, i.e. the propulsion speed, which can be calculated by formula (9). The reward matrix obtained after redesign as follows:
[0050]
[0051] After systematically designing the reward matrix, we can use the redesigned and According to formula (2), the Q-table is learned and updated until it converges, and Q f -table and Q r -table, respectively expressed in matrix form Q f (s,a) and Q r (s,a). At the same time, the AUV learns the optimal target strategy and for:
[0052]
[0053]
[0054] according to The optimal forward path of the AUV can be obtained Implement IPP tasks:
[0055]
[0056] according to The optimal return path of the AUV can be obtained for:
[0057]
[0058] Step 2.3: Hybrid path planning method realizes the automatic return function of AUV
[0059] The automatic return function is designed as follows:
[0060] The AUV follows the optimal forward path At each step of the journey, the remaining energy E of the AUV at the current position p is calculated using formula (15) according to the path the AUV has traveled. r :
[0061]
[0062] Among them, e i For subpath segments energy consumption.
[0063] Find the best path forward The next path point p′ on the path. Using the learned Q r -table plans the optimal return path from p′ to the starting point according to and Calculate the minimum energy consumption E of the AUV from the current position p to the next path point p′ and from the next path point p′ back to the starting point m .
[0064] E r With E m Compare to determine whether the AUV’s energy reserve is sufficient. r ≥E m , then the energy is sufficient, and the AUV goes to the next path point p′ to continue sampling. At this time, the current position of the AUV becomes p′. Otherwise, the AUV stops sampling and starts from the converged Q r -Find the return path P with the lowest energy consumption from the current position p back to the starting point in the table r At this point, the path that the AUV has traveled from the starting point to the current point is the final forward path P f .
[0065] Connect P f and P r The final planned closed round-trip trajectory P is formed. In the above process, only one learning is required to obtain the converged Q r -table, from which the return path can be easily searched, providing convenience for realizing the automatic return function.
[0066] The beneficial effects of the present invention are:
[0067] The proposed Q-learning hybrid path planning method based on probabilistic roadmap can be used for AUV to efficiently sample environmental information. First, the probabilistic roadmap is combined with the Q-learning process to reduce the dimension of the problem to be solved and reduce the computational burden. Then, the reward matrix in Q-learning is systematically designed for the sampling problem in the ocean current field and temperature field, that is, the information path planning problem. In addition, considering the limited energy reserve of AUV, the automatic return function of AUV is designed. The proposed algorithm only needs one learning to achieve the characteristic of multiple planning, which provides great convenience for the realization of this function. The present invention has great application potential on AUVs with limited computing resources and decision-making time. The present invention can not only solve the multi-objective optimization problem, but also has the advantage of high computational efficiency, which minimizes the energy consumption of AUV, takes into account the maximization of the value of AUV sampling information, and realizes obstacle avoidance at the same time, plans a safe and optimal path for AUV, and can realize the function of automatic return of AUV.
[0068] In addition, the hybrid path planning algorithm proposed in the present invention is not only limited to the scenarios considered in the present invention, but is also applicable to other general information path planning problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a schematic diagram showing the environment space where the AUV is located and the path form;
[0070] Figure 2 Schematic diagram of the construction process of the probability roadmap;
[0071] Figure 3 Schematic diagram of the constructed probability roadmap;
[0072] Figure 4 For E o Schematic diagram of the path planning results for 2500J;
[0073] Figure 5 This is a schematic diagram of Monte Carlo simulation results;
[0074] Figure 6(a)E o 2200J scene;
[0075] Figure 6(b)E o 2000J scene;
[0076] Figure 7 For E o Schematic diagram of the path planning results for 500J. DETAILED DESCRIPTION
[0077] The present invention proposes an AUV information path planning method based on learning and sampling, taking into account the existence of obstacles, the influence of ocean currents on the movement of AUVs and the sampling of environmental information. The present invention solves the problem of planning a safe path for AUVs that meets the budget constraint and maximizes the sampling information in an environment where an ocean current field exists. The present invention uses the Q-learning algorithm in reinforcement learning, uses the prior knowledge of environmental information and ocean current fields to design a reward matrix, learns and updates the Q-table based on the reward matrix until convergence, and constructs an information-rich optimal safe path for the AUV by traversing the converged Q-table. In addition, in order to improve the efficiency of the Q-learning method, a probabilistic roadmap is combined with it to reduce the dimensions of the state space and Q-table in the Q-learning method. The path between any two position nodes can be searched in the Q-table that has been completely learned to converge. Using this feature, the present invention designs a function for the AUV to automatically return when its energy is insufficient.
[0078] The technical solution adopted by the present invention to solve its technical problem is:
[0079] The present invention comprises the following specific steps:
[0080] Step 1: Use Q-learning for AUV path planning:
[0081] Q-learning is a value-based reinforcement learning method that implicitly learns the target policy π by explicitly learning a value function Q(s,a). The main idea of the Q-learning algorithm is to construct a state set S and an action set A, and then establish a Q-table that stores Q values to evaluate the quality of the selected actions. Preferred actions will be rewarded, otherwise they will be punished. Specifically, the AUV is in state s t Execute action a t+1 , and receive real-time reward value r t+1 =R(s t ,a t+1 ), where R is the reward matrix. The reward matrix R has states S as rows and actions A as columns. R(s i ,a j )(i,j=1,2,…,N) represents the transition from the current state s i Execute action a j Reaching state s j The reward value obtained after the transition. When two states cannot be transferred, the corresponding matrix element is set to -1. When two states can be transferred, if the state s j If is the target state, set the matrix element to 10, otherwise set it to 0. The reward matrix R is as follows:
[0082]
[0083] Through the process of learning and updating the Q-table, the AUV can learn a target strategy π:S→A, which maps the state set S to the action set A, and the AUV will choose a series of actions from the current state to the target state. The optimal target strategy π * It can guide the AUV to choose the action that maximizes the expected value of the cumulative reward Q. According to the learned optimal strategy π * , the AUV can safely reach the target state in the most energy-efficient way.
[0084] For the AUV path planning problem, the state space S is the set of all possible positions of the AUV, and the action space A is the set of all possible movements of the AUV. The Q value is the AUV at a certain time t, at position s t (s t ∈S) takes an action a t (a t ∈A) The expectation of the future cumulative reward of moving to another location is defined as:
[0085]
[0086] Where π is the target policy, represents the expected operation, r i (i=t+1, t+2,…, t+m) represents the reward value obtained by the AUV at the future time i. t =r t+1 +γr t+2 +γ 2 r t+3 +…γ m-1 r t+m represents the cumulative discounted reward value at the current time t in the future m moments. The reward at the future moment is multiplied by the discount coefficient γ,γ 2 ,…,γ m-1 Reflected at the current moment. The discount coefficient γ∈[0,1) indicates the degree of foresight of the AUV. The closer γ is to 1, the more foresighted the AUV is, that is, the more the AUV will consider the impact of its action choices on the future.
[0087] Use the temporal difference method to learn the target policy π. The learning and updating process of the cumulative reward expected Q value in the Q-table is:
[0088]
[0089] Where α is the learning rate, s′ is the next state reached after executing action a in state s. The value of Q(s,a) represents the quality of the target policy π that selects action a in state s.
[0090] For the entire learning process, first initialize the Q-table to an all-zero matrix of the same size as the reward matrix R, and then use formula (2) to iteratively update the Q-table. In each iteration, an action a is selected from the randomly selected initial state according to the behavior strategy (for example, ∈greedy, that is, randomly selecting an action with a probability of ∈, and selecting an action according to the target strategy π with a probability of (1-∈). After executing action a, the next state s' is obtained. When the value of R(s,a) is -1, a new iteration is performed. Otherwise, use formula (2) to update the value of Q(s,a) and reach state s'. Repeat the above process until the target state is reached and the iteration is terminated. If the convergence condition of the Q-table is met, the entire learning process ends.
[0091] The learning and updating of Q-table is the core of Q-learning algorithm. Through this process, the converged Q * , and learn the optimal target strategy π for the AUV * From formula (2), we can see that the convergence condition of Q-table is that for each state s and action a:
[0092] R(s,a)+γmax a ,Q(s',a')=Q(s,a) (3)
[0093] Or it can be expressed as:
[0094] |R(s,a)+γmax a ,Q(S',a')-Q(s,a)|<δ (4)
[0095] Where δ is a very small positive constant. When the conditions in formula (4) are met, the Q-table can be considered to be convergent. At this time, based on the convergent Q-table, that is, Q * , the optimal target strategy π * It is expressed as:
[0096]
[0097] Using the optimal target strategy π * Select actions in sequence to realize the path planning of the AUV from the starting state to the target state. The obtained state sequence corresponds to the position of the AUV in space. Since the reward matrix is specially designed for the AUV to reach the target position, the AUV is based on π * The selected action will eventually achieve the planning goal of the shortest path. The optimal path P composed of the obtained state sequence * It is expressed as:
[0098]
[0099] in, represents the optimal path P * The path points on the , n is the number of path points, Indicates that from the path point To waypoint The path representation is as follows: Figure 1 shown.
[0100] Step 2: AUV informative path planning method based on learning and sampling
[0101] In the previous step, the principle of Q-learning algorithm to realize the shortest path planning of AUV was introduced. This step will explain in detail the AUV informative path planning method based on learning and sampling. First, in order to improve the computational efficiency, a Q-learning hybrid path planning method based on probabilistic roadmap is proposed. Secondly, in order to solve the IPP problem considering the influence of ocean currents, the reward matrix of Q-learning in the hybrid path planning method is specially designed. Finally, it is introduced how the proposed hybrid path planning method can conveniently realize the automatic return function of AUV.
[0102] Step 2.1, Q-learning hybrid path planning method based on probabilistic roadmap
[0103] The efficiency of Q-table learning and updating in Q-learning is related to the dimensions of the state space and action space. When the state space and action space are large, the convergence speed of the algorithm will be very slow. To address this problem, the present invention introduces a probabilistic roadmap method based on sampling to reduce the spatial dimension of Q-learning. The basic idea of the probabilistic roadmap method is to construct a network graph containing possible paths through random sampling and search for paths from it. The probabilistic roadmap method mainly includes two stages: the graph construction stage and the graph search stage.
[0104] In the graph construction phase, a roadmap is constructed to represent the working environment around the AUV. First, the environment is initialized as an empty undirected graph G(S,A), where the vertex set S represents a set of collision-free AUV position nodes, that is, the state space in Q-learning. The edge set A represents the set of collision-free paths, that is, the action space in Q-learning. Then, the roadmap is constructed using the uniform random sampling (URS) method and the K nearest neighbor (KNN) algorithm. Using the URS method, collision-free nodes s are sampled in free space. i (i=1,2,…,N) and add them to the vertex set S, such as Figure 2 As shown in the leftmost figure. Then use the KNN algorithm to search for s ik neighbor nodes of node s i Connect each of its k neighbor nodes to generate links to build a roadmap. At the same time, check whether the links collide with any obstacles, add the non-collision links to the edge set A, otherwise delete the links, such as Figure 2 As shown in the middle figure. Thus, the constructed low-dimensional collision-free probability roadmap is obtained, as shown in Figure 2 As shown in the rightmost picture.
[0105] In the graph search phase, the Q-learning algorithm is integrated with the generated probabilistic roadmap. The probabilistic roadmap is used as the input of the Q-learning algorithm to construct the reward matrix R and Q-table. The set of randomly sampled nodes in the probabilistic roadmap is set as the state space in Q-learning. More specifically, assuming that the constructed probabilistic roadmap is as follows Figure 3 As shown, the number of randomly sampled nodes is N, and each node s in the graph i (i=1,2,…,N) represents a state, i.e. the position of the AUV. State s1 is the initial state, state s N is the target state. Action a j (j=1,2,…,N) is represented by the arrows on each edge in the figure. It can be seen that the dimension of the state space is greatly reduced compared to the original environment map. Correspondingly, the dimensions of the reward matrix and Q-table are also greatly reduced, thereby improving the efficiency of the algorithm.
[0106] Although the probabilistic roadmap can reduce the number of states in Q-learning, it may lead to incomplete spatial coverage. Since the nodes in the probabilistic roadmap are randomly generated, the path or optimal path cannot be guaranteed to exist in every run. If this happens, a new probabilistic roadmap will be constructed using the resampling method. In addition, the number of sampled nodes can be appropriately increased to improve spatial coverage, but the efficiency of the algorithm will be sacrificed. In practical applications, the relevant algorithm parameters should be reasonably designed.
[0107] Step 2.2: Hybrid path planning method to solve the IPP problem in ocean current field
[0108] Since the reward matrix does not care about the information value of each state position and the ocean current, the hybrid path planning method in step 2.1 cannot be directly applied to the IPP problem in the ocean current field. Inspired by the principle that the Q-learning algorithm selects the path with the highest Q value, the above-mentioned Q-learning hybrid algorithm based on the probabilistic roadmap is further adopted to achieve IPP through the systematic design of the reward matrix R.
[0109] The AUV needs to follow the forward path P fSampling environmental information, while considering the impact of ocean currents on its energy consumption. For a constructed probability route map, such as Figure 3 As shown, first we get an initial reward matrix at this time, There are only three element values, namely -1, 0 and 10. Then, the reward matrix is adjusted using the known flow field and environmental information data. Consider the state s h The sampling information value at And by the state s l Transfer to state s h Energy consumption The non-negative values in the initial reward matrix Redesigned to:
[0110]
[0111] Where ρ and ω are positive constant weight coefficients. Energy consumption Calculate by the following formula (8):
[0112]
[0113] Among them, P v is the propulsion power of the AUV, which is related to the propulsion speed of the AUV is proportional to the cube of i is the AUV along the subpath segment The time it takes to travel, k is the drag coefficient of the AUV, which is determined by the design of the AUV itself, and the path point p i Corresponding to state s l , path point p i+1 Corresponding to state s h , is the AUV in the sub-path segment The speed relative to the seabed when traveling on the surface can be determined by and ocean current speed The vector synthesis is obtained, that is:
[0114]
[0115] In formula (7), there is a combined reward r ie =ρr i -ωr e . Due to r i and r e The units and magnitudes are different. i and r e Perform dimensionless processing, that is, normalization. Use the Min-Max normalization method to make r i and re The value of is in the range [0,1]. At the same time, in order to avoid r ie Negative reward values appear in r ie It is also normalized to be in the range [0,1], that is:
[0116]
[0117] Due to r ie The range is [0,1], which is smaller than the initial reward matrix The value of 10 at the target position in the middle does not affect the purpose of the AUV approaching the target area. By properly designing the values of ρ and ω, a reasonable trade-off between information collection and energy consumption can be achieved. The redesigned reward matrix as follows:
[0118]
[0119] The AUV moves along the forward path P f During the sampling process, if the energy reserve is insufficient, it is necessary to return along the return path P in the most energy-efficient way. r Return to the starting point. Since on the return path P r The AUV does not perform sampling, and only considers the impact of ocean currents on the energy consumption of the AUV. Therefore, the design of the reward matrix is different. From formula (8), it can be seen that when the propulsion speed of the AUV is constant, the energy consumption of the AUV is proportional to the navigation time. Therefore, the value of the reward matrix is redesigned using the inverse of the AUV navigation time. The shorter the navigation time, the less energy consumption, and the higher the reward value obtained by the AUV. Establish the initial reward matrix
[0120]
[0121] The AUV needs to return to the starting point, so the original starting state s1 is set as the target state, and the reward matrix The value of the corresponding position in is 10. Then, based on the known ocean current field data, as well as the propulsion speed and direction of the AUV, the navigation time of the AUV is calculated as:
[0122]
[0123] Among them, Δt i,j is the AUV position from state s i (The two-dimensional coordinates of space are [x i ,y i ]) to state position s j (The two-dimensional coordinates of space are [x j ,y j ]) The time it takes to sail, l cellFor the environment space (such as Figure 1 The length of the unit grid in is the speed of the AUV relative to the seabed, which can be calculated by formula (9). The reward matrix obtained after redesign for:
[0124]
[0125] After systematically designing the reward matrix, we can use the redesigned and According to formula (2), the Q-table is learned and updated until it converges, and Q f -table and Q r -table, respectively expressed in matrix form Q f (s,a) and Q r (s,a). At the same time, the AUV learns the optimal target strategy and for:
[0126]
[0127]
[0128] according to The optimal forward path of the AUV can be obtained Implement IPP tasks:
[0129]
[0130] according to The optimal return path of the AUV can be obtained for:
[0131]
[0132] Step 2.3: Hybrid path planning method realizes the automatic return function of AUV
[0133] Considering the limitation of AUV's own energy reserve, AUV may not be able to reach the target area with the greatest information value. AUV performing sampling mission can sense its own energy reserve and automatically return to the starting point when the energy is insufficient. The learned Q-table provides great convenience for the design of this function, which will be introduced in this step.
[0134] The automatic return function is designed as follows:
[0135] The AUV follows the optimal forward path At each step of the journey, according to the path that the AUV has traveled, the remaining energy E of the AUV at the current position p is calculated using (16): r :
[0136]
[0137] Among them, e i For subpath segments energy consumption.
[0138] Find the best path forward The next path point p′ on the path. Using the learned Q r -table plans the optimal return path from p′ to the starting point according to and Calculate the minimum energy consumption E of the AUV from the current position p to the next path point p′ and from the next path point p′ back to the starting point m .
[0139] E r With E m Compare to determine whether the AUV’s energy reserve is sufficient. r ≥E m , then the energy is sufficient, and the AUV goes to the next path point p′ to continue sampling. At this time, the current position of the AUV becomes p′. Otherwise, the AUV stops sampling and starts from the converged Q r -Find the return path P with the lowest energy consumption from the current position p back to the starting point in the table r At this point, the path that the AUV has traveled from the starting point to the current point is the final forward path P f .
[0140] Connect P f and P r The final planned closed round-trip trajectory P is formed. In the above process, only one learning is required to obtain the converged Q r -table, from which the return path can be easily searched, providing convenience for realizing the automatic return function.
[0141] Step 3, simulation results and discussion
[0142] In this step, simulation results are given to demonstrate the feasibility and effectiveness of the proposed hybrid path planning method. In the simulation experiment, the size of the two-dimensional environment space is 10×10, and the size of each unit grid is 1km1km. Therefore, there are a total of 100 locations in the environment space. Using the URS method, different numbers of nodes are sampled to construct a probabilistic roadmap. Temperature gradient data is used as known sampling information. The starting point coordinates of the AUV are set to [1,1], and the target node is set to the center point of the 3×3 area with the richest temperature gradient information in the entire environment [8,8]. These two points will be added to the collision-free roadmap in the same way as generating a probabilistic roadmap. The two-dimensional simulation environment is as follows: Figure 1 As shown, the dark gray polygons represent the distribution of static obstacles, the asterisks represent the starting positions, the dots represent the target positions, the background is the distribution of temperature gradient information data, and the arrows on the background represent ocean current vectors.
[0143] The relevant parameters are designed as follows: the drag coefficient of AUV k = 3.425, the propulsion speed of AUV is constant at 0.5 m / s; the learning parameters in Q-learning are ∈0.9, γ = 0.8, α = 0.2; the weight coefficients in the reward function are ρ = 1.5, ω = 0.5.
[0144] 3.1 Implementation of IPP in the environment of temperature field and ocean current field
[0145] At the initial energy reserve E o In different cases, the number of sampling nodes was set to N75 and multiple simulations were performed. The results are as follows:
[0146] (1) The AUV has sufficient energy reserve E o 2500J, at which point the AUV can complete the sampling mission and return to the starting point. The forward path and return path planned using the proposed learning and sampling-based path planning method are shown in Figure 2. Figure 4 As shown. The background area with a larger absolute value of the temperature gradient in the figure has a higher information value. When the AUV samples along the forward path, it considers three factors: safe arrival at the target location, information acquisition, and energy consumption. In order to reduce energy consumption, the AUV tries to follow the ocean current and pass through areas with high information value as much as possible. When returning, the AUV only considers energy consumption. In the figure, when the AUV returns from the target point, in order to follow the ocean current, it will travel a short distance to the upper left, but in order to return to the starting point, it will turn downward and close to the starting point. Since the overall trend of the ocean current field from the target point to the starting point is opposite to the direction of travel of the AUV, the AUV cannot use the ocean current when returning, and can only minimize energy consumption as much as possible.
[0147] Since the sampled nodes are randomly generated when constructing the probabilistic roadmap, different random seeds are set and 100 Monte Carlo simulations are performed. The path planning results are as follows: Figure 5 As shown in the figure, it can be seen that sometimes the planned path may not be optimal, which is due to the randomness of sampling in the probabilistic roadmap method. To solve this problem, the random seed can be set manually to select better planning results. In addition, the distribution trend of random nodes can be designed so that they tend to be distributed in areas with higher information value.
[0148] (2) Insufficient energy reserve of AUV E o 2200J or 2000J, at this time the energy reserve of the AUV is not enough to support it to reach the target area, and it cannot complete the entire sampling process. It will return to the starting point halfway. In Figure 6(a), the initial energy reserve of the AUV is 2200J, and when it almost reaches the target area along the optimal sampling path, it will return to the starting point. Otherwise, it will run out of energy and cannot return. The sampling area passed by the AUV is mainly an area rich in information. When returning, the AUV first drives a distance to the lower right to use the ocean current to reduce energy consumption, instead of driving directly to the lower left. Then, the AUV drives to the lower left and returns to the starting point. In Figure 6(b), the AUV has a lower initial energy reserve of 2000J, the forward path becomes shorter, and the AUV returns to the starting point halfway.
[0149] (3) AUV has very little energy reserve, only E o 500J, at this time the AUV cannot depart due to insufficient energy. Figure 7 The AUV’s initial energy is not enough to support it to take the first step, although it can ride the ocean current on the way back. Although in practical applications, the AUV may be able to set off for a short distance, the resolution of the map and the use of probabilistic roadmaps will lead to different decisions.
[0150] 3.2 Comparison between hybrid path planning algorithm and Q-learning
[0151] In this step, the proposed hybrid path planning method is compared with the separate Q-learning algorithm in terms of running time, energy consumption, and information gain to demonstrate the superiority of the proposed algorithm.
[0152] Table 1 compares the running time of the algorithms, including learning time and planning time.
[0153] Table 1
[0154]
[0155] From the comparison of learning time, it can be seen that the hybrid path planning algorithm is more efficient than the Q-learning algorithm in computation. At the same time, reducing the number of randomly sampled nodes N in the probabilistic roadmap can also improve efficiency. However, when the number of sampled nodes is too small, the completeness of the algorithm cannot be guaranteed. From the comparison of planning time, it can be seen that the planning time of each algorithm is roughly the same under different initial energy reserves. This is because the process of Q-learning learning and updating the Q-table takes more time, while the process of searching for a path based on the initial energy takes very little time.
[0156] Table 2 compares the algorithms in terms of energy consumption and information gain obtained.
[0157] Table 2
[0158]
[0159] When the initial energy reserve E o At 5000J, the AUV can complete the sampling task and return to the starting point regardless of which algorithm is used. When using a single Q-learning algorithm, the information obtained by the AUV along the forward path is the richest, although it requires more energy and running time. For the hybrid path planning algorithm, the N75 case is better than the other two cases. Therefore, it is necessary to balance the efficiency and performance of the algorithm by reasonably setting the number of sampling nodes.
[0160] The present invention proposes a Q-learning hybrid path planning method based on a probabilistic roadmap in a temperature field environment with ocean currents, which is used for AUV to efficiently sample environmental information. First, the Q-learning algorithm is used to solve the general AUV path planning problem. In order to reduce the computational burden, the probabilistic roadmap is combined with the Q-learning process to reduce the dimension of the problem. Then, the reward matrix in Q-learning is systematically designed for the sampling problem in the ocean current field and temperature field, that is, the informative path planning problem. In addition, considering the limited energy reserve of AUV, the automatic return function of AUV is designed. The proposed algorithm only needs one learning to achieve the characteristic of multiple planning, which provides great convenience for the realization of this function. Combined with actual marine environmental data, the effectiveness of the proposed hybrid algorithm is verified by simulation. Through the simulation of various scenarios, AUV can well complete the informative path planning task. Compared with the single Q-learning algorithm, the proposed hybrid algorithm has higher computational efficiency. Therefore, the algorithm has great application potential on AUVs with limited computing resources and decision-making time. Moreover, the proposed hybrid algorithm is not only limited to the considered scenario but also applicable to other general informative path planning problems.
Claims
1. A learning and sampling based AUV informative path planning method, characterized by The specific steps include: Step 1: Use Q-learning for AUV path planning: Step 1.1, AUV is in state s t Execute action a t+1 , and receive real-time reward value r t+1 =R(s t , a t+1 ), where R is the reward matrix; the reward matrix R has state S as rows and action A as columns, R(s i , a j ) indicates that from the current state s i Execute action a j Reach the next state s j The reward value obtained after that; where i, j = 1, 2, ..., N; The reward matrix R is as follows: When two states cannot be transferred, the corresponding matrix elements are set to -1. When two states can be transferred, if the state s j If it is the target state, set the matrix element to 10, otherwise set it to 0; Step 1.2, through the process of learning and updating the Q-table to store Q values, the AUV can learn a target strategy π: S→A, which maps the state set S to the action set A. The AUV will choose a series of actions from the current state to the target state accordingly. The optimal target strategy π * It can guide the AUV to choose the action that maximizes the expected Q value of the cumulative reward, so that the AUV can safely reach the target state in the most energy-efficient way; For the AUV path planning problem, the state space S is the set of all AUV positions, and the action space A is the set of all AUV moves; the Q value is the AUV at a certain time t, at position s t (s t ∈S) takes an action a t (a t ∈A) The expectation of the future cumulative reward of moving to another location is defined as: Where π is the target policy, represents the expected operation, r i (i=t+1, t+2, ..., t+m) represents the reward value obtained by the AUV at the future time i; G t =r t+1 +γr t+2 +γ 2 r t+3 +…+γ m-1 r t+m represents the cumulative discounted reward value at the current time t in the future m moments. The reward at the future moment is multiplied by the discount coefficient γ, γ 2 , …, γ m-1 Reflect on the present moment; Step 1.3, use the time difference method to learn the target strategy π; the learning and updating process of the cumulative reward expected Q value in the Q-table is: Where α is the learning rate, s′ is the next state reached after executing action a in state s, a′ is the action performed by s′, and the value of the value function Q(s, a) represents the quality of the target strategy π that selects action a in state s; Through the learning and updating process of the cumulative reward expected Q value in the Q-table, the converged Q * , and learn the optimal target strategy π for the AUV * ; Using the optimal target strategy π * Select actions in sequence to realize the path planning of the AUV from the starting state to the target state. The obtained state sequence corresponds to the position of the AUV in space. AUV according to π * The selected action will eventually achieve the planning goal of the shortest path; the optimal path P composed of the obtained state sequence * It is expressed as: in, represents the optimal path P * The path points on the , n is the number of path points, Indicates that from the path point To waypoint subpath segment of ; Step 2: AUV informative path planning method based on learning and sampling; Step 2.1, Q-learning hybrid path planning method based on probabilistic roadmap; The probabilistic roadmap method consists of two phases: the graph construction phase and the graph search phase; In the graph construction phase, a roadmap is constructed to represent the working environment around the AUV. First, the environment is initialized as an empty undirected graph G(S, A), where the vertex set S represents a set of collision-free AUV position nodes, i.e., the state space in Q-learning; the edge set A represents a set of collision-free paths, i.e., the action space in Q-learning. Secondly, the uniform random sampling (URS) method and K nearest neighbor (KNN) algorithm are used to construct the roadmap. Using the URS method, collision-free nodes s are sampled in free space. i , i = 1, 2, ..., N, and add it to the vertex set S; Then, the KNN algorithm is used to search for s i k neighbor nodes of node s i Connect to its k neighbor nodes respectively, generate lines to build a roadmap; at the same time, check whether the lines collide with any obstacles, add the non-collision lines to the edge set A, otherwise delete the lines; Finally, a constructed low-dimensional collision-free probability roadmap is obtained; In the graph search phase, the Q-learning algorithm is integrated with the generated probabilistic roadmap. The probabilistic roadmap is used as the input of the Q-learning algorithm to construct the reward matrix R and Q-table. The set of randomly sampled nodes in the probabilistic roadmap is set as the state space in Q-learning. Step 2.2, hybrid path planning method to solve the IPP problem in ocean current field; The AUV moves along the path P f Sampling environmental information, while considering the impact of ocean currents on its energy consumption; for the constructed probabilistic roadmap, an initial reward matrix is obtained at this time There are only three element values, namely -1, 0 and 10; then the reward matrix is adjusted using the known flow field and environmental information data Redesign, considering the state s h The sampling information value at and by the state s l Transfer to state s h Energy consumption The non-negative values in the initial reward matrix Redesigned to: Among them, ρ and ω are positive constant weight coefficients, and energy consumption Calculate by the following formula (8): Among them, P v is the propulsion power of the AUV, which is related to the propulsion speed of the AUV is proportional to the cube of i is the AUV along the subpath segment The time it takes to travel, k is the drag coefficient of the AUV, which is determined by the design of the AUV itself, and the path point p i Corresponding to state s l , path point p i+1 Corresponding to state s h , is the AUV in the sub-path segment The speed relative to the seabed when traveling on the surface is and ocean current speed The vector synthesis of is: In formula (7), there is a combined reward r ie =ρr i -ωr e , for r i and r e Perform dimensionless processing, that is, normalize, and use the Min-Max normalization method to make r i and r e The value of is in the range [0, 1], and r ie It is also normalized to be in the range [0, 1]; By properly designing the values of ρ and ω, a reasonable trade-off between information collection and energy consumption is achieved, and the reward matrix is redesigned. as follows: The AUV moves along the forward path P f During the sampling process, if the energy reserve is insufficient, it is necessary to return along the return path P in the most energy-efficient way. r Return to the starting point; use the inverse of the AUV's voyage time to redesign the value of the reward matrix. The shorter the voyage time, the less energy consumption, and the higher the reward value obtained by the AUV. Establish the initial reward matrix as follows: The AUV returns to the starting point, so the original starting state s1 is set as the target state, and the reward matrix The value of the corresponding position in is 10; then, based on the known ocean current field data, as well as the propulsion speed and direction of the AUV, the navigation time of the AUV is calculated as: Among them, Δt i,j is the AUV position from state s i , the two-dimensional coordinates of space are [x i ,y i ], to state position s j , the two-dimensional coordinates of space are [x j ,y j ], the time spent sailing, l cell is the length of the unit grid in the environment space, is the speed of the AUV relative to the seabed, i.e., the propulsion speed, which is calculated by formula (9); the reward matrix obtained after redesign is as follows: After systematically designing the reward matrix, we use the redesigned and According to formula (2), the Q-table is learned and updated until it converges, and Q f -table and Q r -table, respectively expressed in matrix form Q f (s, a) and Q r (s, a); AUV learns the optimal target strategy and for: according to Get the optimal forward path of the AUV Implement IPP tasks: according to Get the optimal return path for the AUV for: Step 2.3: Hybrid path planning method realizes the automatic return function of AUV The AUV follows the optimal forward path At each step of the journey, the remaining energy E of the AUV at the current position p is calculated using formula (15) according to the path the AUV has traveled. r : Among them, e i For subpath segments Energy consumption on Find the best path forward The next path point p′ on the r -table plans the optimal return path from p′ to the starting point according to and Calculate the minimum energy consumption E of the AUV from the current position p to the next path point p′ and from the next path point p′ back to the starting point m ; E r With E m Compare and determine whether the energy reserve of AUV is sufficient; if E r ≥E m , then the energy is sufficient, and the AUV goes to the next path point p′ to continue sampling. At this time, the current position of the AUV becomes p′; otherwise, the AUV stops sampling and starts from the converged Q r -Find the return path P with the lowest energy consumption from the current position p back to the starting point in the table r ; At this point, the path that the AUV has traveled from the starting point to the current point is the final forward path P f ; Connect P f and P r The final planned closed round-trip trajectory P is formed.
2. According to claim 1, a method for AUV information path planning based on learning and sampling, characterized in that: In step 1.2, the discount coefficient γ∈[0,1) indicates the degree of foresight of the AUV. The closer γ is to 1, the more foresighted the AUV is, that is, the more the AUV considers the impact of its action selection on the future.
3. According to claim 1, a method for AUV information path planning based on learning and sampling, characterized in that: In step 1.3, for the entire learning process, the Q-table is initialized to an all-zero matrix of the same size as the reward matrix R, and then the Q-table is iteratively updated using formula (2); In each iteration, an action a is selected from the randomly selected initial state according to the behavior strategy; after executing action a, the next state s' is obtained; When the value of R(s, a) is -1, a new iteration is performed; otherwise, the value of Q(s, a) is updated using formula (2) and the state s' is reached; Repeat the above process until the target state is reached, and this iteration is terminated; if the convergence condition of the Q-table is met, the entire learning process ends.
4. A method for AUV information path planning based on learning and sampling according to claim 1 or 3, characterized in that: In step 1.3, from formula (2), we can see that the convergence condition of Q-table is that for each state s and action a: R(s,a)+γmax a′ Q(s′,a′)=Q(s,a) (3) Or expressed as |R(s,a)+γmax a′ Q(s′,a′)-Q(s,a)|<δ (4) Where δ is a very small positive constant. When the conditions in formula (4) are met, the Q-table is convergent. At this time, based on the convergent Q-table, that is, Q * , the optimal target strategy π * It is expressed as: π * (s)=argmax a Q * (s,a) (5)。 5. The AUV informative path planning method based on learning and sampling according to claim 1, characterized in that: In step 2.1, although the probabilistic roadmap can reduce the number of states in Q-learning, there is a problem of incomplete spatial coverage. If this happens, a new probabilistic roadmap is constructed using the resampling method; In addition, the number of sampling nodes can be appropriately increased to improve the spatial coverage.
Citation Information
Patent Citations
A multi-AUV efficient data collection method based on VOI in an underwater wireless sensor network
CN109275099A
AUV (Autonomous Underwater Vehicle) three-dimensional path planning method based on reinforcement learning
CN109540151A