Path planning method based on improved DQN algorithm
By improving the reward function of the DQN algorithm, introducing prior knowledge and priority experience playback strategy, combined with the attenuation ε-greedy strategy, the problems of slow convergence speed, low sampling efficiency and sparse rewards in path planning are solved, and more efficient and accurate path planning is achieved.
Patent Information
- Application Number
- CN202411889064.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-27
AI Technical Summary
In path planning, DQN algorithm has problems such as slow convergence speed, low sampling efficiency and sparse rewards.
By improving the reward function, introducing prior knowledge and priority experience playback strategies, combined with attenuated ε-greedy strategy, the DQN algorithm is optimized to improve the efficiency and accuracy of path planning.
The improved DQN algorithm significantly improves the convergence speed, enhances data utilization efficiency, and improves the quality of path planning through a more reasonable reward mechanism.
Smart Images

Figure CN120043522A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of path planning, and particularly to a path planning method based on improved DQN. Background Art
[0002] Path planning methods have been one of the research hotspots since the birth of robotics, aiming to select a collision-free path to reach an ideal position. Mobile robot path planning mainly solves three problems: 1) enabling the robot to move from the initial point to the target point; 2) using certain algorithms to enable the robot to avoid obstacles and pass through some necessary points to complete corresponding operation tasks; 3) on the premise of completing the above tasks, optimizing the robot's running trajectory as much as possible. For mobile robots performing tasks such as rescue or transportation, path planning is crucial. Traditional path planning methods, such as the A* or Dijkstra algorithms, require a pre-existing map of the robot's environment. However, nowadays, dynamic path planning has become a popular research topic, enabling mobile robots to dispense with pre-set static conditions. Deep reinforcement learning (DRL) is also a research field that has received much attention, and researchers are using it to solve dynamic path planning problems. The optimal path is a key requirement for robots to accurately and efficiently complete tasks. In the field of autonomous mobile robot path planning, common methods include offline and online ones. In the offline scenario, path planning is based on a prior map of the environment, which is reasonably constructed as a maze, and then search algorithms such as a* and Dijkstra are used for path planning. This method is called the global path planning method. On the contrary, in the online scenario, the robot continuously processes sensor perception data to navigate to the target, and these methods focus on local path planning to ensure the full completion of the obstacle avoidance task. Artificial potential field, convolutional neural network, and reinforcement learning can be regarded as common means for online path planning of mobile robots. In recent years, due to the immunity of reinforcement learning to environmental mutations and dynamic obstacles, it has received extensive attention in the field of online path planning.
[0003] DRL focuses on the performance of software agents in the environment to maximize cumulative rewards. Through trial and error, the agent interacts with the environment and learns the optimal policy to solve sequential decision-making problems in fields such as natural science, social science, and engineering. The sequential decision-making process is as follows: The agent is provided with some initial observations of the environment and is required to select some actions from a given set of possible actions. The environment responds by switching to another state and generating a reward signal (a scalar number), which is regarded as a basic truth estimate of the agent's performance. This process continues, and the agent makes choices based on its observations and the environment and reacts to the next state and reward signal. The only goal of the agent is to maximize the cumulative reward.
[0004] The DQN algorithm is a deep reinforcement learning algorithm. When facing a complex and changing environment, the DQN algorithm has a powerful self-learning and self-improving ability. At the same time, by introducing a convolutional neural network, it solves the dimensionality disaster problem existing in reinforcement learning in complex environments. Therefore, DQN has been widely applied in path planning problems. The DQN algorithm is a combination of deep learning and reinforcement learning. It has the basic characteristics of reinforcement learning: it needs to start learning from scratch in an unknown environment and explore step by step. This means that every path in path planning needs to be explored, and the entire map needs to be incrementally learned from scratch, which greatly increases the computational amount. The entire map needs to be incrementally learned from scratch, which means that DQN has problems such as slow convergence speed, low sampling efficiency, and sparse rewards. Summary of the Invention
[0005] To solve the problems existing in the DQN algorithm in the above content, the present invention proposes a DQN path planning algorithm and system that improves the reward function and action selection strategy and introduces prior knowledge and prioritized experience replay. To solve the problem of sparse rewards, there are usually three solutions, namely reward shaping, intrinsic curiosity module, and curriculum learning. Reward shaping is to design a more reasonable reward function to guide training. The present invention improves the reward function through the reward shaping method. The reward is divided into five parts: distance reward, avoidance reward, arrival reward, obstacle penalty, and stay penalty. To balance the exploration ability and utilization ability of the model, the present invention designs a decaying ε-greedy strategy. In the early stage of model training, the exploration ability of the model is increased, and in the later stage of model training, the utilization ability of the model is increased. To solve the problem of slow convergence speed, prior knowledge is introduced. Before algorithm training, the simulated annealing algorithm is used to find a collision-free path. When initializing the action value function, a reasonable value greater than 0 is assigned to the action values on this path, avoiding the exploration of the system from scratch, reducing the useless exploration due to randomness, reducing the trial and error from scratch, shortening the entire learning time, and significantly improving the convergence speed. The present invention designs a prioritized experience replay strategy. Since each experience has a different enhancement effect on the model, to improve the utilization efficiency of data, TD-Error is used to define the experience priority, and the data that is more helpful to the model is focused on training. TD-Error represents the difference between the model output and the true value. By selectively sampling data with high ED-Error, the model performance is enhanced, and at the same time, the proportion of low-quality data is reduced.
[0006] To achieve the above object, the present invention provides a path planning method based on an improved DQN algorithm. The steps include:
[0007] Construct a two-dimensional grid network for path planning as the environmental map;
[0008] Based on a two-dimensional grid network environment map and the Markov Decision Process (MDP), a DQN path planning model is constructed;
[0009] The DQN path planning model is improved and optimized to obtain the final model;
[0010] Using the final model, path planning is completed.
[0011] In the two-dimensional grid network map constructed in the above steps, each square has the same size and is evenly distributed. The squares are divided into four types, namely the starting square, the feasible square, the obstacle square, and the end square; the starting square and the end square are in a connected state.
[0012] The DQN path planning model constructed based on the two-dimensional grid network environment map and the Markov Decision Process (MDP) essentially combines the Q-learning algorithm and the deep neural network, uses the neural network to replace the Q-value table, solves the dimensionality disaster problem faced by Q-learning, and greatly improves the generalization ability of the model. As the carrier of the state-action value function, the neural network continuously iterates and updates the parameters ω of the neural network q(s,a,ω) by calculating the value function, so as to continuously approximate the state-action value function, that is:
[0013] q(s,a,ω)≈Q(s,a)
[0014] Among them, s represents the state and a represents the action.
[0015] The methods for improving and optimizing the DQN path planning model include:
[0016] By means of reward shaping, a more reasonable reward function is designed to provide more timely feedback to the agent, replacing the traditional reward function in the DQN path planning model;
[0017] A decaying ε-greedy strategy is designed to balance the exploration ability and exploitation ability of the model, replacing the traditional ε-greedy strategy in the DQN path planning model;
[0018] Prior knowledge is introduced, which is different from the traditional DQN path planning model that starts exploration from scratch, improving the convergence speed of the model;
[0019] TD-Error is used to define the experience priority, and a prioritized experience replay is designed to focus on training high-quality data, replacing the traditional experience replay strategy in the DQN path planning model.
[0020] Among the above improvement strategies, a decaying ε-greedy strategy is designed to balance the exploration ability and exploitation ability of the model, specifically including:
[0021] Decaying ε-greedy makes a minor modification to ε-greedy. ε represents the trade-off between exploration and exploitation. Initially, we hope that the value of ε is high so that we can have a high degree of exploration. It can be understood that the agent can learn more new things. As the agent learns about the environment and the future rewards, ε should decay so that we can make full use of the higher Q-values found by the agent. That is, as time goes by, the value of ε continuously decays and becomes smaller and smaller.
[0022] The formula for the traditional ε-greedy method is as follows:
[0023]
[0024] Among them, a represents the action, s represents the state, A(s) represents the number of actions in the action space, is the state-action value function.
[0025] The probability of choosing the greedy action is greater than or equal to the probability of choosing any other action. When ε = 0, it becomes the greedy strategy; when ε = 1, it becomes a uniform distribution, randomly choosing actions from the action space.
[0026] For the decaying ε-greedy method, assume that there are N episodes in total during training. The algorithm initially sets ε = p init (for example, p init = 0.7), and then gradually decreases to ε = p end (for example, p init = 0.1) as the training progresses. Specifically, in the initial stage of training, let the model explore more freely with a high probability, and then gradually decrease at a rate of γ during training, increasing the probability of full exploitation and decreasing the probability of free exploration. The formula is as follows:
[0027]
[0028] ε ← (p init - p end )δ + p end
[0029] In the above improved strategy, by means of reward shaping, a more reasonable reward function is designed to provide more timely feedback to the agent, specifically including:
[0030] The reward is divided into four parts: distance reward, avoidance reward, arrival reward, obstacle penalty, and stay penalty, as shown in the following formula, where target_distance is the distance of the agent at stage S t+1 stage and S tThe absolute value of the distance difference from the target point in the stage. The closer the agent is to the target point, the higher the reward. On the contrary, the farther it deviates from the target point, the greater the punishment. At the same time, when the agent avoids obstacles, it should also be rewarded to encourage the agent to avoid obstacles. obstacle_distance is St +1 stage and S t The absolute value of the distance difference between the agent and the nearest obstacle in the stage. b is the stall penalty to encourage the agent to find a solution faster.
[0031]
[0032] In the above improved strategy, prior knowledge is introduced, which is different from the traditional DQN path planning model that starts exploration from scratch, improving the convergence speed of the model. Specifically, it includes:
[0033] To solve the problem of slow convergence speed, prior knowledge is introduced. Prior knowledge is knowledge prior to the agent's experience. At the same time, it avoids the system's exploration from scratch. Before algorithm training, first use the simulated annealing algorithm to let the agent find a collision-free path from the initial point to the target point in the static environment. Denote this collision-free path as s=(s 1 , s 2 , …… s n ). When initializing the action value function, assign a reasonable value greater than 0 to the action values on this path, avoiding the system's exploration from scratch, reducing useless exploration due to randomness, reducing trial and error from scratch, shortening the entire learning time, and significantly improving the convergence speed.
[0034] In the above improved strategy, TD-Error is used to define the experience priority, and prioritized experience replay is designed to focus on training high-quality data. Specifically, it includes:
[0035] Since each experience has a different enhancement to the model, to improve the utilization efficiency of data, those data that are more helpful to the model can be focused on training. Use TD-Error to define the experience priority. TD-Error represents the difference between the model output and the true value. By selectively sampling data with high TD-Error to enhance the model performance, while reducing the proportion of low-quality data. Calculate the priority priority i and the sampling rate P(i) of the data. The formulas are as follows, where σ i represents the value of TD-Error:
[0036] priority i =|σ i |
[0037]
[0038] Using the SumTree data structure to store the sampling rate of data can effectively save the search time. The SumTree is a binary tree - type data structure, where all leaf nodes store the data sampling rate P(i), and all parent nodes are the sum of their child nodes.
[0039] The path - planning system based on the improved DQN algorithm designed by the present invention includes four modules, namely: a map - building module, a model - building module, an algorithm - optimization module, and a path - planning module.
[0040] The above - mentioned map - building module is used to build a two - dimensional grid network environment map.
[0041] The above - mentioned model - building module is used to build a path - planning model that can be used in the improved DQN algorithm of the present invention.
[0042] The above - mentioned algorithm - optimization module is used to optimize the traditional DQN algorithm, apply the innovation points of the present invention to improve the algorithm, and obtain the final path - planning model.
[0043] The above - mentioned path - planning module is used to use the final path - planning model to complete the path - planning task.
[0044] The advantages of the present invention are as follows:
[0045] Based on the traditional DQN path - planning algorithm, the present invention uses the method of reward shaping to construct a more reasonable reward function to provide more timely feedback for the intelligent agent; designs a decaying ε - greedy strategy to balance the exploration ability and exploitation ability of the model, replacing the traditional ε - greedy strategy in the DQN path - planning model; introduces prior knowledge, different from the traditional DQN path - planning model that starts exploration from 0, to improve the convergence speed of the model; uses TD - Error to define the experience priority, designs a prioritized experience replay, focuses on training high - quality data to improve the model performance, and at the same time reduces the proportion of low - quality data. Brief Description of the Drawings
[0046] To more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings used:
[0047] Figure 1 It is a schematic flow chart of the method of the embodiment of the present invention;
[0048] Figure 2 It is a 10×10 grid network map;
[0049] Figure 3 It is the reinforcement learning process;
[0050] Figure 4 It is a brief structural diagram of the DQN model;
[0051] Figure 5 is a reinforcement learning process with a prioritized experience replay buffer;
[0052] Figure 6 is a schematic diagram of the SumTree data structure; Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0055] Embodiment 1
[0056] As Figure 1 shown, it is a schematic diagram of the method flow of this embodiment, and the steps include:
[0057] step1. Construct a two-dimensional grid network environment map for path planning.
[0058] As Figure 2 shown, it is a 10×10 grid network environment map. Each square has the same size and is evenly distributed. The squares are divided into four types, namely the starting square, the feasible square, the obstacle square, and the ending square; the starting square and the ending square are in a connected state.
[0059] step2. Model using the MDP characteristics.
[0060] The framework of using the reinforcement learning method to solve the path planning problem is to learn how to maximize the benefit of the sequential decision-making problem, which can be represented by the Markov decision process (MDP).
[0061] MDP is a framework that can be composed of a five-tuple (S, A, P, R, γ), where (S, A, R) are the three sets included in the MDP, representing respectively:
[0062] S: Set of states
[0063] A: Set of actions
[0064] R: Set of rewards
[0065] P is the probability distribution included in the MDP, and there are two probability distributions, namely:
[0066] State transition probability: p(s1|s,a) represents the probability of transferring to state s1 when action a is executed in state s. Reward probability: p(r|s,a) represents the probability of obtaining reward r when action a is executed in state s.
[0067] γ ∈ (0, 1): Discount factor, which is used to control the agent to pursue a shorter path to reach the goal in the path planning problem. If γ is smaller, the agent will be more "myopic"; on the contrary, if γ is larger, the agent will be more "far-sighted"; if γ is 1, the agent ignores the time steps and treats the rewards of all time steps equally.
[0068] Policy: Policy Π(a, s) represents the probability of choosing action a in state s.
[0069] MDP also has an important property, namely memoryless property, and the mathematical formula is as follows:
[0070] p(s t+1 |a t+1 ,s t ,...,a 1 ,s 0 ) = p(s t+1 |a t+1 ,s t )
[0071] p(r t+1 |a t+1 ,s t ,...,a 1 ,s 0 ) = p(r t+1 |a t+1 ,s t )
[0072] The purpose of reinforcement learning to solve the path planning problem is to finally find an optimal action sequence in a given environment to maximize the cumulative reward of the agent. For a given action policy π, the cumulative reward is defined, and the formula is as follows:
[0073]
[0074] In the formula, Gt is called the cumulative reward, which represents the sum of discounted rewards from time step t to the end of the action sequence.
[0075] Since the action sequences of the agent in the same state may be different, the cumulative reward of a certain state is an expected value, rather than a definite value. Define a state value function V π(s) quantifies the expected cumulative reward of state s under a given policy π, and the formula is as follows:
[0076] V π (s) = E π [r t+1 |S t = s] + γE[G t+1 |S t = s]
[0077]
[0078] Among them, V π (s) represents the expected cumulative reward of the agent starting from state s under a given policy π. Further generalize it to action a, and change the state value function to the state-action value function q(s t , a t ) to describe the cumulative reward, and the formula is as follows:
[0079] q π (s, a) = E[G t |S t = s, A t = a]
[0080] The algorithm of reinforcement learning needs to obtain all the state-action value functions of a given state, select the maximum value from them, and then obtain the optimal policy. The formula is as follows:
[0081]
[0082] The brief reinforcement learning process is as Figure 3 shown.
[0083] Step 3. Based on the two-dimensional grid network environment map, construct a path planning model for the DQN algorithm.
[0084] The brief structure diagram of the DQN model is as Figure 4As shown, deep neural networks have been introduced into reinforcement learning algorithms and achieved great success in both applications and methods, with neural networks playing a crucial role. Traditional reinforcement learning algorithms rely on constructing and updating Q-tables to achieve policy optimization. However, reinforcement learning based on Q-tables ultimately cannot escape the problem of limited capacity. When the state space and action space are large, traditional reinforcement learning algorithms are no longer feasible. With the birth of DQN, deep learning and reinforcement learning are integrated. Its essence is the combination of the Q-learning algorithm and deep neural networks. By replacing the Q-value table with a neural network, it solves the dimensionality disaster problem encountered by Q-learning and effectively improves the generalization ability of the model. Using a deep convolutional neural network to represent q(s,a) not only solves the problem of limited capacity, but also the deep convolutional neural network has excellent generalization ability, so there is no need to train each state-action value function.
[0085] DQN is an algorithm based on Q-learning, aiming to minimize the objective function as follows:
[0086]
[0087] The specific process is as follows:
[0088] In the two-dimensional grid network environment map, each grid corresponds to a state `s`. The agent starts from the current state s t , selects an action a according to the decaying ε-greedy policy t , and then transfers to a new state s t+1 , and at the same time obtains a reward r t . The state-action sequence (s t , a t , r t , s t+1 ) formed during the interaction between the agent and the environment is regarded as an experience and stored in the experience pool. There are 5 actions during the interaction between the agent and the environment, namely up, down, left, right, and staying still. Among them, the four actions of up, down, left, and right only move one grid distance each time. The settings of the reward function and the decaying ε-greedy are shown in step4.
[0089] During the training process, a batch of samples are drawn from the experience pool as training data, and then the TD-Error is calculated to update the network parameters. The TD-Error is expressed as follows:
[0090]
[0091] Two networks are introduced, namely the main network and the target network If the two networks are separated, the objective function becomes the following formula:
[0092]
[0093] At the beginning, ω T is the same as ω. Fix ω T and keep it unchanged. Then, take out some samples from the replay set and update the parameter ω using the gradient descent algorithm. After iterating a certain number of times, a new ω is obtained, that is, the parameter ω in the main network is updated. Subsequently, assign ω to the parameter ω in the target network T , and then keep ω T unchanged and continue to update ω. Keep iterating. Finally, both ω T and ω can converge to the optimal value.
[0094] Step 4. Apply the innovation point to improve the DQN algorithm model to obtain the final path planning model:
[0095] Step 4.1. For the decaying ε-greedy strategy applied to action selection, the specific implementation is as follows:
[0096] The decaying ε-greedy makes a minor modification based on ε-greedy. ε represents the trade-off between exploration and exploitation. At the beginning, we hope that the value of ε is high, so that we can have a high degree of exploration. It can be understood that the agent can learn more new things. As the agent understands the environment and the future rewards, ε should decay, so that the higher Q-values found by the agent can be fully utilized. That is, as time goes by, the value of ε continuously decays and becomes smaller and smaller.
[0097] The formula for the traditional ε-greedy method is as follows:
[0098]
[0099] where a represents the action, s represents the state, A(s) represents the number of actions in the action space, is the state-action value function.
[0100] The probability of selecting a greedy action is greater than or equal to the probability of selecting any other action. When ε = 0, it becomes a greedy strategy; when ε = 1, it becomes a uniform distribution and randomly selects an action from the action space.
[0101] For the decaying ε-greedy method, assume that there are N episodes in the training. The algorithm initially sets ε = p init (e.g., pinit = 0.7), and then gradually decreases to ε = p during training end (e.g., p init = 0.1). Specifically, in the initial stage of training, let the model explore more freely with a high probability, and then gradually decrease at a rate of γ during training, increase the probability of full utilization, and decrease the probability of free exploration. The formula is as follows:
[0102]
[0103] ε ← (p init - p end )γ + p end
[0104] Step4.2. The improvement of the reward function is as follows:
[0105] The reward is divided into five parts: distance reward, avoidance reward, arrival reward, obstacle penalty, and stay penalty, as shown in the following formula. Among them, target_distance represents the absolute value of the difference in distance between the agent and the target point in the S t+1 phase and the S t phase. The closer the agent is to the target point, the higher the reward obtained; conversely, the farther it deviates from the target point, the greater the penalty. At the same time, when the agent avoids obstacles, it should be given a reward to encourage the agent to avoid obstacles. obstacle_distance is the absolute value of the difference in distance between the agent and the nearest obstacle in the S t+1 phase and the S t phase, and b is the stall penalty, which is used to encourage the agent to find a solution more quickly.
[0106]
[0107] When the agent reaches the target, it gets a reward of +30, and when it collides with an obstacle, it gets a reward of -30. Experiments show that setting k = -0.2, μ = 0.003, and b = -0.005 can promote the rapid convergence of the model.
[0108] Step4.3 The introduction of prior knowledge is specifically implemented as follows:
[0109] Before the algorithm starts training, first use the simulated annealing algorithm to prompt the agent to find a collision-free path from the initial point to the target point in the static environment, and record this collision-free path as s = (s 1 , s 2 , …… s n) When initializing the action value function, assign a reasonable value greater than 0 to the action values on this path to avoid the system's exploration starting from zero, reduce the ineffective exploration caused by randomness, reduce the attempts and errors from scratch, shorten the overall learning time, and significantly improve the convergence speed.
[0110] Step4.4 For the prioritized experience replay strategy, the specific implementation is as follows:
[0111] A brief reinforcement learning process with a prioritized experience replay buffer is shown in Figure 5 . Given that each experience has different enhancement effects on the model, to improve the utilization efficiency of data, the TD-Error is used to define the priority of the experience, and the data that can provide greater help to the model is focused on for training. The TD-Error reflects the deviation between the model output and the true value. By selectively sampling data with high TD-Error, the model performance is strengthened, and at the same time, the proportion of low-quality data is reduced.
[0112] Calculate the priority priority of the data i and the formula for calculating the sampling rate P(i) of the data is as follows, where σ i represents the value of the TD-Error:
[0113] priority i =|σ i |
[0114]
[0115] A schematic diagram of the SumTree data structure is shown in Figure 6 . Using the SumTree data structure to store the sampling rate of data can efficiently save the search time. The SumTree belongs to a binary tree type of data structure. All leaf nodes are used to store the data sampling rate P(i), and the value of all parent nodes is the sum of the values of their child nodes.
[0116] Table 1
[0117]
[0118] The model parameters of this embodiment are set as shown in Table 1.
[0119] Step5. Use the final model to complete path planning.
[0120] After the above steps, use the trained final model to complete path planning.
Claims
1. The path planning method based on the improved DQN algorithm includes the following steps: Construct a two-dimensional grid network as an environment map for path planning; Based on the two-dimensional grid network environment map and Markov decision process (MDP), a DQN path planning model is constructed; Improve and optimize the DQN path planning model to obtain the final model; Using the final model, path planning is completed.
2. The path planning method based on the improved DQN algorithm according to claim 1, characterized in that: In a two-dimensional grid network map, each square is the same size and evenly distributed. The squares are divided into four types: starting squares, feasible squares, obstacle squares, and end squares. In fact, the squares and the end squares are connected.
3. The path planning method based on the improved DQN algorithm according to claim 1, characterized in that: MDP is a framework that can be composed of a five-tuple (S, A, P, R, γ), where (S, A, R) are the three sets contained in MDP, representing: S: state set A: Action Collection R: Reward Collection P is the probability distribution contained in the MDP. There are two probability distributions: State transition probability: p(s1|s,a) represents the probability of executing action a in state s and transitioning to state s1; Reward probability: p(r|s,a) represents the probability of getting reward r by performing action a in state s; Strategy: Strategy Π(a, s) represents the probability of selecting action a in state s; The reinforcement learning algorithm needs to obtain all the state-action value functions of a given state, and select the maximum value from them to get the optimal strategy. The formula is as follows:
4. The path planning method based on the improved DQN algorithm according to claim 1, characterized in that: Methods for optimizing DQN include: applying the decaying ε-greedy strategy to select actions, improving the reward function, using the priority experience replay strategy, and introducing prior knowledge to replace the ε-greedy strategy, reward function, and experience replay strategy in the traditional DQN path planning model.
5. The path planning method based on the improved DQN algorithm according to claim 4, characterized in that: For the strategy of attenuating ε-greedy applied to action selection, the specific implementation is as follows: Decayed ε-greedy is a slight modification of ε-greedy. ε represents the trade-off between exploration and exploitation. At the beginning, we hope that the value of ε is high, so that we can have a high degree of exploration, which can be understood as the agent can learn more new things. As the agent understands the environment and the future rewards, ε should decay, so that the higher Q value found by the agent can be fully utilized. That is, as time goes on, the value of ε continues to decay and becomes smaller and smaller. The traditional ε-greedy method formula is as follows: Among them, a represents action, s represents state, and A(s) represents the number of actions in the action space. is the state-action value function, The probability of selecting a greedy action is greater than or equal to the probability of selecting any other action. When ε = 0, it becomes a greedy strategy; when ε = 1, it becomes a uniform distribution, randomly selecting actions from the action space. For the decaying ε-greedy method, assuming that there are N episodes in training, the algorithm initially sets ε = p init (For example, p init =0.7), and then gradually decreases to ε = p with training. end (For example, p init =0.1). Specifically, in the initial training process, the model is allowed to explore more freely with a high probability, and then gradually decreases at a rate γ as the training progresses, increasing the probability of full utilization and reducing the probability of free exploration. The formula is as follows: ε←(p init -p end )γ+p end。 6. The path planning method based on the improved DQN algorithm according to claim 4, characterized in that: The improvements to the reward function are as follows: The reward is divided into five parts: distance reward, avoidance reward, arrival reward, obstacle penalty and retention penalty, as shown in the following formula, where target_distance represents the agent's distance in S t+1 Stage and S t The absolute value of the distance difference from the target point in the stage. The closer the agent is to the target point, the higher the reward; conversely, the farther it deviates from the target point, the greater the penalty. At the same time, when the agent avoids obstacles, it should be rewarded to encourage it to avoid obstacles. obstacle_distance is S t+1 Stage and S t The absolute value of the distance difference between the agent and the nearest obstacle in the stage, b is the stall penalty, which is used to encourage the agent to find a solution more quickly, The agent receives a reward of +30 when it reaches the target and a reward of -30 when it collides with an obstacle. Experiments show that setting k = -0.2, μ = 0.003, and b = -0.005 can promote rapid convergence of the model.
7. The path planning method based on the improved DQN algorithm according to claim 4, characterized in that: The introduction of prior knowledge is specifically implemented as follows: Before the algorithm is trained, the simulated annealing algorithm is first used to enable the agent to find a collision-free path from the initial point to the target point in the static environment. This collision-free path is recorded as s = (s1, s2, ... s n ), when initializing the action value function, a reasonable value greater than 0 is assigned to the action value on this path, so as to avoid the system's exploration from scratch, reduce invalid exploration caused by randomness, reduce trial and error from scratch, shorten the overall learning time, and significantly improve the convergence speed.
8. The path planning method based on the improved DQN algorithm according to claim 4, characterized in that: For the priority experience replay strategy, the specific implementation is as follows: Given that each experience has different enhancing effects on the model, in order to improve the efficiency of data utilization, TD-Error is used to define the priority of experience, focusing on training data that can provide greater help to the model. TD-Error reflects the deviation between the model output and the true value. The model performance is enhanced by selectively sampling data with high TD-Error, while reducing the proportion of low-quality data and calculating the priority of data. i The formula for calculating the sampling rate P(i) of the data is as follows, where σ i Indicates the value of TD-Error: priority i =|σ i | Using the SumTree data structure to store data sampling rates can efficiently save search time. SumTree is a binary tree type data structure. All leaf nodes are used to store data sampling rates P(i), and the value of all parent nodes is the sum of the values of their child nodes.
Citation Information
Cited By
Disaster reduction and rescue path planning method and system based on Internet of Things
CN121031933A
Radiation sensing path planning method based on deep Q network and course learning
CN121052482A
Shared bicycle scheduling variable VRP search method based on deep reinforcement learning driving
CN121936782A