A mobile robot path planning method based on improved ant colony algorithm

CN122835397APending Publication Date: 2026-09-29QIQIHAR UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611035410.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]然而,当移动机器人处于复杂环境中时,传统蚁群算法在搜索初期由于信息素分布均匀,路径选择缺乏有效引导,搜索行为呈现出较强的无序随机性,导致大量无效路径被反复探索,影响算法的收敛速度

Benefits of technology

本发明通过对蚂蚁前期每次迭代后的精英路径进行准对立学习,可以进一步丰富路径多样性,降低算法陷入局部最优的概率;自适应参数调节策略,能够在迭代前期减弱信息素正反馈作用、增强启发信息引导以提升路径探索能力,后期强化信息素积累以加速收敛;启发函数的优化能够使蚂蚁更倾向于探索转弯角度更小的新节点;回退机制,能够降低路径搜索过程中陷入死锁的概率以及避免出现震荡路径;利用冗余节点删除策略优化最终路径,能够减少路径中不必要的转折点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122835397A_ABST
    Figure CN122835397A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of mobile robot path planning, and particularly relates to a mobile robot path planning method based on an improved ant colony algorithm. The method first adopts a grid method to model a working environment of a mobile robot; subsequently, candidate node selection in an original method is optimized, the planned path is prevented from passing through a diagonal obstacle, an improved heuristic function is used to make ants more inclined to explore new nodes with small turning angles, a dynamic parameter is adjusted to balance global search capability and local optimization capability of the algorithm, a backtracking mechanism is introduced to reduce the probability of falling into a deadlock in a path search process in a complex environment, a quasi-antagonistic learning is conducted on elite paths after each iteration in an early stage to further enrich path diversity and reduce the probability of falling into a local optimum of the algorithm; finally, the final path planned by the method is subjected to secondary optimization to reduce unnecessary turning points in the path. Experimental results show that the application can effectively shorten the path length, improve path smoothness and algorithm convergence efficiency, and has good environmental adaptability and application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to mobile robot path planning, specifically involving an ant colony algorithm path planning method based on a fusion of elite quasi-oppositional learning and multi-strategy improvement. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence and automatic control technologies, numerous benefits have been brought about in reducing manpower input, improving production efficiency, and saving time. Mobile robots are increasingly widely used in military, medical services, scientific research and exploration, intelligent manufacturing, warehousing and logistics, and autonomous driving. Reliable, safe, and efficient mobile robots are an important research area in robotics. Path planning, as a key problem in mobile robot research, aims to find an optimal or near-optimal path from the starting point to the target point in an environment with obstacles, ensuring safety, no collisions, and satisfying multiple constraints.

[0003] Commonly used algorithms in global path planning for mobile robots include traditional algorithms and intelligent heuristic algorithms. Traditional algorithms include A* algorithm, RRT algorithm, and artificial potential field method. These algorithms perform well in simple graph environments, but their computational cost increases significantly in complex environments, making it difficult to balance computational efficiency with path optimality. Intelligent heuristic algorithms include ant colony optimization (ACO), particle swarm optimization (PSO), and gray wolf algorithm. Among these, ant colony optimization (ACO) has attracted widespread attention due to its strong robustness, positive feedback mechanism, parallel search capability, and ease of integration with other algorithms.

[0004] However, when mobile robots are in complex environments, traditional ant colony algorithms exhibit strong randomness in the initial search phase due to the uniform distribution of pheromones and the lack of effective guidance in path selection. This leads to the repeated exploration of numerous invalid paths, affecting the algorithm's convergence speed. Furthermore, the existence of deadlock problems results in a high ant mortality rate, and the planned paths often traverse diagonal obstacles, posing safety risks. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides a mobile robot path planning method based on an improved ant colony algorithm. This method improves the algorithm's global search capability and ability to escape deadlock by introducing an adaptive parameter adjustment strategy, an elite quasi-opposition learning strategy, and a backoff mechanism, while reducing the probability of it reaching local optima. Furthermore, it utilizes a redundant node deletion strategy to further optimize path quality.

[0006] The technical solution adopted by the present invention to achieve the above objectives includes the following steps: S1. The robot's workspace is divided into multiple uniform two-dimensional grid units using a grid method. Based on the robot's size parameters, obstacle areas in the environment are expanded to ensure a safe distance for the robot during movement. Based on the grid division results, a two-dimensional grid map is constructed and represented as a binary matrix. In this matrix, a value of 0 indicates that the corresponding grid is a passable area, and a value of 1 indicates that the corresponding grid is an obstacle area. The robot's starting position and target position are marked in the two-dimensional grid map for subsequent path planning calculations. S2. Initialize algorithm parameters: number of ants m, maximum number of iterations T, maximum expected pheromone factor. Minimum value of pheromone expectation factor Maximum value of heuristic factor Minimum value of heuristic factor Maximum value of pheromone volatility coefficient Minimum value of pheromone volatility coefficient The parameters are: pheromone concentration range, pheromone intensity Q, number of elite ants K, number of times elite quasi-oppositional learning is performed N, and the initial global pheromone matrix is ​​set to half of the maximum pheromone concentration. S3. During the traversal of the eight neighborhoods around the current node, add constraint judgments to the diagonal nodes. If the orthogonal grids on both sides of the candidate node in the diagonal direction are obstacles, it is forbidden to add the diagonal node to the candidate node set. When constructing the path using the state transition function, add the node visit penalty and the corner penalty to the heuristic function. S4. When encountering a deadlock or when the number of node accesses exceeds the threshold, a rollback strategy is executed. After each iteration, it is determined whether to execute the elite quasi-opposition learning process, and then the pheromone matrix is ​​updated. S5. If the maximum number of iterations has not been reached, return to step S2 and re-execute steps S2 to S4. If the maximum number of iterations has been reached, perform secondary optimization on the optimal path to remove redundant nodes and turning points before outputting.

[0007] Furthermore, the adaptive hyperparameters in step S2 are updated according to equations (1), (2), and (3): (1) (2) (3) Where t represents the current iteration number, T represents the maximum iteration number, r is a random number in the interval (-0.1, 0.1), α(t) is the expected pheromone factor value of the current iteration, which determines the degree of influence of pheromones in path selection, β(t) is the heuristic factor value of the current iteration, which reflects the guiding role of heuristic information in path selection, and ρ(t) is the pheromone evaporation coefficient of the current iteration, which determines the rate of pheromone evaporation in the path.

[0008] Furthermore, the heuristic function in step S3 is expressed as follows: (4) (5) (6) (7) in, This represents the Euclidean distance between the current node i and the candidate node j. Let represent the Euclidean distance between candidate node j and target point g, and let cosθ represent the cosine of the angle between the current point and the candidate node and the target point. This indicates the number of times the candidate node has been accessed.

[0009] Furthermore, the state transition function in step S3 is expressed as: (8) in, It is the set of next nodes that ant k can choose when it is located at node i. It represents the pheromone concentration on path (i,j) in the current iteration. This is the heuristic value for path (i,j) in the current iteration. The pheromone update in the path is performed according to equations (9), (10), and (11): (9) (10) (11) Where ρ is the pheromone evaporation coefficient, ρ∈ (0,1), It is the pheromone left by ant k when it passes through path (i,j) in the current iteration. Q is a constant representing the pheromone intensity. It is the total length of the search path of the k-th ant in the current iteration.

[0010] Instead of the taboo table used in traditional ant colony optimization (ACO) algorithms, this algorithm records the number of times each node is visited. When a path becomes deadlocked or a node's visit count exceeds a threshold, ants are allowed to backtrack and reselect a new path. To enrich path diversity and reduce premature convergence in the early stages of iteration, an elite quasi-oppositional learning strategy is introduced. If the current iteration count is less than or equal to the number of times elite quasi-oppositional learning is executed, after each iteration, the paths in the path set are sorted in ascending order of cost value. The top K paths with the lowest cost are selected as elite solutions, and the path at the end of the sorted set is the worst solution for the current iteration. For each elite path, the opposing positions of all nodes except the start and end points are determined. Random reconstruction is performed within the feasible neighborhood of each node and its opposing position to generate new candidate paths. The cost of the new path is calculated. If it is better than the worst solution of the current iteration, the worst path in the path set is replaced by the new path after elite quasi-oppositional learning, and it participates in the subsequent pheromone update process. This process is performed sequentially on the K elite paths. If the worst solution in this iteration has been updated, the previous solution in the path set that is better than the worst solution becomes the new worst solution.

[0011] To ensure the feasibility of the path after elite quasi-opposition learning, the starting and ending points of the fixed path are not involved in the elite quasi-opposition learning process; only the intermediate nodes of the path are processed. The validity of each node in the new path is checked, nodes existing on obstacles are removed, and then the connectivity between any two nodes is checked. If the path between two nodes will not collide with obstacles, large-step exploration can be achieved; if there is a possibility of collision between the path and obstacles, an 8-neighborhood greedy search is used to fill in the discontinuities in the path between the two nodes, ensuring the feasibility and safety of the new path after elite quasi-opposition learning.

[0012] Furthermore, the elite quasi-oppositional learning process in step S4 is represented as follows: (12) (13) Where lb and ub represent the lower and upper bounds of the search space, respectively, x∈ [lb,ub], Let x be the opposite of x, and γ be a random quasi-opposite factor, γ∈[0.25,0.75]. It is the quasi-random opposite of x.

[0013] After iteration and the algorithm outputs the final path, a secondary optimization is performed to remove redundant nodes and turning points. Starting from the path's starting point, the Bresenham algorithm is used to determine the connectivity between non-adjacent nodes. While ensuring the path doesn't collide with obstacles, intermediate redundant nodes are removed to simplify the path. For consecutive nodes on the same straight line, intermediate nodes can be directly deleted. For turning points, it's necessary to further determine whether the connecting path between the current node and its subsequent nodes will collide with obstacles. If there's no collision risk, the turning point is considered redundant and deleted; otherwise, it's retained to ensure path feasibility. The optimized path not only reduces the number of nodes but also the number of turns, better meeting the robot's actual operational needs.

[0014] The beneficial effects of this invention are as follows: This invention enriches path diversity and reduces the probability of the algorithm getting stuck in local optima by performing quasi-oppositional learning on the elite paths of ants after each iteration in the early stages. The adaptive parameter adjustment strategy weakens the positive feedback effect of pheromones and enhances the guidance of heuristic information to improve path exploration ability in the early stages of iteration, and strengthens pheromone accumulation to accelerate convergence in the later stages. The optimization of the heuristic function makes ants more inclined to explore new nodes with smaller turning angles. The backoff mechanism reduces the probability of deadlock during path search and avoids oscillating paths. The optimization of the final path using the redundant node deletion strategy reduces unnecessary turning points in the path. Attached Figure Description

[0015] Figure 1 is a flowchart of the method of the present invention.

[0016] Figure 2 is a graph showing the variation of adaptive parameters in the method of the present invention.

[0017] Figure 3 shows the path planning results of the method of the present invention and the traditional ant colony algorithm in a 20×20 grid map environment.

[0018] Figure 4 shows the iterative convergence curves of the method of the present invention and the traditional ant colony algorithm in a 20×20 grid map environment.

[0019] Figure 5 shows the path planning results of the method of this invention and the traditional ant colony algorithm in a 30×30 grid map environment.

[0020] Figure 6 shows the iterative convergence curves of the method of the present invention and the traditional ant colony algorithm in a 30×30 grid map environment. Detailed Implementation

[0021] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments or drawings are only used to illustrate the present invention and do not limit the scope of the present invention.

[0022] As shown in Figure 1, a mobile robot path planning method based on an improved ant colony algorithm is described in this embodiment, which includes the following steps: S1. The robot's workspace is divided into multiple uniform two-dimensional grid units using a grid method. Based on the robot's size parameters, obstacle areas in the environment are expanded to ensure a safe distance for the robot during movement. Based on the grid division results, a two-dimensional grid map is constructed and represented as a binary matrix. In this matrix, a value of 0 indicates that the corresponding grid is a passable area, and a value of 1 indicates that the corresponding grid is an obstacle area. The robot's starting position and target position are marked in the two-dimensional grid map for subsequent path planning calculations. S2. Initialize algorithm parameters: number of ants m=60, maximum number of iterations T=100, maximum expected pheromone factor. =3. Minimum value of pheromone expectation factor =1, Maximum value of the heuristic factor =6. Minimum value of heuristic factor =3. Maximum value of pheromone volatility coefficient =0.3, minimum pheromone evaporation coefficient =0.1, pheromone concentration range [0.01, 5], pheromone intensity Q=1, elite ant number K=6, elite quasi-opposition learning execution times N=30, the initial global pheromone matrix is ​​set to half of the maximum pheromone concentration. The curve of the adaptive parameter value in this step changing with the number of iterations is shown in Figure 2, and it is updated according to the following formula: (1) (2) (3) Where t represents the current iteration number, T represents the maximum iteration number, r is a random number in the interval (-0.1, 0.1), α(t) is the expected pheromone factor value of the current iteration, which determines the degree of influence of pheromone in path selection, β(t) is the heuristic factor value of the current iteration, which reflects the guiding role of heuristic information in path selection, and ρ(t) is the pheromone evaporation coefficient of the current iteration, which determines the rate of pheromone evaporation in the path; S3. During the traversal of the eight neighbors of the current node, constraints are added to nodes in the diagonal direction. If both orthogonal grids on both sides of a candidate node in the diagonal direction are obstacles, the diagonal node is prohibited from being added to the candidate node set. When constructing the path using the state transition function, penalties for node visits and corner turns are added to the heuristic function. The optimized heuristic function in this step is expressed as follows: (4) (5) (6) (7) in, This represents the Euclidean distance between the current node i and the candidate node j. Let represent the Euclidean distance between candidate node j and target point g, and let cosθ represent the cosine of the angle between the current point and the candidate node and the target point. This represents the number of times the candidate node has been visited; the state transition function when the ant selects the next node in this step is expressed as: (8) in, It is the set of next nodes that ant k can choose when it is located at node i. It represents the pheromone concentration on path (i,j) in the current iteration. This is the heuristic value for path (i,j) in the current iteration. The pheromones in the path are dynamically updated according to the following formula: (9) (10) (11) Where ρ is the pheromone evaporation coefficient, ρ∈ (0,1), It is the pheromone left by ant k when it passes through path (i,j) in the current iteration. Q is a constant representing the pheromone intensity. It is the total length of the search path of the k-th ant in the current iteration; S4. When encountering a deadlock or when the number of node accesses exceeds the threshold, a rollback strategy is executed. After each iteration, it is determined whether to execute the elite quasi-opposition learning process, and then the pheromone matrix is ​​updated. S5. If the maximum number of iterations has not been reached, return to step S2 and re-execute steps S2 to S4. If the maximum number of iterations has been reached, perform secondary optimization on the optimal path to remove redundant nodes and turning points before outputting.

[0023] In this embodiment, the tabu list in the traditional ant colony algorithm is removed, and instead, the number of visits to each node is recorded. When the path gets deadlocked or the number of visits to a node exceeds a threshold during exploration, ants are allowed to backtrack and choose a new exploration path. To enrich path diversity and reduce the probability of premature convergence in the early stages of iteration, an elite quasi-oppositional learning strategy is introduced. If the current iteration number is less than or equal to the number of elite quasi-oppositional learning executions, after each iteration, the paths in the path set are sorted in ascending order of cost value, and the top K paths with the lowest cost are selected as elite solutions. The path at the end of the sorted set is the worst solution for the current iteration. For each elite path, the opposing positions of all nodes except the starting and ending points are determined. Random reconstruction is performed within the feasible neighborhood of each node and its opposing position to generate new candidate paths, and the cost of the new path is calculated. If the new path is better than the worst solution of the current iteration, the worst path in the path set is replaced by the new path after elite quasi-oppositional learning, and it participates in the subsequent pheromone update process. This process is performed sequentially on the K elite paths. If the worst solution in this iteration has been updated, the previous solution in the path set that is better than the worst solution becomes the new worst solution.

[0024] In this embodiment, to ensure the feasibility of the path after elite quasi-opposition learning, the starting and ending points of the fixed path are not involved in the elite quasi-opposition learning process; only the intermediate nodes of the path are processed. The validity of each node in the new path is checked, nodes existing on obstacles are removed, and then the connectivity between every two nodes is checked. If the path between two nodes will not collide with obstacles, large-step exploration can be achieved; if the path between two nodes may collide with obstacles, an 8-neighborhood greedy search is used to supplement the discontinuity of the path between the two nodes, ensuring the feasibility and safety of the new path after elite quasi-opposition learning.

[0025] In this embodiment, the elite quasi-oppositional learning process is represented as follows: (12) (13) Where lb and ub represent the lower and upper bounds of the search space, respectively, x∈ [lb,ub], Let x be the opposite of x, and γ be a random quasi-opposite factor, γ∈[0.25,0.75]. It is the quasi-random opposite of x.

[0026] In this embodiment, after the algorithm iterates and outputs the final path, a secondary optimization is performed to remove redundant nodes and turning points. Starting from the path's starting point, the Bresenham algorithm is used to determine the connectivity between non-adjacent nodes. While ensuring the path does not collide with obstacles, intermediate redundant nodes are deleted to simplify the path. For consecutive nodes on the same straight line, intermediate nodes can be directly deleted. For turning points, it is necessary to further determine whether the connecting path between the current node and its subsequent nodes will collide with obstacles. If there is no collision risk, the turning point can be considered redundant and deleted; otherwise, it must be retained to ensure path feasibility. The path after secondary optimization not only reduces the number of nodes but also reduces the number of turns, better meeting the actual operational needs of the robot.

[0027] The feasibility and superiority of the algorithm of this invention were further verified through simulation experiments conducted in a two-dimensional grid map environment. The results are shown in Figures 3 to 6, with maps of 20×20 and 30×30 grid environments, respectively. Each grid represents an actual spatial area of ​​1m×1m. The algorithm of this invention was compared with the traditional ant colony algorithm. Both algorithms were run independently 30 times in two different map environments, and the optimal path length, the average optimal path length, and the number of turns were selected as evaluation indicators. In the 20×20 grid map, the starting and ending coordinates of the path were set to (0.5, 0.5) and (19.5, 19.5), respectively; in the 30×30 grid map, the starting and ending coordinates of the path were set to (0.5, 0.5) and (29.5, 29.5), respectively.

[0028] Table 1. Simulation results under two raster map environments 20×20 JOIN US 30.3827.52 30.95 ± 0.5727.52 ± 0.00 143 30×30 EQUATE 47.7041.97 48.47±0.7742.33±0.36 2310 Simulation results show that both the traditional ant colony algorithm and the method of this invention can plan effective paths in map environments of different sizes. However, the path planned by the method of this invention is superior, with fewer turns and faster convergence. Specifically, in a 20×20 map environment, the optimal path length planned by the method of this invention is reduced by 9.4% compared to the traditional ant colony algorithm, and the path does not involve frequent turns or crossing diagonally adjacent obstacles, achieving stable convergence during the elite quasi-opposite learning phase. In a 30×30 map environment, the optimal path length planned by the method of this invention is reduced by 12% compared to the traditional ant colony algorithm, further demonstrating the significant advantage of the method of this invention in complex map environments.

[0029] The above description is merely an explanation of the principles and specific embodiments of the present invention and is not intended to limit the invention. Those skilled in the art can make various improvements and variations to the present invention. Therefore, any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A path planning method for a mobile robot based on an improved ant colony algorithm, characterized in that, Includes the following steps: S1. The robot's workspace is divided into multiple uniform two-dimensional grid units using a grid method. Based on the robot's size parameters, obstacle areas in the environment are expanded to ensure a safe distance for the robot during movement. Based on the grid division results, a two-dimensional grid map is constructed and represented as a binary matrix. In this matrix, a value of 0 indicates that the corresponding grid is a passable area, and a value of 1 indicates that the corresponding grid is an obstacle area. The robot's starting position and target position are marked in the two-dimensional grid map for subsequent path planning calculations. S2. Initialize algorithm parameters: number of ants m, maximum number of iterations T, maximum expected pheromone factor. Minimum value of pheromone expectation factor Maximum value of heuristic factor Minimum value of heuristic factor Maximum value of pheromone volatility coefficient Minimum value of pheromone volatility coefficient The parameters are: pheromone concentration range, pheromone intensity Q, number of elite ants K, number of times elite quasi-oppositional learning is performed N, and the initial global pheromone matrix is ​​set to half of the maximum pheromone concentration. S3. During the traversal of the eight neighborhoods around the current node, add constraint judgments to the diagonal nodes. If the orthogonal grids on both sides of the candidate node in the diagonal direction are obstacles, it is forbidden to add the diagonal node to the candidate node set. When constructing the path using the state transition function, add the node visit penalty and the corner penalty to the heuristic function. S4. When encountering a deadlock or when the number of node accesses exceeds the threshold, a rollback strategy is executed. After each iteration, it is determined whether to execute the elite quasi-opposition learning process, and then the pheromone matrix is ​​updated. S5. If the maximum number of iterations has not been reached, return to step S2 and re-execute steps S2 to S4. If the maximum number of iterations has been reached, perform secondary optimization on the optimal path to remove redundant nodes and turning points before outputting.

2. The method according to claim 1, characterized in that, The adaptive hyperparameters in step S2 are updated according to equations (1), (2), and (3): (1) (2) (3) Where t represents the current iteration number, T represents the maximum iteration number, r is a random number in the interval (-0.1, 0.1), α(t) is the expected pheromone factor value of the current iteration, which determines the degree of influence of pheromone in path selection, β(t) is the heuristic factor value of the current iteration, which reflects the guiding role of heuristic information in path selection, and ρ(t) is the pheromone evaporation coefficient of the current iteration, which determines the rate of pheromone evaporation in the path.

3. The method according to claim 1, characterized in that, The heuristic function in step S3 is expressed as follows: (4) (5) (6) (7) in, This represents the Euclidean distance between the current node i and the candidate node j. θ represents the Euclidean distance between the candidate node and the target point, and cosθ represents the cosine of the angle between the current point and the candidate node and the target point. This indicates the number of times the candidate node has been accessed.

4. The method according to claim 1, characterized in that, The state transition function in step S3 is expressed as: (8) in, It is the set of next nodes that ant k can choose when it is located at node i. It represents the pheromone concentration on path (i,j) in the current iteration. It is the heuristic value of path (i,j) in the current iteration, and the pheromone update in the path is performed according to equations (9), (10), and (11): (9) (10) (11) Where ρ is the pheromone evaporation coefficient, ρ∈ (0,1), It is the pheromone left by ant k when it passes through path (i,j) in the current iteration. Q is a constant representing the pheromone intensity. It is the total length of the search path of the k-th ant in the current iteration.

5. The method according to claim 1, characterized in that, The elite quasi-oppositional learning process in step S4 is represented as follows: (12) (13) Where lb and ub represent the lower and upper bounds of the search space, respectively, x∈ [lb,ub], Let x be the opposite of x, and γ be a random quasi-opposite factor, γ∈[0.25,0.75]. It is the quasi-random opposite of x.