Robot path planning method and system based on artificial bee colony and reinforcement learning
By combining artificial bee colony and reinforcement learning methods, the artificial bee colony algorithm is used to generate an elite solution set and encode it as prior knowledge for reinforcement learning. A multi-objective reward function is designed to solve the efficiency and quality problems of robot path planning in complex environments, generating short and smooth paths that are suitable for industrial robots and logistics distribution.
Patent Information
- Application Number
- CN202511493967.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing robot path planning methods struggle to quickly generate short and smooth paths in complex environments. Traditional algorithms suffer from high computational complexity and poor real-time performance, while single reinforcement learning methods offer only blind initial exploration and are unlikely to find the globally optimal solution.
By combining artificial bee colony algorithm and reinforcement learning, an elite solution set is generated by initializing the artificial bee colony algorithm and encoded as the prior knowledge of the reinforcement learning agent. A reward function that integrates path length and smoothness is designed, and the policy is fine-tuned using the Q-learning algorithm to generate the optimal path.
It enables the rapid generation of short and smooth paths in complex environments, improving the efficiency and quality of path planning. It is suitable for scenarios such as automated scanning of complex workpiece surfaces and logistics distribution by industrial robots.
Smart Images

Figure CN120996078A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of path planning, in particular to a robot path planning method and system based on artificial bee colony and reinforcement learning. BACKGROUND
[0002] The existing robot path planning methods can be mainly divided into traditional algorithms, heuristic algorithms and emerging machine learning methods. Traditional graph search algorithms, such as A* algorithm and Dijkstra algorithm, can guarantee to find the shortest path, but their computational complexity increases exponentially with the complexity of the environment and the resolution of the map, and they have poor real-time performance and are difficult to directly optimize complex indicators such as path smoothness. Heuristic intelligent optimization algorithms, such as genetic algorithm (Genetic Algorithm, GA), particle swarm optimization algorithm (Particle Swarm Optimization, PSO) and the artificial bee colony algorithm (Artificial Bee Colony, ABC) involved in the present application, search for global optimal solutions by simulating natural phenomena, and show good global search ability and flexibility in solving complex path planning problems. However, such algorithms also have inherent defects: their convergence speed is often slow, especially when the solution space dimension is high and the constraints are complex, a large number of iterative calculations are needed to converge; in addition, the algorithm is prone to wander around the local optimal solution in the later stage, the population diversity decreases, resulting in reduced optimization efficiency.
[0003] In recent years, reinforcement learning (Reinforcement Learning, RL) has attracted widespread attention in the field of path planning due to its strong sequential decision-making ability. Through interaction with the environment and the use of reward functions, RL agents can learn efficient path strategies. However, deep reinforcement learning often requires a large amount of sampling and data training, and the learning process is slow; while table-based methods such as Q-learning face the "curse of dimensionality" problem when the state space is large. More importantly, the performance of RL algorithms is highly dependent on the initial strategy and exploration efficiency, and in the initial stage of completely random exploration, the learning efficiency is low and it is difficult to guarantee convergence to the global optimal solution.
[0004] Therefore, one of the outstanding problems in the prior art is that a single global heuristic algorithm (such as ABC) can perform extensive exploration, but has slow convergence and insufficient local development capability in the later stage; while a single reinforcement learning method is good at local fine decision-making, but its initial exploration is blind, the global search efficiency is low, and it is difficult to optimize multiple objectives (such as the shortest path and the fewest turning points) at the same time. How to combine the advantages of the two types of algorithms and design a hybrid optimization method that can quickly and efficiently plan a path that is both short and smooth has become a technical problem that needs to be solved in the field.
[0005] Therefore, there is a need for an artificial bee colony and reinforcement learning based robot path planning method and system that can guarantee the quality and efficiency of path planning and quickly generate an optimized path that is both short and smooth in a complex environment. SUMMARY
[0006] The main purpose of the present application is to provide an artificial bee colony and reinforcement learning based robot path planning method and system to solve the problem that the prior art cannot quickly generate a path that is both short and smooth in a complex environment.
[0007] To achieve the above purpose, the present application provides an artificial bee colony and reinforcement learning based robot path planning method, which specifically comprises the following steps: S1, initializing the artificial bee colony algorithm population according to the known environment information, the starting point and the target point.
[0008] S2, iteratively optimizing the population based on the search mechanism of the artificial bee colony algorithm.
[0009] S3, when the preset optimization process termination condition is met, selecting a batch of high-quality solutions from the final population to form an elite solution set.
[0010] S4, encoding the elite solution set as prior knowledge of the reinforcement learning agent to initialize the value function of the reinforcement learning, and setting the state space, action space and reward function of the reinforcement learning that combines the path length and smoothness objectives.
[0011] S5, based on the initialized value function, running the reinforcement learning algorithm for policy search and optimization to fine-tune the path strategy.
[0012] S6, outputting the optimal path after the reinforcement learning algorithm converges.
[0013] Further, step S1 specifically comprises the following steps: S1.1, each feasible path is composed of a series of node sequences and is represented as: ; wherein, is the starting point, is the target point, and is a node in the environment.
[0014] S1.2, performing global search using the artificial bee colony algorithm; initializing the ABC population, the population size is , each honey source represents a path containing nodes; the encoding method of the honey source adopts integer encoding, and the initialization formula of the th honey source is: ; in, nectar source The first in One node; The generation rules are as follows: ; in, It is the set of all nodes in the environment, excluding the start and end points; , This represents the total number of nodes. This means randomly selecting a node from the set of nodes.
[0015] Furthermore, step S2 specifically includes the following steps: S2.1, using the fitness function to comprehensively consider path length and smoothness, the fitness function... The expression is: ; in, The total length of the path is expressed by the following formula: ; in, Indicates Euclidean distance; Representing a path The number of inflection points; This is the smoothness penalty coefficient.
[0016] S2.2, Smoothness is reflected by the number of inflection points. The number of inflection points is determined by calculating the included angle between the vectors of three consecutive points. Let three consecutive points... , , This forms two vectors. and The cosine of the angle between two vectors for: ; like ,but This is determined to be an inflection point; among them, For the preset threshold, The larger the value, the weaker the strength of the inflection point determination.
[0017] S2.3, based on fitness function Neighborhood search and population update are performed; the formula for generating new nectar sources during the hired bee stage is: ; in, For new honey sources; This is the current honey source; randomly selected field as a honey source; a random number in the range of [-1, 1].
[0018] Further, step S3 specifically comprises: When the number of iterations of the ABC algorithm reaches a preset maximum value , or the optimal fitness of the population improves by less than a threshold value in succession , the search of the ABC algorithm is terminated; the highest fitness non-repeated paths in the final population are selected to form an elite solution set ; the selection formula is: ; wherein, is a fitness threshold value, is the i-th path.
[0019] Further, step S4 specifically comprises the following steps: S4.1, for each path in the elite solution set , traverse the state-action pairs of each path , wherein, is the current node state, and the action is the next node selected; the value of the state-action pair is initialized as a priority value : ; wherein, is the action value function.
[0020] S4.2, set the environment of reinforcement learning, the state space is the set of all nodes; the action space is the set of adjacent nodes reachable from the state ; the reward function is a fusion of path length and smoothness objective, and the reward function is designed as: ; wherein, is the moving distance; is an indicator function, which returns 1 if a turning point is generated at , and otherwise returns 0; is a turning point penalty weight.
[0021] S4.3, a Q-learning algorithm is used for policy fine-tuning, and the value update formula is: ; in, and These represent the current state and action, respectively. Represents the execution of actions The next state after that; Indicates the next state Below, one of the action variables among all possible actions that the agent can perform; Represents the next state Of all possible actions The largest value; Indicates an immediate reward; The learning rate; This is the discount factor.
[0022] Furthermore, step S5 specifically includes the following steps: S5.1, Q-learning employs an ε-greedy strategy to balance exploration and exploitation. The decision-making process of the ε-greedy strategy is described as follows: ; in, , representing the exploration rate; Indicates that the current The action variable with the largest value.
[0023] S5.2, when The change in the value function is less than the convergence threshold. ,Right now: ; Or reach the maximum number of training rounds The learning process will terminate at that time. Among them, For the number of iterations, For the first After the next iteration Value function.
[0024] Furthermore, step S6 specifically includes: From the starting point Initially, the agent chooses the option with the greatest [potential] at each step. Value action : ; Until the target point is reached Generate the final path list .in, The first in the optimal path list Each node.
[0025] The application further provides a robot path planning system based on artificial bee colony and reinforcement learning, comprising: A population initialization module is configured to initialize an artificial bee colony algorithm population according to known environment information, a starting point and a target point, wherein each individual in the population represents a feasible path from the starting point to the target point composed of path points; A global search optimization module is configured to iteratively optimize the population based on a search mechanism of the artificial bee colony algorithm, wherein the quality of the individual is evaluated by a fitness function that comprehensively considers path length and path smoothness, and the population is updated accordingly; An elite solution set construction module is configured to select a batch of high-quality solutions from the final population to form an elite solution set when the global search optimization module meets a preset optimization process termination condition; A reinforcement learning initialization module integrates a reinforcement learning algorithm and is configured to encode the elite solution set as prior knowledge of a reinforcement learning agent to initialize a value function, set a state space, an action space and a reward function that combines path length and smoothness objectives; A strategy fine-tuning module is configured to run the reinforcement learning algorithm based on the initialized value function to perform strategy search and optimization and fine-tune the path strategy; A path output module is configured to output the optimal path determined after convergence, which meets the requirements of shortest length and fewest inflection points.
[0026] The application has the following beneficial effects: 1. The application realizes the organic combination of global exploration and local development through a hierarchical optimization architecture, which not only leverages the global search capability of the artificial bee colony algorithm but also utilizes the local optimization advantage of the reinforcement learning algorithm, effectively avoiding the problem of local optimum.
[0027] 2. The application designs a multi-objective fitness function and a reward function that combine path length and smoothness, which can significantly reduce the number of inflection points while ensuring the shortest path, generate a smoother path that better meets the motion characteristics of the robot, and improve motion efficiency and service life.
[0028] 3. The application proposes an elite solution set screening and conversion mechanism, which converts high-quality solutions obtained by the artificial bee colony algorithm into prior knowledge of the reinforcement learning algorithm, significantly improves the initial performance and learning efficiency of the reinforcement learning algorithm, and speeds up the overall convergence speed.
[0029] 4. The application solves the problem of slow convergence speed or low solution quality of a single algorithm in complex path planning and is suitable for industrial robots equipped with multi-line laser scanners to perform automatic scanning operations on complex workpiece surfaces, and has application value in logistics distribution, intelligent inspection and other scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0030] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart of a robot path planning method based on artificial bee colony and reinforcement learning according to the present invention is shown.
[0031] Figure 2 This illustration shows an application scenario of an industrial robot equipped with a multi-line laser scanner scanning a workpiece, as described in an embodiment of the present invention.
[0032] Figure 3 This is a schematic diagram illustrating the generation of sampling points on the surface of the workpiece under test in a scanning and measurement robot environment as described in this embodiment of the invention.
[0033] Figure 4 This is a schematic diagram of the planned scanning viewpoint set on the surface of the workpiece under test in a scanning measurement robot environment involved in the embodiments of the present invention.
[0034] Figure 5 This is a schematic diagram of the optimal path generated by the robot path planning method based on artificial bee colony and reinforcement learning in an embodiment of the present invention.
[0035] Figure 6 This is a schematic diagram showing the final scanning effect of an industrial robot equipped with a multi-line laser scanner on the surface morphology of a workpiece, generated according to the robot path planning method based on artificial bee colony and reinforcement learning as described in an embodiment of the present invention.
[0036] The reference numerals in the above figures are explained as follows: 1. Industrial robot; 2. Multi-line laser scanner; 3. Binocular tracking camera; 4. Workpiece; 5. Computer. Detailed Implementation
[0037] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] Example 1 like Figure 1 The robot path planning method based on artificial bee colony and reinforcement learning shown includes the following steps: S1, initializing a population of artificial bee colony algorithm according to known environment information, a starting point and a target point.
[0039] S2, iteratively optimizing the population based on a search mechanism of the artificial bee colony algorithm.
[0040] S3, when a preset optimization process termination condition is met, selecting a batch of high-quality solutions from the final population to form an elite solution set.
[0041] S4, encoding the elite solution set as prior knowledge of a reinforcement learning agent to initialize a value function of the reinforcement learning, and setting a state space, an action space and a reward function of the reinforcement learning which is fused with a path length and a smoothness target.
[0042] S5, based on the initialized value function, running a reinforcement learning algorithm to search and optimize a strategy to fine-tune a path strategy.
[0043] S6, outputting an optimal path after the reinforcement learning algorithm converges.
[0044] Specifically, step S1 specifically includes the following steps: S1.1, each feasible path is composed of a series of node sequences and is represented as: ; wherein, is a starting point, is a target point, and is a node in the environment.
[0045] S1.2, performing global search by using an artificial bee colony algorithm (ABC); initializing an ABC population, and the population size is , each honey source (i.e., individual) represents a path and contains nodes; the encoding mode of the honey source adopts integer encoding, and the initialization formula of the honey source is: ; wherein, is the 1st node in the honey source , is the 2nd node in the honey source , and so on; is the node in the honey source ; the generation rule of is: ; wherein, A set of all nodes in the environment except the start node and the end node; , A total number of nodes; A randomly selected node from the node set.
[0046] The search mechanism based on the artificial bee colony algorithm iteratively optimizes the population, wherein the advantages and disadvantages of individuals are evaluated by a fitness function that comprehensively considers path length and path smoothness, and the population is updated accordingly.
[0047] Specifically, step S2 specifically includes the following steps: S2.1, the iteration process of the ABC algorithm evaluates the advantages and disadvantages of each honey source using a multi-objective fitness function, which comprehensively considers path length and smoothness using a fitness function, and the fitness function is The expression is: ; Wherein, represents the total length of the path, and the calculation formula is: ; Wherein, represents the Euclidean distance; represents the number of inflection points of the path ; is a smoothness penalty coefficient, used to adjust the weight of path length and smoothness.
[0048] S2.2, smoothness is reflected by the number of inflection points, and the number of inflection points is determined by calculating the vector angle of three consecutive points, assuming that three consecutive points , , form two vectors and , and the cosine value of the angle between the two vectors is: ; If (where is a preset threshold, generally set to 30°~45°), then is determined as an inflection point; wherein, is a preset threshold, the greater, the weaker the inflection point determination strength.
[0049] S2.3, the ABC algorithm performs neighborhood search and population update based on the fitness function through the cooperative search mechanism of employed bees, onlookers and scouts; the new honey source generation formula in the employed bee stage is: ; Wherein, is a new honey source; is a current honey source; is a randomly selected domain honey source; is a random number in the range of [-1, 1].
[0050] Specifically, step S3 is specifically: when the number of iterations of the ABC algorithm reaches a preset maximum value , or the optimal fitness of the population improves by less than a threshold value within a continuous generation, is a positive integer. Terminate the ABC algorithm search; select the top non-repeating paths with the highest fitness from the final population to form an elite solution set ; the selection formula is: ; wherein, is a fitness threshold, is the path.
[0051] Specifically, step S4 specifically includes the following steps: S4.1, the elite solution set is encoded as prior knowledge of the reinforcement learning agent, used to initialize values. For each path in the elite solution set, traverse the state-action pairs of each path , wherein, is the current node state, and the action is the next node selected; the value of the state-action pair is initialized to a priority value : ; wherein, is the action value function.
[0052] S4.2, set the environment of the reinforcement learning, the state space is the set of all nodes; the action space is the set of adjacent nodes reachable from the state ; the reward function is a fusion of path length and smoothness objectives, and the reward function is designed as: ; wherein, is the movement distance; is an indicator function that returns 1 if a turning point is produced at , and 0 otherwise. penalty weight for the inflection point.
[0053] The reward function that combines the path length and smoothness objectives is defined as: the immediate reward value obtained by the agent after performing an action, the numerical value of which is negatively related to the path segment length increased by the current action and negatively related to whether the current action introduces an inflection point.
[0054] S4.3, the policy is fine-tuned using the Q-learning algorithm, The value update formula is: ; wherein, and represent the current state and action, respectively; represent the next state after performing the action ; represents the maximum value of all possible actions in the next state ; ; ; ; represents the immediate reward; is the learning rate; is the discount factor.
[0055] Specifically, step S5 specifically includes the following steps: S5.1, the Q-learning adopts the ε-greedy strategy to balance exploration and utilization, and the decision-making process of the ε-greedy strategy is expressed as: ; wherein, , represents the exploration rate, indicating the probability of the agent randomly selecting an action for exploration; represents the action variable that maximizes the current value.
[0056] S5.2, when the value function change is less than the convergence threshold , that is: ; or the maximum number of training rounds is reached , the learning process is terminated. Wherein, is the number of iterations, is the value function after the th iteration.
[0057] Specifically, step S6 is specifically:
[0058] From the starting point Initially, the agent chooses the option with the greatest [potential] at each step. Value action : ; Until the target point is reached Generate the final path list .in, The first in the optimal path list Each node. Both of the following conditions must be met: 1. Shortest length: 2. Fewest inflection points: .
[0059] To verify the method of this invention, path planning was performed during the scanning operation of a workpiece with freeform surface features on an industrial robot equipped with a multi-line laser scanner. For example... Figure 2 As shown, a six-axis industrial robot 1 has a multi-line laser scanner 2 mounted on its end flange. A workpiece 4, with complex free-form surface features, is placed directly below it. A binocular tracking camera 3 is located on the right, used to identify reflective points on the multi-line laser scanner to determine its position and orientation in three-dimensional space. A computer 5 communicates with both the multi-line laser scanner and the industrial robot. The industrial robot moves according to a robot path planning method based on a hybrid optimization mechanism of artificial bee colony-reinforcement learning, obtaining the optimal path point set. Simultaneously, the multi-line laser scanner transmits the scanned data to the computer in real time for display, ultimately verifying the effectiveness of the invention.
[0060] First, the surface features of the workpiece are identified and divided using a surface sampling method, resulting in a set of small surface patches. Then, the center points of each small surface patch are obtained, such as... Figure 3 As shown, the normal vector of the point with respect to the workpiece surface is calculated; finally, based on the optimal scanning distance of the laser scanner, the final set of scanning viewpoints, i.e., the set of path points to be planned, is obtained, as shown. Figure 4 As shown.
[0061] The method provided by this invention processes the set of points to be planned to obtain the final path, such as... Figure 5 As shown. According to Figure 5 The generated scanning path ultimately yields the scanning results of the workpiece surface morphology, with significant effects and good reconstruction results, such as... Figure 6 As shown.
[0062] Example 2 A robot path planning system based on artificial bee colony and reinforcement learning includes: The population initialization module is configured to initialize a population of the artificial bee colony algorithm according to known environmental information, a starting point and a target point, wherein each individual in the population represents a feasible path from the starting point to the target point composed of path points.
[0063] The global search optimization module is configured to iteratively optimize the population based on a search mechanism of the artificial bee colony algorithm, wherein the quality of the individual is evaluated by a fitness function that comprehensively considers path length and path smoothness, and the population is updated accordingly.
[0064] The elite solution set construction module is configured to select a batch of high-quality solutions from the final population to form an elite solution set when the global search optimization module meets a preset optimization process termination condition.
[0065] The reinforcement learning initialization module integrates a reinforcement learning algorithm and is configured to encode the elite solution set as prior knowledge of a reinforcement learning agent, to initialize a value function, and to set a state space, an action space and a reward function that combines path length and smoothness objectives.
[0066] The policy fine-tuning module is configured to run the reinforcement learning algorithm to search for and optimize a policy based on the initialized value function, to fine-tune the path policy.
[0067] The path output module is configured to output the optimal path determined after convergence, which meets the requirements of shortest length and fewest turning points.
[0068] Of course, the above description is not a limitation on the present application, and the present application is not limited to the above examples. Changes, modifications, additions or replacements made by those skilled in the art within the scope of the present application should also be within the scope of the present application.
Claims
1. A robot path planning method based on artificial bee colony and reinforcement learning, characterized in that, Specifically, the steps include the following: S1. Initialize the artificial bee colony algorithm population based on the known environmental information, starting point, and target point; S2, based on the search mechanism of the artificial bee colony algorithm, iteratively optimizes the population; S3, when the preset optimization process termination condition is met, select a batch of high-quality solutions from the final population to form an elite solution set; S4 encodes the elite solution set into prior knowledge of the reinforcement learning agent to initialize the value function of reinforcement learning, and sets the state space, action space and reward function that integrates path length and smoothness objectives of reinforcement learning. S5, based on the initialized value function, runs a reinforcement learning algorithm to search for and optimize policies, so as to fine-tune the path policy; S6 outputs the optimal path after the reinforcement learning algorithm converges.
2. The robot path planning method based on artificial bee colony and reinforcement learning according to claim 1, characterized in that, Step S1 specifically includes the following steps: S1.1, each feasible path It consists of a series of node sequences, represented as: ; in, Starting point For the target point, For nodes in the environment; S1.2, use the artificial bee colony algorithm for global search; initialize the ABC population with a population size of... Each nectar source represents a path, containing The number of nodes; the encoding method for honey sources uses integer encoding, the first node... The initialization formula for a honey source is: ; in, nectar source The first in One node; The generation rules are as follows: ; in, It is the set of all nodes in the environment, excluding the start and end points; , This represents the total number of nodes. This means randomly selecting a node from the set of nodes.
3. A robot path planning method based on artificial bee colony and reinforcement learning according to claim 1, characterized in that, Step S2 specifically includes the following steps: S2.1, using the fitness function to comprehensively consider path length and smoothness, the fitness function... The expression is: ; in, The total length of the path is expressed by the following formula: ; in, Indicates Euclidean distance; Representing a path The number of inflection points; This is the smoothness penalty coefficient; S2.2, Smoothness is reflected by the number of inflection points. The number of inflection points is determined by calculating the included angle between the vectors of three consecutive points. Let three consecutive points... , , This forms two vectors. and The cosine of the angle between two vectors for: ; like ,but This is determined to be an inflection point; among them, For the preset threshold, The larger the value, the weaker the strength of the inflection point determination; S2.3, based on fitness function Neighborhood search and population update are performed; the formula for generating new nectar sources during the hired bee stage is: ; in, For new honey sources; This is the current honey source; The honey source is randomly selected from the area; It is a random number in the range [-1, 1].
4. A robot path planning method based on artificial bee colony and reinforcement learning according to claim 1, characterized in that, Step S3 is as follows: When the number of iterations of the ABC algorithm reaches the preset maximum value Or the optimal fitness of the population in a continuous Intra-generation improvement less than threshold When the time is right, terminate the ABC algorithm search; select the population with the highest fitness from the final population. A set of non-repeating paths constitutes the elite solution set. The screening formula is: ; in, For fitness threshold, For the first One path.
5. A robot path planning method based on artificial bee colony and reinforcement learning according to claim 1, characterized in that, Step S4 specifically includes the following steps: S4.1, for each path in the elite solution set traverse each path State-Action Pairs ,in, Current node state, action To select the next node; the state-action pair The value is initialized to a priority value. : ; in, The action value function; S4.2, Setting up the reinforcement learning environment, state space The set of all nodes; action space From state The set of reachable neighboring nodes; the reward function is a fusion of path length and smoothness objectives. Designed as follows: ; in, The distance traveled; For indicator functions, if in If an inflection point is found, return 1; otherwise, return 0. Penalty weight for inflection point; S4.3 uses the Q-learning algorithm for policy fine-tuning. The value update formula is: ; in, and These represent the current state and action, respectively. Represents the execution of actions The next state after that; Indicates the next state Below, one of the action variables among all possible actions that the agent can perform; Represents the next state Of all possible actions The largest value; Indicates an immediate reward; The learning rate; This is the discount factor.
6. A robot path planning method based on artificial bee colony and reinforcement learning according to claim 1, characterized in that, Step S5 specifically includes the following steps: S5.1, Q-learning employs an ε-greedy strategy to balance exploration and exploitation. The decision-making process of the ε-greedy strategy is described as follows: ; in, , representing the exploration rate; Indicates that the current The action variable with the largest value; S5.2, when The change in the value function is less than the convergence threshold. ,Right now: ; Or reach the maximum number of training rounds When the learning process is terminated, the learning process is terminated; among them, For the number of iterations, For the first After the next iteration Value function.
7. A robot path planning method based on artificial bee colony and reinforcement learning according to claim 1, characterized in that, Step S6 is as follows: From the starting point Initially, the agent chooses the option with the greatest [potential] at each step. Value action : ; Until the target point is reached Generate the final path list ;in, The first in the optimal path list Each node.
8. A robot path planning system based on artificial bee colony and reinforcement learning, utilizing the method described in any one of claims 1-7, characterized in that, include: The population initialization module is used to initialize the artificial bee colony algorithm population based on known environmental information, starting point and target point. Each individual in the population represents a feasible path from the starting point to the target point, which is composed of path points. The global search optimization module is used to iteratively optimize the population based on the search mechanism of the artificial bee colony algorithm. The quality of individuals is evaluated by a fitness function that comprehensively considers path length and path smoothness, and the population update is guided accordingly. The elite solution set construction module is used to select a batch of high-quality solutions from the final population to form an elite solution set when the global search optimization module meets the preset optimization process termination conditions. The reinforcement learning initialization module integrates reinforcement learning algorithms to encode the elite solution set into prior knowledge of the reinforcement learning agent, which is used to initialize the value function and set the state space, action space, and reward function that combines path length and smoothness objectives. The policy fine-tuning module is used to run reinforcement learning algorithms to search for and optimize policies based on the initialized value function, so as to fine-tune the path policy. The path output module is used to output the optimal path after convergence. The optimal path satisfies the requirements of shortest length and fewest inflection points.
Citation Information
Patent Citations
Multi robot path planning method based on multi-target artificial bee colony algorithm
CN104808665A
Mobile robot path planning method based on reinforcement learning
CN110794832A
Mobile robot path planning method based on improved artificial bee colony algorithm
CN114662638A
Flexible job shop batch scheduling optimization method and device and electronic equipment
CN115292950A
Intelligent path planning method in dynamic environment
CN115373393A
Cited By
Scheduling optimization method for hybrid flow shop with automated guided vehicle
CN122172755A
Multi-machine scheduling method, device and equipment for train inspection robot based on genetic algorithm and reinforcement learning and medium
CN122264489A