A path planning method based on machine learning in a degradable environment
By combining A* and DQN algorithms, semantic segmentation and fuzzy evaluation technology are used to solve the problem of performance degradation of lidar in a degradable environment, and more efficient and safe path planning is achieved.
Patent Information
- Application Number
- CN202510330342.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In a degraded environment, the performance degradation of lidar results in a decrease in perceptual accuracy and reliability, affecting the accuracy and safety of path planning.
Combining the A* algorithm and the DQN algorithm, the A* algorithm is used to quickly generate initial paths, providing prior information for DQN, and reducing training time and computing resource consumption. At the same time, semantic segmentation technology and fuzzy evaluation method were introduced to evaluate environmental areas and increase the weight of degraded areas, and guide the robot to drive in information-rich areas.
It improves the accuracy and safety of the robot's path planning in a degradable environment, enhances the robustness and practicality of the system, and reduces the risk of path planning failure.
Smart Images

Figure CN119845286B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of path planning, machine learning, and in particular to a path planning method based on machine learning in a degradable environment. Background Art
[0002] In recent years, mobile robots are gradually becoming part of people's daily lives. They integrate multiple intelligent functions such as environmental perception, decision-making and planning, and behavior control, and have shown great application potential in scenarios such as hotels and factories. For example, in hotels, robots can take on tasks such as customer guidance and item delivery; in factories, robots can replace repetitive labor, thereby significantly improving production efficiency. As mobile robot technology continues to mature, its key role in future social services will become increasingly prominent, and will have a profound impact on improving quality of life and work efficiency.
[0003] At the same time, path planning, as one of the core technologies of robot autonomous navigation, plays an important role in the fields of autonomous driving, service robots and automatic guided vehicles. Existing path planning algorithms mainly rely on the real-time perception of environmental geometry and dynamic obstacles, usually through devices such as laser radar, visual sensors or inertial units to obtain environmental information and calculate the optimal path from the starting point to the target point.
[0004] However, in certain environments such as long straight corridors and vast plains, the performance of LiDAR will be significantly degraded, affecting its perception accuracy and reliability. This degradation is mainly due to the singleness of environmental features. In open areas or areas with sparse features, the point cloud data collected by LiDAR is relatively sparse, and it is difficult to effectively extract sufficient environmental features, resulting in a decrease in positioning accuracy. Especially in degraded areas, due to the insufficient point cloud density, the accurate identification and matching of environmental feature points becomes extremely difficult, which further has a negative impact on positioning and map construction. In addition, in scenes lacking structural features, the information obtained by LiDAR is highly similar and lacks directional constraints. This characteristic not only reduces the accuracy and completeness of map construction, but may also cause the path planning algorithm to misjudge obstacles, significantly increasing the risk of path planning failure.
[0005] In the existing technology, the A* algorithm, as a classic heuristic path planning method, has demonstrated good performance in a variety of complex environments and can effectively handle multi-obstacle and dynamic obstacle scenarios. However, when faced with changes in environmental data caused by lidar degradation, the performance of the A* algorithm will be significantly limited. In contrast, the Deep Q-Network (DQN) algorithm can continuously update strategies by interacting with the environment, thereby better adapting to the changing environment. However, the DQN algorithm takes a long time to train, and it is difficult to quickly generate the optimal path in a complex environment, which limits its application in scenarios with high real-time requirements. Summary of the invention
[0006] Aiming at the degradation problem encountered by laser radar in an easily degraded environment, the present invention proposes an innovative path planning method to prevent the mobile robot from entering the degraded area, so as to improve the autonomous planning performance of the mobile machine in the degraded environment. Specifically, the present method combines the traditional A* algorithm and the DQN algorithm, and uses the A* algorithm to quickly generate the initial path, providing reliable prior information for DQN, thereby significantly reducing the training time and computing resource consumption. The DQN algorithm is used to achieve local decision optimization, so that the robot can adapt to local environmental changes in real time. At the same time, the present invention introduces semantic segmentation technology to divide the environment into degraded areas and non-degraded areas, and evaluates these areas in combination with the fuzzy evaluation method. By increasing the weight of the degraded area, the robot is further guided to travel in the information-rich area. The improved algorithm not only improves the path planning accuracy and safety of the robot in an environment with uneven laser radar perception and positioning, but also enhances the overall robustness and practicality of the system.
[0007] The specific implementation steps are as follows:
[0008] Implementation step 1: point cloud data acquisition and information map generation.
[0009] Step 1.1: LiDAR point cloud data acquisition and preprocessing.
[0010] Use LiDAR to perform omnidirectional scanning in the robot's moving environment, obtain point cloud images, and extract obstacle distribution and terrain information. De-noise and downsample the point cloud data to improve data quality.
[0011] Step 1.2: Semantic segmentation and information map generation.
[0012] Step 1: Establish factor set and evaluation set. Evaluate the impact of different semantic classes on robot positioning and determine the weights of different semantic classes. First, establish factor set , where the elements Represents the i-th factor affecting lidar positioning, and establishes an evaluation set , where the elements Represents the jth evaluation level.
[0013] Step 2: Fuzzy comprehensive evaluation. For a single element , we perform single element evaluation based on actual observations or expert knowledge and obtain a fuzzy evaluation vector:
[0014] ,
[0015] in Indicated in factors The evaluation results are The degree of membership;
[0016] All single-element evaluation vectors are formed as rows of a matrix Fuzzy comprehensive evaluation matrix Specifically:
[0017] ,
[0018] Step 3: Calculate the weight vector. Use the hierarchical analysis method to calculate the weight and form the weight vector :
[0019] ,
[0020] in, Represents the relative weight of the i-th element in the comprehensive evaluation.
[0021] Step 4: Fuzzy comprehensive evaluation calculation. The fuzzy weight of each factor and fuzzy evaluation matrix Combined, we get the comprehensive fuzzy evaluation vector :
[0022] ,
[0023] in, represents standard matrix multiplication, Indicates that in the evaluation set The comprehensive membership of each evaluation level.
[0024] Step 5: Determine the numerical evaluation. According to the comprehensive fuzzy evaluation vector , determine the specific numerical evaluation (Evaluation score) is:
[0025] ,
[0026] in, is the fuzzy evaluation vector The jth element in For evaluation set Each rating level in The value of the mapping, q is the evaluation set The total number of elements in .
[0027] Step 6: Create an information map. Use real-time semantic segmentation technology to dynamically perceive and classify the environment and generate a 2D semantic projection map. Define a sliding window of fixed size with the robot's current position as the center, and record the center position of each grid in the sliding window. , calculate and store the evaluation score , forming an information map for subsequent path planning.
[0028] Implementation step 2: Global path planning.
[0029] Step 2.1: Initialize the grid map and parameters. Create a grid map, set the starting point and target point, and initialize the open list and closed list.
[0030] Step 2.2: Set the starting point and calculate the initial cost. Add the starting point to the open list and set the initial cost value for it. The A* cost function is:
[0031] ,
[0032] in, is the cumulative cost from the starting point to the current node, is the heuristic estimate of the distance from the current node to the target point, both of which are calculated using the Euclidean distance.
[0033] Step 2.3: Node expansion and search: Use the A* algorithm to search for the optimal path through node expansion and heuristic cost function.
[0034] Step 2.4: Check the target node. Determine whether the current node n is the target node. If so, trace back to the parent node from the target node step by step to build the optimal path from the starting point to the end point. If the current node n is not the target node, continue to execute step 2.3 until the target point is found.
[0035] Step 2.5: Path backtracking and output. After finding the target node, reconstruct the optimal path from the starting point to the end point by backtracking the parent node. The planning result is a series of grid points, which constitute a path set , specifically expressed as:
[0036] ,
[0037] in, is the two-dimensional coordinate of the kth point.
[0038] Implementation step 3: Design reward function. The reward function design mainly focuses on safety, shortest path, and efficiency to ensure that the agent can complete the task efficiently and safely in a degraded environment.
[0039] Step 3.1: Design of reward function considering safety. In order to ensure that the robot avoids degradation areas in a degradation-prone environment, a reward function based on evaluation scores is designed:
[0040] ,
[0041] in, is the safety reward function, is the risk penalty coefficient, is the parameter of the penalty growth rate, is the evaluation score.
[0042] Step 3.2: Consider the design of the reward function with the shortest path. When the current position of the agent becomes farther from the target position, a negative reward is obtained. When the current position of the agent becomes closer to the target position, a certain degree of reward is obtained. The closer the distance, the higher the reward. In other cases, there is no reward:
[0043] ,
[0044] in, is the path reward function, is the current position of the agent To the target location The distance is the maximum distance from the agent to the target position, It is the positive amplification factor, which is used to adjust the reward amplitude.
[0045] Step 3.3: Design of reward function considering efficiency. In order to prevent the agent from wandering in the local area during the movement, the agent is encouraged to reach the target as soon as possible, which is specifically expressed as:
[0046] ,
[0047] in, is the efficiency reward function, is the penalty coefficient (positive coefficient), is an indicator function. When the position of the tth step is the same as that of the previous k steps, the indicator function returns 1, otherwise it returns 0. is the position of the robot at step t.
[0048] Taking safety, shortest path and efficiency into consideration, the agent should take safety as the primary goal when planning the path, while taking into account the shortest path and planning efficiency. According to actual needs and expert experience, different importance weights are given to the safety, path and efficiency reward functions, among which safety has the highest weight, path is second, and efficiency is the lowest. The final optimized reward function can be expressed as:
[0049] ,
[0050] Where R is the reward function, , , They are safety, path and efficiency reward functions, is the weight, satisfying and .
[0051] Implementation step 4: Design loss function. The DQN loss function is used to measure the gap between the predicted Q value of the current network and the target Q value, thereby improving the quality of the strategy.
[0052] L ( i ) = 1 N ∑ i = 1 N F ⋅ [ R i + c max a ′ Q t arg and ( s ′ i , a ′ ; i − ) − Q ( s i , a i ; i ) ] 2 ,
[0053] in is the total loss function, N represents the number of experience samples, It’s an instant reward. is the discount factor, is the target network parameter, is the current network parameter, is the maximum expected return, is the current network for the sample The predicted Q value of .
[0054] The loss function updates its parameters through the gradient descent algorithm, so that the predicted Q value continues to approach the target Q value and improve the quality of the strategy.
[0055] Implementation step 5: Model training. During the model training phase, the agent selects the optimal action based on the current state and continuously adjusts the parameters of the Q network based on environmental feedback, so that the path planning and obstacle avoidance behaviors are closer to the optimal strategy in the actual scenario. In order to improve the stability and efficiency of training, the experience replay mechanism is used to store interaction data, and batch sampling is used to reduce the correlation between data. In addition, the regular synchronization operation of the target network effectively reduces gradient oscillations and ensures the convergence and reliability of the DQN model. After continuous optimization, DQN can generate the optimal strategy that takes into account both global planning and local obstacle avoidance needs, enabling the agent to achieve efficient and safe autonomous planning in a degraded environment. The specific training steps are as follows:
[0056] Step 1: State initialization and path planning. Build an information map and use the A* algorithm to plan the global path and define the state of the agent. , which includes the following: environmental information collected by the LiDAR, information map, global path set obtained by A* planning, and discrete action space (forward, backward, left turn, right turn, etc.). At the same time, the DQN network parameters and experience replay pool are initialized.
[0057] Step 2: Experience collection. The agent explores the environment, collects interaction data such as state transitions and rewards, and stores the data in the experience replay pool.
[0058] Step 3: Experience replay. Randomly sample data from the replay pool to build training batches, reduce data correlation, and improve training stability.
[0059] Step 4: Calculate the target Q value. Using the target network (parameters are ) to calculate the target value to ensure that the supervision signal is more stable and help update the subsequent network parameters. The target value calculation formula is:
[0060] ,
[0061] in, is the target value, It’s an instant reward. is the discount factor, are the parameters of the target network.
[0062] Step 5: Optimize network parameters. Calculate the difference between the predicted Q value and the target Q value based on the loss function, and use the gradient descent algorithm to update the behavior network parameters. .
[0063] Step 6: Update the target network and adjust the strategy. Regularly synchronize the behavior network parameters to the target network to reduce gradient oscillation during training and further improve the optimization efficiency of the strategy.
[0064] Implementation step 6: Final path fusion and generation. By matching the current state with the nearest node in the global path and smoothly connecting the dynamically generated local path segments, a continuous, smooth and safe global path is finally generated.
[0065] The present invention proposes an innovative solution to the problem of path planning of intelligent agents in degraded environments, aiming to achieve safe and efficient path planning of intelligent agents in degraded environments. Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0066] 1. The present invention uses the DQN network combined with the evaluation scores in the information map to make judgments, so that the intelligent agent can effectively avoid degraded areas, significantly improving the safety of path planning;
[0067] 2. The present invention combines A* and DQN algorithms to reduce training time and computing resource consumption, and improves the efficiency of path planning of intelligent agents in easily degraded environments;
[0068] 3. The path planning method designed by the present invention is highly scalable and applicable, and can enable the intelligent agent to autonomously plan a safe and optimal path in a degradable environment. This method has significant application value and practical benefits in the fields of transportation and service industries.
[0069] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 This is the overall flow chart of the path planning method based on machine learning;
[0071] Figure 2 It is the overall framework of the DQN algorithm. DETAILED DESCRIPTION
[0072] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0073] Attached Figure 1 It is the overall flow chart of the present invention. This example provides a path planning method based on machine learning in a degraded environment, which specifically includes the following processes: point cloud data acquisition and information map generation, global path planning, reward function design, cost function design, model training, and final path fusion and generation.
[0074] Attached Figure 2 The process of the DQN algorithm includes four main parts: the external environment, the current network, the target network, and the experience replay pool. These parts work together to realize the learning and decision-making process of the DQN algorithm. The key steps are: the agent receives the current state from the external environment , and selects an action a based on the current network. After executing the action, the agent receives a reward r and observes the next state ,experience It is stored in the experience replay pool. During the training phase, a batch of experience samples are randomly drawn from the experience replay pool and used to calculate the loss function, which measures the predicted Q value of the current network. and target Q value The gap between them is calculated by minimizing the loss function and using the gradient descent method to update the parameters of the current network. The parameters of the current network are periodically soft-updated to the target network to maintain the stability of the target network.
[0075] For ease of understanding, specific examples are given for easily degradable environments and robots. Taking an automated warehouse handling robot as an example, the easily degradable environment is composed of symmetrical long straight shelves, etc., and the robot is a handling robot.
[0076] The specific implementation steps are as follows:
[0077] Implementation step 1: Obtain point cloud data and create an information map.
[0078] Step 1.1: LiDAR point cloud data acquisition and preprocessing.
[0079] Use LiDAR to perform omnidirectional scanning in an automated warehouse environment, obtain point cloud images, and extract obstacle distribution and terrain information. De-noise and downsample the point cloud data to improve data quality.
[0080] Step 1.2: Semantic segmentation and information map generation.
[0081] Step 1: Establish factor set and evaluation set. Quantitatively evaluate the impact of different semantic classes on robot positioning performance and determine the weights of different semantic classes. First, establish factor set , where the elements Represents the i-th factor affecting lidar positioning, and establishes an evaluation set , where the elements Represents the jth evaluation level.
[0082] Step 2: Fuzzy comprehensive evaluation. For a single element We perform single-element evaluation based on actual observations or expert knowledge and obtain a fuzzy evaluation vector:
[0083] ,
[0084] in, Indicated in factors The evaluation results are The degree of membership;
[0085] All single-element evaluation vectors are formed as rows of a matrix Fuzzy comprehensive evaluation matrix Specifically:
[0086] ,
[0087] Step 3: Calculate the weight vector. Use the hierarchical analysis method to calculate the weight and form the weight vector for:
[0088] ,
[0089] in, Represents the relative weight of the i-th factor in the comprehensive evaluation.
[0090] Step 4: Fuzzy comprehensive evaluation. Single factor evaluation matrix Combined, the comprehensive fuzzy evaluation vector is obtained :
[0091] ,
[0092] in, represents standard matrix multiplication, Indicates that in the evaluation set The comprehensive membership of the jth evaluation level.
[0093] Step 5: Determine the numerical evaluation. According to the comprehensive fuzzy evaluation vector , determine the specific numerical evaluation (Evaluation score) is:
[0094] ,
[0095] in, is the fuzzy evaluation vector The jth element in For evaluation set Each rating level in The value of the mapping, q is the evaluation set The total number of elements in .
[0096] Step 6: Create an information map. Use real-time semantic segmentation technology to dynamically perceive and classify the environment and generate a 2D semantic projection map. Define a sliding window of fixed size with the robot's current position as the center, and record the center position of each grid in the sliding window. , calculate and store the comprehensive score , forming an information map for subsequent path planning.
[0097] Implementation step 2: The robot performs overall path planning based on the map and uses the A* search algorithm to calculate the optimal path.
[0098] Step 2.1: Initialize the grid map and parameters. Create a grid map, set the starting point and target point, and initialize the open list and closed list.
[0099] Step 2.2: Set the starting point and calculate the initial cost. Add the starting point to the open list and set the initial cost value for it. The A* cost function is:
[0100] ,
[0101] in, The cumulative cost from the starting point to the current node is the cumulative distance, is the heuristic estimate of the distance from the current node to the target point, both of which are calculated using the Euclidean distance.
[0102] Step 2.3: Node expansion and search: Use the A* algorithm to search for the optimal path through node expansion and heuristic cost function.
[0103] Step 2.4: Check the target node. Determine whether the current node n is the target node. If so, trace back to the parent node from the target node step by step to build the optimal path. If the current node n is not the target node, continue to execute step 2.3 until the target point is found.
[0104] Step 2.5: Path backtracking and output. After finding the target node, reconstruct the optimal path from the starting point to the end point by backtracking the parent node. The planning result is a series of grid points, which constitute the path set output represented as:
[0105] ,
[0106] in, is the two-dimensional coordinate of the kth point.
[0107] Implementation step 3: Design reward function. The reward function design mainly focuses on safety, shortest path, and efficiency to ensure that the agent can complete the task efficiently and safely in a degraded environment.
[0108] Step 3.1: Consider the design of reward function for safety. In order to ensure that the robot avoids degradation areas in a degradation-prone environment, a penalty function based on risk score is designed:
[0109] ,
[0110] in, is the safety reward function, is the risk penalty coefficient, is the parameter of the penalty growth rate, is the evaluation score.
[0111] Step 3.2: Consider the design of the reward function with the shortest path. When the current position of the robot becomes farther from the target position, a negative reward is obtained. When the current position of the robot becomes closer to the target position, a certain degree of reward is obtained. The closer the distance, the higher the reward. There is no reward in other cases:
[0112] ,
[0113] in, is the path reward function, is the current position of the agent To the target location The distance is the maximum distance from the agent to the target position, It is the positive amplification factor, which is used to adjust the reward amplitude.
[0114] Step 3.3: Design of reward function considering efficiency. In order to prevent the agent from wandering in the local area during the movement, the agent is encouraged to reach the target as soon as possible, which is specifically expressed as:
[0115] ,
[0116] in, is the efficiency reward function, is the positive penalty coefficient, is an indicator function. When the position of the tth step is the same as that of the previous k steps, the indicator function returns 1, otherwise it returns 0. is the position of the robot at step t.
[0117] Taking safety, shortest path and efficiency into consideration, the agent takes safety as the primary goal, while taking into account task completion and efficiency requirements. According to actual needs and expert experience, we directly assign different importance weights to the safety, path and efficiency reward functions, among which safety has the highest weight, path is second, and efficiency has the lowest. The final optimized reward function is expressed as:
[0118] ,
[0119] Where R is the reward function, , , They are safety, goal and efficiency reward functions, is the weight, satisfying and .
[0120] Implementation step 4: Design loss function. The DQN loss function is used to measure the gap between the predicted Q value of the current network and the target Q value, thereby improving the quality of the strategy.
[0121] L ( i ) = 1 N ∑ i = 1 N F ⋅ [ R i + c max a ′ Q t arg and ( s ′ i , a ′ ; i − ) − Q ( s i , a i ; i ) ] 2 ,
[0122] in is the total loss function, N represents the number of experience samples, It’s an instant reward. is the discount factor, is the target network parameter, is the current network parameter, is the maximum expected return, is the current network for the sample The predicted Q value of .
[0123] The loss function updates its parameters through the gradient descent algorithm, so that the predicted Q value continues to approach the target Q value and improve the quality of the strategy.
[0124] Implementation step 5: Model training. During the model training phase, the robot selects the optimal action based on the current state, and continuously adjusts the Q network parameters based on environmental feedback, so that the path planning and obstacle avoidance behaviors are closer to the optimal strategy in the actual scenario. To improve the stability and efficiency of training, the system uses an experience replay mechanism to store interaction data and reduces the correlation between data through batch sampling. In addition, the regular synchronization operation of the target network effectively reduces gradient oscillations and ensures the convergence and reliability of the DQN model. After continuous optimization, DQN can generate a navigation strategy that takes into account both global planning and local obstacle avoidance needs, enabling the robot to achieve efficient and safe autonomous driving in a degraded environment. The specific training steps are as follows:
[0125] Step 1: Define the state of the agent by building an information map and planning the global path using the A* algorithm , which includes the following: environmental information collected by the lidar, information map, global path set and discrete action space (forward, backward, left turn, right turn, etc.) obtained by A* planning, and initializes the DQN network parameters and experience replay pool.
[0126] Step 2: Experience collection: The agent explores the environment, collects interaction data such as state transfer and rewards, and stores the data in the experience replay pool.
[0127] Step 3: Experience replay, construct training batches by randomly sampling data in the replay pool, reducing data correlation and improving training stability.
[0128] Step 4: Calculate the target Q value and use the target network (parameters are ) calculates the target value to ensure that the supervision signal is more stable, which helps to update the subsequent network parameters.
[0129] Step 5: Optimize network parameters, calculate the gap between the predicted Q value and the target Q value based on the loss function, and use the gradient descent algorithm to update the behavior network parameters .
[0130] Step 6: Update the target network and adjust the strategy. Regularly synchronize the behavior network parameters to the target network to reduce gradient oscillation during training. At the same time, dynamically adjust the exploration rate so that the robot can balance exploration and utilization during training, further improving the optimization efficiency of the strategy.
[0131] Implementation step 6: Path fusion and generation. By matching the current position to the nearest node in the global path and smoothly connecting the local paths generated by DQN, a continuous, smooth and safe path is finally generated.
Claims
1. A path planning method based on machine learning in a degradable environment, characterized in that: The following steps are involved: Step S1: LiDAR point cloud data processing and information map generation: Step S11: Obtaining point cloud data of the laser radar; Step S12: performing denoising and downsampling preprocessing on the point cloud data; Step S13: using the trained semantic segmentation model to perform semantic segmentation on the preprocessed point cloud data. The data for model training is generated by manual annotation. The semantic segmentation model divides the data into degraded areas and non-degraded areas. The degraded areas refer to areas where the lidar perception capability is limited, and the non-degraded areas refer to areas where the lidar perception accuracy is normal. Step S14: Score the semantic segmentation results based on the fuzzy evaluation method, calculate the weight of each semantic class through the hierarchical analysis method, generate the evaluation score in combination with the fuzzy comprehensive evaluation matrix, and map the scoring results to the grid map to form an information map; Step S2: Global path planning: Step S21: Create a grid map based on the information map, set the starting point and the target point, and initialize the open list and the closed list; Step S22: Use the A* algorithm to perform path planning, and its cost function is defined as: in, The cumulative cost from the starting point to the current node is the cumulative distance, is the heuristic estimate from the current node to the target point, both of which are calculated by the Euclidean distance; Step S23: Select the node with the lowest cost according to the cost function value to expand until the target point is found, and backtrack to generate a path set consisting of two-dimensional coordinate points. , expressed as: in, It is The two-dimensional coordinates of a point; Step S3: Local path planning: Step S31: Determine the global path set based on the information map constructed in step S1 Whether each node in is in the degenerate region; Step S32: For a node in a degenerate region, select the next node that is adjacent to the node on the global path and is in a non-degenerate region as a local target point; Step S33: Generate a local path using the DQN algorithm, whose state input includes the lidar environment information, information map, global path set and evaluation score, design a reward function based on safety, shortest path and efficiency, calculate the loss function, train the experience replay pool, and output the local path; Step S4: Final path fusion and generation: Step S41: Match the nearest node in the global path set according to the current position of the agent; Step S42: smoothly connect the local path generated by DQN in step S3 with the global path to form a continuous, smooth and safe optimal path.
2. The path planning method based on machine learning in a degradable environment according to claim 1, characterized in that: The scoring of the semantic segmentation result based on the fuzzy evaluation method described in step S14 specifically includes the following steps: Step 1: Establish factor set , where the element represents the i-th factor affecting the laser radar positioning, and the evaluation set is established , where the element represents the jth evaluation level; Step 2: Perform single-element evaluation on a single factor based on expert knowledge to generate a fuzzy evaluation vector; Step 3: Use the hierarchical analysis method to calculate the weight of each factor and form a weight vector ; Step 4: Combine the weight vector with the fuzzy evaluation matrix to obtain the comprehensive fuzzy evaluation vector ; Step 5: Determine the numerical score F based on the comprehensive fuzzy evaluation vector and evaluate the risk level of each semantic class.
3. The path planning method based on machine learning in a degradable environment according to claim 1, characterized in that: The reward function design described in step S33 specifically includes the following steps: The reward function design for security is as follows: in, is the safety reward function, is the risk penalty coefficient, is the penalty growth rate parameter, and F is the evaluation score; The reward function mechanism design for the shortest path is as follows: in, is the path reward function, is the current location To the target location The distance is the maximum distance to the target location, is the positive amplification factor, which adjusts the reward amplitude; The design of the efficient reward function mechanism is as follows: in, is the efficiency reward function, is the penalty coefficient and is a positive coefficient, is the indicator function, when Step and forward If the step positions are the same, the indicator function returns 1, otherwise it returns 0. For the The position of the step, T is the total number of steps; Among all the reward functions designed above, when the three reward functions conflict, different reward coefficients are designed to determine the priority of the reward function. The final optimized reward function is expressed as: in, is the reward function, , , They are safety, path and efficiency reward functions, is the weight, satisfying and .
Citation Information
Patent Citations
Robot fleet management and additive manufacturing for value chain networks
AU2021401816A1
Motion planning method based on machine learning in complex environment
CN116551703A