Fusion type robot path planning method based on channel guidance

By introducing a channel-guided TRD structure integrating RRT-star and DRL in robot path planning, combining greedy algorithms and inverse solution modules of neural networks, the problem of low efficiency of robot path planning in the existing technology is solved, and the effect of efficiently optimizing paths in complex industrial application scenarios is achieved.

CN120141516APending Publication Date: 2025-06-13DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510289032.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing robot path planning methods are difficult to efficiently optimize paths in complex industrial application scenarios, traditional methods are difficult to utilize past experience, and neural network-based methods are less efficient in exploring in complex scenarios.

Method used

The fusion robot path planning method based on channel guidance is adopted, combined with the TRD structure of RRT-star and DRL, and trained through channel rewards, endpoint distance rewards and collision rewards, optimized the robot path, and converted into the path of the robot joint space through the greedy algorithm and the inverse solution module of the neural network.

Benefits of technology

It improves the efficiency of robot path planning, can quickly find feasible paths in complex environments, reduce collisions with obstacles, optimize path length, combines efficient global exploration and local optimization, and surpasses the performance of a single algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120141516A_ABST
    Figure CN120141516A_ABST
Patent Text Reader

Abstract

According to the fusion type robot path planning method based on channel guidance, an RRT-star module provides a feasible path to guide DRL early-stage exploration, a DRL module for channel guidance encourages a robot to move along a channel and guides the robot to bypass an obstacle, the path length of a robot operation space is optimized, and the path length of the robot operation space is optimized; an inverse solution module fusing a greedy algorithm and a neural network calculates loss according to the target pose and the current pose predicted by the neural network, the input of the neural network is adjusted in the gradient descent direction through a greedy strategy, and when the neural network tries to adjust the near-end joint of the robot to be free of fruit, the neural network is started; and then trying to adjust the near-end joint and the adjacent joints of the robot until the path of the robot operation space is converted into the path of the joint space, so that the robot path planning task is completed. The RRT-star module overcomes the problem that pure reinforcement learning is low in exploration efficiency in a complex environment, the channel-guided DRL module reduces collision between the robot and an obstacle, the inversion solution module fusing the greedy algorithm and the neural network expands and adjusts adjacent joints when the near-end joints are preferentially adjusted and necessary, and calculation redundancy is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and robot control, and particularly relates to a fusion robot path planning method based on channel guidance. Background Art

[0002] With the continuous development and maturity of robot technology, its application in the industrial field has become increasingly extensive. Compared with human labor, industrial robots do not feel fatigue or lose interest when facing cumbersome work tasks, and greatly improve the accuracy and work efficiency of operations, thus effectively promoting the development of enterprises and social productivity. Therefore, robots have developed rapidly all over the world and in various industries. With the advent of high-tech such as artificial intelligence, people are no longer limited to having robots perform single mechanized operations, but hope to endow robots with a certain degree of "intelligence" so that they can perform operations without human instructions, perceive the environment, and then make decisions through autonomous learning, so that robots can adapt to complex unstructured environments and continuously improve work efficiency. Among them, reinforcement learning, as a branch of the field of machine learning, can prompt robots to interact with the environment and obtain the decision with the maximum cumulative reward through a reward and punishment mechanism, so as to achieve system optimization.

[0003] In the prior art, traditional robot path planning methods are difficult to utilize past experience, and the generated paths are difficult to be continuously optimized; the robot path planning method based on neural networks has low exploration efficiency when facing complex scenarios. Therefore, there is an urgent need for a method to improve the efficiency of robot path planning to adapt to complex industrial application scenarios. Summary of the Invention

[0004] The technical problem to be solved by the technical solution of the present invention is: how to combine the advantages of traditional robot path planning methods and robot path planning methods based on neural networks to improve the efficiency of robot path planning.

[0005] To solve the above technical problem, the technical solution of the present invention provides a fusion robot path planning method based on channel guidance, including the following steps:

[0006] Create a digital virtual system of the robot operation space according to the shape and pose information of the robot and obstacles.

[0007] In the robot operation space of the digital virtual system of the robot operation space, plan a feasible path from the starting position P start of the robot to the ending position P end and define it as a plurality of paths F c , and use the node position P that is closest to the current position P curr of the robotshor Construct a channel centered at and with a radius of u, in combination with P curr , P shor , u and the end position P end , and train with channel rewards, end - point distance rewards, and collision rewards to obtain an optimized single path X.

[0008] Define a robot parameter model based on the joint angles Θ, poses T, and pose tolerance error δ of the robot kin According to the center point c of each joint i of the robot i , the axial unit vector and the half - side length define the OBB parameters, disassemble the optimized single path X to obtain target positions and target postures in multiple robot operating spaces, define the target pose matrix of the current node according to the target positions and target postures, and generate positive step - size candidate inputs for each dimension of an arbitrary random neural network input set K and negative step - size candidate inputs

[0009] Assume that the first layer only allows adjusting the collision - proximal joints, and respectively combine the positive step - size candidate input and the negative step - size candidate input to solve for a collision - free inverse solution with the initial connection weights of a random neural network.

[0010] If a collision - free inverse solution is not obtained, enter the second layer, reduce the learning rate of the neural network according to the exponential decay rule. The second layer allows adjusting the 2 joints adjacent to the collision - proximal joint as the center, and again respectively combine the positive step - size candidate input and the negative step - size candidate input to solve for a collision - free inverse solution with the initial connection weights.

[0011] According to the robot parameter model, convert the predicted joint angles corresponding to the positive step - size candidate input and the negative step - size candidate input into the poses in the robot operating space corresponding to the positive step - size candidate input and the negative step - size candidate input .

[0012] According to the poses in the robot operating space corresponding to the positive step - size candidate input and the negative step - size candidate input , respectively combine the target positions and target postures, calculate the position loss and the posture loss. According to the energy consumption coefficient matrix W, the moment of inertia l i , the joint angular velocity w i , the friction coefficient μ i , the mass m iand linear velocity v i , calculate the energy loss, and calculate the safety threshold loss based on the collision joint set, the signed distance from the obstacle point p to the OBB parameters and the safety threshold.

[0013] Assign the position loss, attitude loss, energy loss and safety threshold loss to the corresponding weight coefficients to obtain the positive step size candidate input and negative step size candidate input The corresponding overall loss function, between two candidate inputs and The one with smaller loss is selected as the new input, and the initial connection weights of the neural network are updated along the direction of the gradient descent of the overall loss function.

[0014] Update positive step size candidate input and negative step size candidate input Iterate again using the new input and the updated initial connection weights until the value of the overall loss function is less than the loss function threshold loss_threshold or the number of iterations ANN_step is equal to the iteration threshold ANN_step_threshold, and the inverse solution of the current node is obtained.

[0015] Continue to solve the inverse solution of the next node of the current node, and iterate again until all nodes in a single path X are inversely solved. At this time, the inverse solutions corresponding to all nodes are obtained, and finally the path Θ of the robot joint space is formed.

[0016] Preferably, according to the current position P of the robot curr , distance from current position P curr The nearest node position P shor And channel radius u, construct the channel reward function for training and optimizing the optimized path X, the formula is as follows:

[0017]

[0018] Among them, dis1 represents P curr and P shor and the distance between them, A and B are positive constants used to adjust the maximum range of channel rewards in each round.

[0019] Preferably, according to the current position P of the robot curr , end position P end And the channel radius u, construct the terminal distance reward function for training and optimizing the optimized path X, the formula is as follows:

[0020]

[0021] Among them, dis2 represents P curr and Pend The distance between, where C is a normal constant used to adjust the maximum range of the end - point distance reward in each round.

[0022] Preferably, according to the current position P of the robot curr , a collision reward function for training and optimizing to obtain the optimized path X is constructed, and the formula is as follows:

[0023]

[0024] where E is a negative constant.

[0025] The technical solution of the present invention proposes a channel - guided fusion robot path planning method, which uses the channel - guided fusion RRT - star and DRL's TRD structure to plan the path of the robot's operating space. The RRT - star module provides a feasible path to guide DRL for preliminary exploration. The channel - guided DRL module encourages the robot to move along the channel, guides the robot to bypass obstacles, and further optimizes the path length to obtain the path of the robot's operating space. The inverse - solution module that combines the greedy algorithm and the neural network calculates the loss according to the target pose and the current pose predicted by the neural network, and adjusts the input of the neural network along the gradient descent direction under the action of the greedy strategy. The neural network first tries to adjust the proximal joints of the robot. If a suitable inverse solution cannot be found, it then tries to adjust the proximal joints and their adjacent joints of the robot until the path in the robot's operating space is transformed into the path in the joint space, completing the robot path planning task.

[0026] The RRT - star module provides a fast and feasible initial path, overcoming the problem of low exploration efficiency of pure reinforcement learning in complex environments. The channel - guided DRL module can guide the robot to move along the channel, reduce the collision between the robot and obstacles, and improve the training efficiency. The inverse - solution module that combines the greedy algorithm and the neural network preferentially adjusts the proximal joints and only expands to adjust adjacent joints when necessary, reducing computational redundancy. This method combines the advantages of multiple technologies to form a closed - loop of "generation - optimization - execution", combines efficient global exploration and local optimization, and surpasses single algorithms in terms of path quality, computational efficiency, and environmental adaptability, providing a new solution for the path planning task of redundant - degree - of - freedom robots in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is the interaction flowchart of each module of a channel - guided fusion robot path planning method provided by an embodiment of the present invention;

[0028] Figure 2 It is the flowchart of the channel - guided fusion RRT - star and DRL's TRD structure provided by an embodiment of the present invention;

[0029] Figure 3 It is a flowchart of the inverse solution module that combines the greedy algorithm and neural network provided by the embodiment of the present invention. Specific embodiments

[0030] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0031] As Figure 1 shown, the embodiment of the present invention provides a channel-guided fusion robot path planning method, including the following steps:

[0032] Step 1: To avoid life and property losses caused by improper actions during physical debugging, a digital virtual system of the robot operation space restored 1:1 is built in the simulation environment. This digital virtual system of the robot operation space includes information such as the shapes and poses of the robot and obstacles.

[0033] Step 2: The TRD structure that combines RRT-star guided by channels and DRL outputs the path of the robot operation space. The TRD structure includes an RRT-star module and a DRL module guided by channels. The RRT-star module provides multiple feasible paths to guide the exploration in the early stage of DRL. The DRL module guided by channels further optimizes the path length and outputs a single optimized path. The specific steps are as follows:

[0034] Step 2-1: Build the RRT-star module and initialize the parameters of RRT-star. The RRT-star module plans multiple feasible paths F start from the starting position P end of the robot to the ending position P c , and put these paths into the memory bank of DRL to guide the exploration in the early stage of DRL. The expression of a single path F c is as follows:

[0035] F c ={(x c1 , y c1 , z c1 ),…,(x cn , y cn , z cn ),…,(x cN , y cN , z cN )}

[0036] where N is the total number of nodes.

[0037] Step 2-2: Based on the path provided by the RRT-star module, with P shor as the center and u as the radius, construct a channel, and build a DRL module using the SAC algorithm guided by the channel. Its reward function consists of three parts: a channel reward function, an end-point distance reward function, and a collision reward function.

[0038] The reward function of the DRL module guided by the channel is defined as:

[0039] r curr = f(P curr , P shor ) + g(P curr , P end ) + colli(P curr )

[0040] where r curr is the current reward function of the DRL module, which encourages the robot to move along the channel, guides the robot to explore the shortest path, and learn the optimal strategy. f(P curr , P shor ) is the channel reward function, g(P curr , P end ) is the end-point distance reward function, and colli(P curr ) is the collision reward function.

[0041] Specifically, the channel reward function is defined as:

[0042]

[0043] where dis1 represents the distance between P curr and P shor , u is the radius of the channel, and A and B are positive constants used to adjust the maximum range of the channel reward in each round. If P curr is outside the channel, the robot receives a negative reward. On the contrary, if P curr is inside the channel, the robot receives a positive reward. The channel reward encourages the robot to move inside the channel. The farther the robot is from the channel, the smaller the reward it gets.

[0044] The end-point distance reward function is defined as:

[0045]

[0046] where dis2 represents the distance between P curr and P end , and C is a positive constant used to adjust the maximum range of the end-point distance reward in each round. If P curr is outside the channel, the end-point distance reward is 0. If Pcurr Inside the channel, the reward is calculated based on the distance from the end point. The end - point distance reward function creates a non - uniform reward distribution based on the distance from the end point, encouraging the robot to approach the end point. The closer the robot is to the end point, the greater the reward it obtains.

[0047] The collision reward function is defined as:

[0048]

[0049] where \(E\) is a negative constant. The collision reward function includes a collision - detection function. If the robot collides with an obstacle, it will receive a negative reward; otherwise, the reward is 0.

[0050] DRL is trained under the guidance of the reward function. After the training is completed, the optimized single path \(X\) output is defined as:

[0051] \(X=\{(x 1 ,y 1 ,z 1 ),…,(x m ,y m ,z m ),…,(x M ,y M ,z M )\}\)

[0052] where \(M\) is the total number of nodes. The process of the TRD structure that fuses the channel - guided RRT - star and DRL is as Figure 2 shown.

[0053] Step 3: The inverse - solution module that fuses the greedy algorithm and the neural network transforms the path in the robot's operational space into a path in the robot's joint space.

[0054] Step 3 - 1: Build an inverse - solution module that fuses the greedy algorithm and the neural network, and initialize the robot's DH parameter model, joint OBB parameters, neural - network parameters, and the robot's target pose.

[0055] Let the robot's DH parameter model be:

[0056] \(f(\Theta)=T + \varepsilon\)

[0057] \(\|\varepsilon\|<\delta\) kin

[0058] where \(\Theta\) is the joint angle of the robot, \(T\) is the pose of the robot, and \(\delta\) kin is the pose tolerance error.

[0059] The OBB parameters of each joint \(i\) in the local coordinate system are defined as:

[0060]

[0061] Among them, c i is the center point of the joint, is the axial unit vector, is the semi-side length.

[0062] The input of the neural network is defined as K=(k 1 , k 2 , k 3 , k 4 , k 5 , k 6 ), with the initial value being a random number, defining the connection weights of the neural network, and the initial value being a random number.

[0063] The target pose matrix of the current node of the robot is defined as:

[0064]

[0065] Among them, x target , y target , z target is the target position of the current node in the robot's operating space, is the target attitude of the current node in the robot's operating space.

[0066] Step 3-2: Adjust the input of the neural network based on the greedy algorithm. For each dimension b of the input vector K, generate two candidate inputs and where e b is the unit vector of the b-th dimension, and λ is the step size.

[0067] Step 3-3: Hierarchical adjustment of the neural network. The th layer allows adjustment centered on the collision joint, with the adjacent joints before and after.

[0068] The learning rate of the neural network decays exponentially with the number of iterations and is defined as:

[0069]

[0070] Among them, η 0 is the initial value, is the current layer number, and Φ is the decay coefficient.

[0071] The first layer allows adjustment of the collision joint itself. If the required inverse solution is obtained, the loop is exited; if the required inverse solution cannot be obtained, enter the second layer, and at the same time adjust the proximal joint and the 2 adjacent joints before and after, and so on, and solve the loop. The inverse solution refers to a set of joint angles, and many sets of joint angles together form the path of the robot joint space.

[0072] After the solution is completed, the neural network outputs the predicted joint angles, and the joint angles are defined as:

[0073]

[0074] Step 3-4: For the obstacle point p, calculate the signed distance to the OBB, and the signed distance is defined as:

[0075]

[0076] Step 3-5: According to the DH parameter model of the robot, convert the predicted joint angles into the pose in the robot operating space, and the pose in the robot operating space is defined as:

[0077]

[0078] where, is the position in the robot operating space predicted by the neural network, is the attitude in the robot operating space predicted by the neural network.

[0079] Step 3-6: Calculate the loss, and the loss is divided into position loss, attitude loss, energy loss, and safety threshold loss. The position loss is defined as:

[0080]

[0081] where, ||*|| is the norm of the vector, that is, the square root of the sum of the squares of each element of the vector.

[0082] The attitude loss is defined as:

[0083]

[0084] where, ||*|| is the Frobenius norm of the matrix, that is, the square root of the sum of the squares of each element of the matrix.

[0085] The energy loss is defined as:

[0086] loss energy =||W(Θ 1 -Θ 2 )|| 2

[0087] W=(w 1 ,w 2 ,w 3 ,w 4 ,w 5 ,w 6 ,w 7 )

[0088]

[0089] Among them, W is the energy consumption coefficient matrix, l i is the moment of inertia, w i is the joint angular velocity, μ i is the friction coefficient, m i is the mass, v i is the linear velocity.

[0090] The safety threshold loss is defined as:

[0091]

[0092] Among them, cjoint is the set of collision joints, d(p, OBB j ) is the signed distance from the obstacle point p to the OBB parameter, d safe is the safety threshold.

[0093] The overall loss function is defined as:

[0094] loss = α * loss position + β * loss orientation + γ * loss energy + ρ * loss secure

[0095] Among them, α, β, γ, ρ are weight coefficients.

[0096] Select the one with the smaller loss among the two candidate inputs and as the input of the neural network and update the network.

[0097] Step 3 - 7: Return to Step 3 - 2 and perform iterative solution;

[0098] Step 3 - 8: When the value of the overall loss function of the neural network loss is less than the loss function threshold loss_threshold or the number of iterations ANN_step is equal to the iteration number threshold ANN_step_threshold, stop the loop; at this time, the inverse solution of one node is solved;

[0099] Step 3 - 9: Return to Step 3 - 1, take the next node as the target pose of the robot, and solve the inverse solution of the next node;

[0100] Step 3 - 10: When the inverse solutions of all nodes are obtained, the path in the robot joint space is obtained and defined as:

[0101] Θ = {(θ 11 , θ 12 , θ 13 , θ 14 , θ 15, θ 16 , θ 17 ), …,

[0102] (θ h1 , θ h2 , θ h3 , θ h4 , θ h5 , θ h6 , θ 17 ), …, (θ H1 , θ H2 , θ H3 , θ H4 , θ H5 , θ H6 , θ H7 )}

[0103] where H is the total number of nodes. The process of the inverse solution module that combines the greedy algorithm and the neural network is as Figure 3 shown.

[0104] Step 4: Import the path in the joint space into the physical robot, and the robot moves along this path.

[0105] A channel-guided fusion robot path planning method proposed by an embodiment of the present invention uses a channel-guided TRD structure that fuses RRT-star and DRL to plan the path of the robot's operating space. The RRT-star module provides a fast and feasible initial path, overcoming the problem of low exploration efficiency of pure reinforcement learning in complex environments. The channel-guided DRL module can guide the robot to move along the channel, reduce the collision between the robot and obstacles, and improve the training efficiency. The inverse solution module that combines the greedy algorithm and the neural network preferentially adjusts the proximal joints and only expands and adjusts the adjacent joints when necessary, reducing computational redundancy. This method combines the advantages of multiple technologies to form a closed loop of "generation - optimization - execution", combining efficient global exploration and local optimization, and surpassing single algorithms in terms of path quality, computational efficiency, and environmental adaptability, providing a new solution for the path planning task of redundant degree-of-freedom robots in complex scenarios.

Claims

1. A channel-guided fusion robot path planning method, characterized in that: The following steps are involved: Create a digital virtual system of the robot's operating space based on the shape and posture information of the robot and obstacles; In the robot operation space of the digital virtual system of the robot operation space, the robot is planned to move from its starting position P start To the end position P end The feasible path is defined as multiple paths F c , and the robot's current position P curr The nearest node position P shor As the center, build a channel with u as the radius, combined with P curr , P shor , u and end position P end , the optimized single path X is obtained by training with channel reward, terminal distance reward and collision reward; According to the robot's joint angle θ, posture T and posture tolerance error δ kin0 Define the robot parameter model, according to the center point c of each joint i of the robot i , axial unit vector and half side length Define OBB parameters, disassemble the optimized single path X, obtain the target position and target posture in the operation space of multiple robots, define the target posture matrix of the current node according to the target position and target posture, and generate positive step candidate inputs for each dimension of any random neural network input set K and negative step size candidate input Assume that the first layer only allows the adjustment of the proximal joint of the collision, and combines the positive step length candidate input and negative step size candidate input Solve with random initial connection weights of the neural network to obtain a collision-free inverse solution; If no collision-free inverse solution is obtained, the second layer is entered, and the learning rate of the neural network is reduced according to the exponential decay rule. The second layer allows adjustment of the two adjacent joints centered on the proximal joint of the collision, and again combines the positive step size candidate input and negative step size candidate input Solve with the initial connection weights to obtain the collision-free inverse solution; According to the robot parameter model, the positive step length candidate is input and negative step size candidate input Corresponding predicted joint angle Convert to positive step size candidate input and negative step size candidate input Corresponding to the robot's position in the operating space; According to the positive step size candidate input and negative step size candidate input Corresponding to the position and posture of the robot in the operation space, the position loss and posture loss are calculated by combining the target position and target posture respectively. According to the energy consumption coefficient matrix W and the moment of inertia l i 、Joint angular velocity w i , friction coefficient μ i 、Mass m i and linear velocity v i , calculate the energy loss, and calculate the safety threshold loss according to the collision joint set, the signed distance from the obstacle point p to the OBB parameters and the safety threshold; Assign the position loss, attitude loss, energy loss and safety threshold loss to the corresponding weight coefficients to obtain the positive step size candidate input and negative step size candidate input The corresponding overall loss function, between two candidate inputs and Select the one with smaller loss as the new input, and update the initial connection weights of the neural network along the direction of the gradient descent of the overall loss function; Update positive step size candidate input and negative step size candidate input Iterate again using the new input and the updated initial connection weights until the value of the overall loss function is less than the loss function threshold loss_threshold or the number of iterations ANN_step is equal to the iteration threshold ANN_step_threshold, and the inverse solution of the current node is obtained; Continue to solve the inverse solution of the next node of the current node, and iterate again until all nodes in a single path X are inversely solved. At this time, the inverse solutions corresponding to all nodes are obtained, and finally the path Θ of the robot joint space is formed.

2. A channel-guided fusion robot path planning method as claimed in claim 1, characterized in that: According to the current position P of the robot curr , distance from current position P curr The nearest node position P shor And channel radius u, construct the channel reward function for training and optimizing the optimized path X, the formula is as follows: Among them, dis1 represents P curr and P shor and the distance between them, A and B are positive constants used to adjust the maximum range of channel rewards in each round.

3. A channel-guided fusion robot path planning method as claimed in claim 1, characterized in that: According to the current position P of the robot curr , end position P end And the channel radius u, construct the terminal distance reward function for training and optimizing the optimized path X, the formula is as follows: Among them, dis2 represents P curr and P end The distance between the two ends, C is a positive constant used to adjust the maximum range of the end distance reward in each round.

4. A channel-guided fusion robot path planning method as claimed in claim 1, characterized in that: According to the current position P of the robot curr , construct a collision reward function for training and optimizing the optimized path X, the formula is as follows: Where W is a negative constant.