Method and system for controlling path of mining transport vehicle

By using the RRT algorithm to plan the global path in the mining transport vehicle path planning, and combining the TD3 algorithm to simulate and evaluate local trajectories, the problem of inaccurate path planning in complex mining areas is solved, and the accuracy and speed of path planning is improved.

CN119937412APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510089924.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing mining transport vehicle path planning technology has the problem of inaccurate path planning in complex mining areas, and the data processing network is insufficient, which affects the accuracy and speed of path planning.

Method used

The rapid expansion random tree algorithm RRT is used to plan the global path in the state space, and the dual-delay depth deterministic strategy gradient algorithm TD3 is used to simulate multiple operating trajectories of mining trucks. By evaluating and selecting the best trajectory, the path control of mining trucks is achieved by combining global paths and local paths.

Benefits of technology

It improves the stability of the data processing network and can accurately plan the transport vehicle path in complex mining areas, ensuring the quality of the path while improving the path planning speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937412A_ABST
    Figure CN119937412A_ABST
Patent Text Reader

Abstract

The invention discloses a mine transport vehicle path control method and system, and relates to the technical field of path planning. Comprising the following steps of: 1, searching a global path from a starting point to a target point according to an existing map, and sampling and planning the global path in a state space by utilizing a fast expansion random tree algorithm RRT; step 2, acquiring a global path, cutting the global path into a local map range, simulating a plurality of running tracks of the mining transport vehicle by using a double-delay depth deterministic strategy gradient algorithm TD3, evaluating the running tracks, selecting an optimal track as a local path of current running, and taking the optimal track as the local path of the current running; and combining the global path and the local path to realize the path control of the mining transport vehicle. The stability of a data processing network is improved, and the process planning speed is improved while the path quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention discloses a path control method and system for a mining transport vehicle, and relates to the technical field of path planning. Background Art

[0002] Mining is an important industry, but it has many problems such as high risk, high work intensity and poor working environment. Autonomous driving can solve the problem of mine transportation. Path planning, as the core part of the autonomous driving framework, faces many challenges, especially in the face of complex mining operating environments and mine terrain. The existing path planning data processing network still needs to improve stability when processing data, and virtual data can easily dilute real data during planning, resulting in inaccurate path planning. Summary of the invention

[0003] In view of the problems of the prior art, the present invention provides a mine transport vehicle path control method and system, which improves the stability of the data processing network, can realize the mine path planning of the transport vehicle in the complex environment of the mine area, and improve the process planning speed while ensuring the path quality.

[0004] The specific scheme proposed by the present invention is:

[0005] The present invention provides a mine transport vehicle path control method for planning the mine transport vehicle path:

[0006] Step 1: Find the global path from the starting point to the target point based on the existing map: Use the Rapid Random Tree Algorithm RRT to sample and plan the global path in the state space;

[0007] Step 2: Obtain the global path, clip the global path to the local map range, use the double-delay deep deterministic policy gradient algorithm TD3 to simulate multiple running trajectories of the mining transport vehicle, evaluate the running trajectories, select the best trajectory as the current running local path, and combine the global path and the local path to realize the path control of the mining transport vehicle.

[0008] Furthermore, in step 1 of the path control method for a mining transport vehicle, a global path is sampled and planned in a state space using a rapid expansion random tree algorithm RRT, including: setting a starting node and a target node, taking the starting node as the root node of the random tree,

[0009] Assume that the current node of the random tree is x, use the random function to randomly sample in the state space to obtain the sampled node, set it as xsample, if xsample is not within the obstacle, find the node with the closest Euclidean distance to xsample, set it as xnearest, and generate a new node along the direction of the line connecting xsample and xnearest with a step size d, set it as xnew,

[0010] If xnew is not within the obstacle and the path connected with xnearest does not collide with the obstacle, then xnew is added to the current candidate node list and the parent node of xnew is designated as xnearest;

[0011] When the Euclidean distance between the latest sampling node and the target node is less than the current step size, the current path planning is determined to be successful, and the parent node is traced back from the last sampling node, and the entire path is traced back based on the parent node.

[0012] Furthermore, step 2 of the path control method for a mining transport vehicle includes: using a double-delay deep deterministic policy gradient algorithm TD3 to train and generate a policy network and a dual Q network, generating an action strategy through the policy network, and determining the action to be taken by the mining transport vehicle under a given state according to the action strategy, thereby generating an operation trajectory; evaluating the Q value under a given state and action through the dual Q network, comparing the Q values ​​corresponding to different actions, judging the quality of the action, and thus evaluating the operation trajectory, selecting the best evaluated operation trajectory as the optimal trajectory, and using the optimal trajectory as the local path of the current operation.

[0013] Furthermore, in step 2 of the path control method for a mining transport vehicle, an experience replay pool is created to store experience data generated by the interaction between the mining transport vehicle and the operating environment, which is used to train a strategy network and a double Q network, wherein the experience data includes observations, actions, rewards, and the next environmental state, wherein the observations refer to the environmental state observed by the mining transport vehicle at a certain moment; the actions refer to the actions selected by the mining transport vehicle under the observations; the rewards refer to the reward values ​​returned by the environment after the selected actions are executed; and the next environmental state refers to the environmental state observed by the mining transport vehicle after the selected actions are executed.

[0014] The present invention also provides a mining transport vehicle path control system, including a planning control module, the planning control module includes a global path planning module and a local path planning module, and the planning control module plans the mining transport vehicle path:

[0015] The global path planning module finds the global path from the starting point to the target point according to the existing map: the global path is sampled and planned in the state space using the Rapidly Expanding Random Tree Algorithm (RRT);

[0016] The local path planning module obtains the global path, clips the global path to the local map range, uses the double-delay deep deterministic policy gradient algorithm TD3 to simulate multiple operating trajectories of the mining transport vehicle, evaluates the operating trajectories, and selects the best trajectory as the current local path. The planning and control module combines the global path and the local path to realize the path control of the mining transport vehicle.

[0017] Furthermore, the global path planning module of the path control system of a mining transport vehicle uses a rapid expansion random tree algorithm RRT to sample and plan a global path in the state space, including: setting a start node and a target node, taking the start node as the root node of the random tree,

[0018] Assume that the current node of the random tree is x, use the random function to randomly sample in the state space to obtain the sampled node, set it as xsample, if xsample is not within the obstacle, find the node with the closest Euclidean distance to xsample, set it as xnearest, and generate a new node along the direction of the line connecting xsample and xnearest with a step size d, set it as xnew,

[0019] If xnew is not within the obstacle and the path connected with xnearest does not collide with the obstacle, then xnew is added to the current candidate node list and the parent node of xnew is designated as xnearest;

[0020] When the Euclidean distance between the latest sampling node and the target node is less than the current step size, the current path planning is determined to be successful, and the parent node is traced back from the last sampling node, and the entire path is traced back based on the parent node.

[0021] Furthermore, the local path planning module of the path control system of a mining transport vehicle uses the dual-delay deep deterministic policy gradient algorithm TD3 training to generate a policy network and a dual Q network, generates an action strategy through the policy network, and determines the action to be taken by the mining transport vehicle under a given state according to the action strategy, thereby generating an operation trajectory; the Q value under a given state and action is evaluated through the dual Q network, the Q values ​​corresponding to different actions are compared, the quality of the action is judged, and the operation trajectory is evaluated, and the operation trajectory with the best evaluation is selected as the optimal trajectory, and the local path of the current operation of the optimal trajectory is used.

[0022] Furthermore, the local path planning module of the path control system of a mining transport vehicle creates an experience replay pool to store the experience data generated by the interaction between the mining transport vehicle and the operating environment, which is used to train the strategy network and the double Q network, wherein the experience data includes observation values, actions, rewards and the next environmental state, the observation value refers to the environmental state observed by the mining transport vehicle at a certain moment; the action refers to the action selected by the mining transport vehicle under the observation value; the reward refers to the reward value returned by the environment after the selected action is executed; the next environmental state refers to the environmental state observed by the mining transport vehicle after the selected action is executed.

[0023] The benefits of the method of the present invention are:

[0024] The rapidly expanding random tree algorithm RRT is used to sample and plan the global path in the state space to obtain the global path. The double-delay deep deterministic policy gradient algorithm TD3 is used to simulate multiple operation trajectories of mining transport vehicles, evaluate the operation trajectories, and select the best trajectory as the current local path. The path control of mining transport vehicles is realized by combining the global path and the local path. This not only improves the stability of the data processing network, but also enables the mining path planning of transport vehicles in complex mining environments, and improves the process planning speed while ensuring the path quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a schematic flow chart of the method of the present invention.

[0026] Figure 2 This is a schematic diagram of the detection path of a mining transport vehicle.

[0027] Figure 3 This is a diagram of the angle reward. DETAILED DESCRIPTION

[0028] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0029] Example 1

[0030] The present invention provides a mine transport vehicle path control method for planning the mine transport vehicle path:

[0031] Step 1: Find the global path from the starting point to the target point based on the existing map: Use the rapidly expanding random tree algorithm RRT to sample and plan the global path in the state space.

[0032] The rapid expansion random tree algorithm RRT is used to sample and plan the global path in the state space, which may specifically include: setting the starting node and the target node, and taking the starting node as the root node of the random tree,

[0033] Assume that the current node of the random tree is x, use the random function to randomly sample in the state space to obtain the sampling node, set it as x sample , if x sample If it is not within the obstacle, find sample The node with the closest Euclidean distance is x nearest , along x sample With x nearest The connection direction generates a new node with a step size d, set to x new ,

[0034] If x new Not within the obstacle and with x nearestIf the connected path does not collide with obstacles, then x new Add to the current candidate node list and specify x new The parent node is x nearest ;

[0035] When the Euclidean distance between the latest sampling node and the target node is less than the current step size, the current path planning is determined to be successful, and the parent node is traced back from the last sampling node, and the entire path is traced back based on the parent node.

[0036] In step 1, a reward function can also be designed to improve the behavior selection of the mining transport vehicle in the path.

[0037] Let the reward R of the final target point be goal , get reward R goal The condition is that the distance between the mining transport vehicle and the target point is less than the step size expanddis at the current time t, for example x new is the current node position of the mining transport vehicle, using the formula:

[0038]

[0039] Get Reward R goal .

[0040] Set the collision penalty R crash :When the sampling generated x new When inside an obstacle, a negative reward R is given collision ; When x new With parent node x nearest The connected path collides with an obstacle and gives a negative reward R intersect ; When x new If the current operating environment is exceeded, a negative reward R will be given. out . It can be expressed by the following formula:

[0041]

[0042] R crash =R collision +R intersect +R out

[0043] Set the collision penalty R crash , the purpose is to avoid collision between mining transport vehicles and obstacles.

[0044] Design time penalty R count With distance penalty R distance . Time penalty R count The mining vehicle will receive a fixed negative reward of -0.1 every time it interacts with the environment; the distance penalty R distanceThen calculate the straight-line distance between the current mining transport vehicle and the target point goal, and encourage the transport vehicle to move toward the target point. It can be expressed by the following formula:

[0045] R count =-0.1

[0046] R diatance -‖x-goal‖×0.01

[0047] Set time penalty R count With distance penalty R distance , speed up the training, guide the transport vehicle to the target point, and avoid the transport vehicle from stagnating or circling in place.

[0048] Design the positive reward R associated with the turning angle of the path angle , where alpha is x new ,x nearest The connected path and x nearest The angle between the path connected to its parent node x. nearest 、x new 、x sample Located on the same straight line, as shown in the attached Figure 3 As shown, this reward encourages the transporter to sample x sample When , alpha is made as large as possible, so that the path turning angle is reduced. It can be expressed by the following formula:

[0049]

[0050] Design a positive reward R related to the turning angle of the path angle , making the planned route smoother.

[0051] The final reward function R consists of five parts: collision penalty, time penalty, distance penalty, angle reward and final reward:

[0052] R=R count +R goal +R angle +R distance +R crash

[0053] Used to improve the behavior of mining vehicles in the path.

[0054] Step 2: Obtain the global path, clip the global path to the local map range, use the double-delay deep deterministic policy gradient algorithm TD3 to simulate multiple running trajectories of the mining transport vehicle, evaluate the running trajectories, select the best trajectory as the current running local path, and combine the global path and the local path to realize the path control of the mining transport vehicle.

[0055] Among them, the local path generation process may include: using the dual-delay deep deterministic policy gradient algorithm TD3 to train and generate a policy network and a dual Q network, generating an action strategy through the policy network, and determining the action that the mining transport vehicle should take under a given state according to the action strategy, thereby generating an operation trajectory; evaluating the Q value under a given state and action through the dual Q network, comparing the Q values ​​corresponding to different actions, judging the quality of the action, and thus evaluating the operation trajectory, selecting the best evaluated operation trajectory as the optimal trajectory, and using the optimal trajectory as the local path of the current operation.

[0056] In step 2, an experience replay pool can also be created to store the experience data generated by the interaction between the mining transport vehicle and the operating environment, which is used to train the strategy network and the double Q network. The experience data includes observations, actions, rewards and the next environment state. The observation value refers to the environment state observed by the mining transport vehicle at a certain moment; the action refers to the action selected by the mining transport vehicle under the observation value; the reward refers to the reward value returned by the environment after the selected action is executed; the next environment state refers to the environment state observed by the mining transport vehicle after the selected action is executed.

[0057] The double-delayed deep deterministic policy gradient algorithm TD3 is used to randomly use experience data from the experience replay pool to train the policy network and the double Q network. This random sampling method helps to break the time correlation between experiences and makes the training process more stable.

[0058] Example 2

[0059] The present invention also provides a mining transport vehicle path control system, including a planning control module, the planning control module includes a global path planning module and a local path planning module, and the planning control module plans the mining transport vehicle path:

[0060] The global path planning module finds the global path from the starting point to the target point according to the existing map: the global path is sampled and planned in the state space using the Rapidly Expanding Random Tree Algorithm (RRT);

[0061] The local path planning module obtains the global path, clips the global path to the local map range, uses the double-delay deep deterministic policy gradient algorithm TD3 to simulate multiple operating trajectories of the mining transport vehicle, evaluates the operating trajectories, and selects the best trajectory as the current local path. The planning and control module combines the global path and the local path to realize the path control of the mining transport vehicle.

[0062] As the information interaction and execution process between the modules of the above-mentioned system are based on the same concept as the embodiment of the method of the present invention, the specific contents can be found in the description of the embodiment of the method of the present invention and will not be repeated here.

[0063] Likewise, the benefits of the system of the present invention are:

[0064] The rapidly expanding random tree algorithm RRT is used to sample and plan the global path in the state space to obtain the global path. The double-delay deep deterministic policy gradient algorithm TD3 is used to simulate multiple operation trajectories of mining transport vehicles, evaluate the operation trajectories, and select the best trajectory as the current local path. The path control of mining transport vehicles is realized by combining the global path and the local path. This not only improves the stability of the data processing network, but also enables the mining path planning of transport vehicles in complex mining environments, and improves the process planning speed while ensuring the path quality.

[0065] It should be noted that not all steps and modules in the above-mentioned processes and system structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.

[0066] The above-described embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or changes made by those skilled in the art based on the present invention are within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A mining transport vehicle path control method, characterized by planning Mine transport vehicle route: Step 1: Find the global path from the starting point to the target point based on the existing map: use the rapid expansion random tree algorithm RRT to sample and plan the global path in the state space; Step 2: Obtain the global path, clip the global path to the local map range, use the double-delay deep deterministic policy gradient algorithm TD3 to simulate multiple running trajectories of the mining transport vehicle, evaluate the running trajectories, select the best trajectory as the current running local path, and combine the global path and the local path to realize the path control of the mining transport vehicle.

2. A mining transport vehicle path control method according to claim 1, characterized in that In step 1, the rapid expansion random tree algorithm RRT is used to sample and plan the global path in the state space, including: setting the starting node and the target node, and taking the starting node as the root node of the random tree. Assume that the current node of the random tree is x, use the random function to randomly sample in the state space to obtain the sampling node, set it as x sample , if x sample If it is not within the obstacle, find sample The node with the closest Euclidean distance is x nearest , along x sample With x nearest The connection direction generates a new node with a step size d, set to x new , If x new Not within the obstacle and with x nearest If the connected path does not collide with obstacles, then x new Add to the current candidate node list and specify x new The parent node is x nearest ; When the Euclidean distance between the latest sampling node and the target node is less than the step size d, the current path planning is determined to be successful, and the parent node is traced back from the last sampling node, and the entire path is traced back based on the parent node.

3. A mining transport vehicle path control method according to claim 1, characterized in that Step 2 includes: using the double-delayed deep deterministic policy gradient algorithm TD3 to train and generate a policy network and a dual Q network, generating an action strategy through the policy network, and determining the action that the mining transport vehicle should take under a given state according to the action strategy, thereby generating an operation trajectory; evaluating the Q value under a given state and action through the dual Q network, comparing the Q values ​​corresponding to different actions, judging the quality of the action, and thus evaluating the operation trajectory, selecting the best evaluated operation trajectory as the optimal trajectory, and setting the local path of the current operation of the optimal trajectory.

4. A mining transport vehicle path control method according to claim 1, characterized in that In step 2, an experience replay pool is created to store the experience data generated by the interaction between the mining transport vehicle and the operating environment, which is used to train the strategy network and the double Q network. The experience data includes observations, actions, rewards and the next environment state. The observations refer to the environment state observed by the mining transport vehicle at a certain moment; the actions refer to the actions selected by the mining transport vehicle under the observations; the rewards refer to the reward values ​​returned by the environment after the selected actions are executed; and the next environment state refers to the environment state observed by the mining transport vehicle after the selected actions are executed.

5. A mining transport vehicle path control system, characterized in that It includes a planning and control module, which includes a global path planning module and a local path planning module. The planning and control module plans the path of the mining transport vehicle: The global path planning module finds the global path from the starting point to the target point according to the existing map: the global path is sampled and planned in the state space using the Rapidly Expanding Random Tree Algorithm (RRT); The local path planning module obtains the global path, clips the global path to the local map range, uses the double-delay deep deterministic policy gradient algorithm TD3 to simulate multiple operating trajectories of the mining transport vehicle, evaluates the operating trajectories, and selects the best trajectory as the current local path. The planning and control module combines the global path and the local path to realize the path control of the mining transport vehicle.

6. A mining transport vehicle path control system according to claim 5, characterized in that the global The path planning module uses the rapidly expanding random tree algorithm RRT to sample and plan the global path in the state space, including: setting the starting node and the target node, taking the starting node as the root node of the random tree, assuming that the current node of the random tree is x, and using the random function to randomly sample in the state space to obtain the sampling node, which is set as x. sample , if x sample If it is not within the obstacle, find sample The node with the closest Euclidean distance is x nearest , along x sample With x nearest The connection direction generates a new node with a step size d, set to x new , If x new Not within the obstacle and with x nearest If the connected path does not collide with obstacles, then x new Add to the current candidate node list and specify x new The parent node is x nearest ; When the Euclidean distance between the latest sampling node and the target node is less than the step size d, the current path planning is determined to be successful, and the parent node is traced back from the last sampling node, and the entire path is traced back based on the parent node.

7. A mining transport vehicle path control system according to claim 5, characterized in that The local path planning module uses the dual-delay deep deterministic policy gradient algorithm TD3 to train and generate a policy network and a dual Q network. The policy network generates an action strategy, and the action strategy determines the action that the mining transport vehicle should take in a given state, thereby generating an operation trajectory. The dual Q network is used to evaluate the Q value under a given state and action, compare the Q values ​​corresponding to different actions, judge the quality of the action, and thus evaluate the running trajectory. The best running trajectory is selected as the optimal trajectory, and the local path of the current running of the optimal trajectory is set.

8. A mining transport vehicle path control system according to claim 5, characterized in that The local path planning module creates an experience replay pool to store the experience data generated by the interaction between the mine transport vehicle and the operating environment, which is used to train the strategy network and the double Q network. The experience data includes observations, actions, rewards and the next environment state. The observation value refers to the environment state observed by the mine transport vehicle at a certain moment; the action refers to the action selected by the mine transport vehicle under the observation value; the reward refers to the reward value returned by the environment after the selected action is executed; the next environment state refers to the environment state observed by the mine transport vehicle after the selected action is executed.

Citation Information

Cited By

  • Track planning method and unmanned vehicle

    CN121877041A