Visual guidance following parking system and method for automatic driving of mine car

By adopting deep learning and reinforcement learning algorithm training subsystems and autonomous driving subsystems in the mine car automatic driving system, the lack of mobility of traditional navigation technology and path planning strategies in the complex environment of open-pit mining areas is solved, and the efficient follow-up and stopping and obstacle avoidance capabilities of mine car are achieved.

CN120096552AActive Publication Date: 2025-06-06SHANDONG JIAOTONG UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510600576.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-06
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Traditional navigation technology and path planning strategies have shortcomings in terms of mobility, making it difficult to effectively adjust dynamically and deal with obstacles in complex environments in open-pit mining areas as they follow the target location.

Method used

A mining vehicle autonomous driving visual guided follow-up docking system is adopted, which includes a training subsystem and an autonomous driving subsystem. The training subsystem uses deep learning and reinforcement learning algorithms to train a network model for autonomous driving through components such as on-board vision module, excitation module, control signal generation module and training module. The autonomous driving subsystem is based on these trained models to perform automatic driving, obstacle avoidance and docking operations of mine cars in real time.

Benefits of technology

It has achieved efficient follow-up and obstacle avoidance capabilities of the mine car automatic driving system in complex environments in open-pit mining areas, making up for the insufficient mobility of traditional navigation technology and path planning strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120096552A_ABST
    Figure CN120096552A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving, and discloses a mine car automatic driving visual guidance following parking system and method.The mine car automatic driving visual guidance following parking system comprises a training subsystem and an automatic driving subsystem, and the training subsystem comprises a vehicle-mounted visual module, a driving control module, a control signal generation module, a training module, an excitation module and a safety management module; the signal output end of the vehicle-mounted vision module is connected with the signal input end of the excitation module, the signal output end of the excitation module is connected with the signal input ends of the training module and the control signal generation module, and the signal output end of the training module is connected with the signal input ends of the safety management module and the control signal generation module. The signal output end of the control signal generation module is connected with the signal input ends of the training module and the driving control module; according to the invention, the problem of insufficient maneuverability of a traditional navigation technology and a path planning strategy can be solved, and the safety trafficability under the condition of sudden obstacles and the autonomous driving ability in a dynamic change environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a mine car autonomous driving vision-guided following docking system and method. Background Art

[0002] Autonomous driving navigation and positioning methods include satellite navigation and positioning, inertial navigation, visual (laser) SLAM navigation, magnetic nail (tape, QR code) navigation, etc. Autonomous driving vehicles mainly use path planning algorithms for short-range docking or parking. For example, the graph search algorithm represents the driving map with a grid and calculates the shortest path from the starting point to the end point through a heuristic algorithm; the sampling algorithm randomly generates multiple sampling points and searches for valid paths in a complex environment point by point; the optimization algorithm outputs an optimal path that meets environmental and vehicle constraints by minimizing or maximizing the cost function. The path planning algorithm requires manual setting of the regional grid map and the docking or parking target position coordinates of the autonomous driving vehicle in advance, and then calculates and generates the vehicle operation trajectory and control strategy. It is not suitable for special situations such as dynamic adjustment of the target position. Summary of the invention

[0003] The purpose of the present invention is to provide a mine car automatic driving vision-guided following docking system and method to solve the problem of insufficient mobility of traditional navigation technology and path planning strategy.

[0004] In order to achieve the above purpose, the technical solutions adopted are as follows: In a first aspect, the present invention provides a mine car automatic driving vision-guided following and docking system, characterized in that the system comprises a training subsystem and an automatic driving subsystem, wherein the training subsystem comprises an on-board vision module, a drive control module, a control signal generating module, a training module, an excitation module and a safety management module; wherein the signal output end of the on-board vision module is connected to the signal input end of the excitation module, the signal output end of the excitation module is connected to the signal input end of the training module and the control signal generating module, the signal output end of the training module is connected to the signal input end of the safety management module and the control signal generating module, and the signal output end of the control signal generating module is connected to the signal input end of the training module and the drive control module; The automatic driving subsystem includes an on-board vision module, a drive control module, a control signal generating module and a safety management module; wherein, the signal output end of the on-board vision module is connected to the signal input end of the control signal generating module and the safety management module, the signal output end of the control signal generating module is connected to the signal input end of the drive control module, and the signal output end of the safety management module is connected to the signal input end of the control signal generating module.

[0005] Furthermore, in the training subsystem, the vehicle-mounted vision module has a built-in RGB camera, a depth camera and a video preprocessing unit, which are used to obtain RGBD data and send the RGBD data to the excitation module; The excitation module has a built-in target detection network and an excitation value calculation unit. The target detection network is used to pre-train and identify weight parameters of the following and stopping target objects and obstacles inside and outside the path according to the input RGBD data, realize target detection of the following and stopping target objects and obstacles inside and outside the path, and send the detection result data to the excitation value calculation unit; the excitation value calculation unit is used to calculate and generate distance data of the mine car deviating from the following and stopping target object according to the detection result data, and then calculate the excitation value and the docking instruction of the approaching following and stopping target object, and calculate the relative distance data between the mine car and the obstacles inside and outside the path, generate obstacle avoidance instructions according to the strategy, and transmit them to the training module in real time; The control signal generation module has a built-in main prediction network, an action instruction generation strategy unit, an index strategy unit and a mode control unit, and is used to generate control instructions and send them to the drive control module.

[0006] Furthermore, the training module has a built-in target network, an experience replay cache unit, a remote monitoring unit and an automatic reset control unit. The experience replay cache unit is used to store RGBD data, excitation values, obstacle avoidance instructions and docking instructions. The target network is used to generate an expected value of the probability of an action instruction according to the data stored in the experience replay cache unit. The remote monitoring unit is used to send and receive manual remote control instructions. The automatic reset control unit is used to generate a reset action instruction according to the driving action instruction, the obstacle avoidance instruction and the docking instruction. The driving control module is used to receive a driving action instruction and control the mining vehicle based on the driving action instruction, wherein the driving action instruction includes a steering instruction and a speed instruction; The security management module has a built-in prediction supervision network, which obtains training data from the built-in experience replay cache unit of the training module and is trained independently of the main prediction network and the target network; In the autonomous driving subsystem, the on-board vision module has a built-in RGB camera, a depth camera and a video preprocessing unit, which are used to obtain RGBD data and send the RGBD data to the control signal generation module and the safety management module; The control signal generation module calculates and generates the probability of the mine car running action instruction, and the action instruction generation strategy unit calculates and generates the mine car driving action instruction; The driving control module is used to receive a driving action instruction and control the mining vehicle based on the driving action instruction, wherein the driving action instruction includes a steering instruction and a speed instruction; The safety management module has a built-in prediction supervision network, which generates obstacle avoidance / docking instructions based on RGBD data and sends them to the control signal generation module. The control signal generation module readjusts the driving direction action instructions based on the obstacle avoidance instructions, or generates a docking operation action instruction based on the docking instruction.

[0007] Furthermore, the control signal generation module, during the training phase, alternately executes the observation mode and the training mode; When the observation mode is executed, RGBD data, incentive value, obstacle avoidance instruction and docking instruction are obtained from the incentive module, the RGBD data is input into the main prediction network, the probability of the mine car driving action instruction is calculated, and the action instruction generation strategy unit is input to calculate and generate the mine car driving action instruction, which is transmitted to the drive control module, and the RGBD data and its corresponding action instruction, incentive value and obstacle avoidance instruction, docking instruction data are sent to the experience replay cache unit; at the same time, based on the obstacle avoidance instruction and the docking instruction, the stop observation action instruction is triggered to drive the mine car to stop running, and stop transmitting the experience data to the experience replay cache unit; When executing the training mode, the main prediction network obtains RGBD data and action instruction data from the experience replay cache unit, calculates the probability of the mine car driving action instruction using the RGBD data, combines the action instruction data in the experience replay cache unit, calculates and generates the action instruction index value through the index strategy unit, and carries out the main prediction network training; The control signal generation module has a built-in mode control unit to control the action instruction generation mode and control the switching of the observation mode and the training mode. The action instruction generation mode includes the main prediction network generating action instructions, receiving and sending manual remote control instructions, receiving and sending obstacle avoidance instructions and docking instructions.

[0008] Furthermore, the control signal generation module has a built-in action instruction generation strategy unit, which generates action instructions in the following manner: Get the state-action value output by the main prediction network , ,in is the speed value, To turn to value; The action instruction generation strategy unit outputs the action instruction as , the custom random action instruction is expressed as , the custom random value is , the action command randomly adjusts the threshold to , , the speed command is calculated by the following formula a v and steering instructions a r : ; in, Indicates the bit index value of the maximum value in the output action instruction. Indicates a custom random speed command. Indicates a custom random turn instruction; During the training phase, as the number of training rounds increases, the , so that Constantly tending to the dependent state-action value, the action instruction randomly adjusts the threshold The reduction strategy is expressed as: ; in, is the initial value, is the pre-set minimum value, is the reduction value for each training round, epoches is the number of training rounds.

[0009] Furthermore, the training module has a built-in target network that forms a deep Q network with the main prediction network. The target network has the same network structure as the main prediction network. The target network receives weight parameters of the main prediction network at a predetermined period and updates itself. The target network provides action instruction probability expected values ​​for the main prediction network training.

[0010] Furthermore, the training module has a built-in experience replay cache unit to store the RGBD data, running action instructions, incentive values, obstacle avoidance instructions and docking instructions transmitted by the control signal generation module; in the training stage, the experience replay cache unit randomly selects experience data in batches, and sends the current frame RGBD data and the current frame corresponding running action instruction data in the experience data to the main prediction network in the control signal generation module; sends the current frame corresponding incentive value and the next frame RGBD data in the experience data to the target network; the experience replay cache unit randomly selects experience data in batches, and sends the current frame RGBD data and the current corresponding obstacle avoidance instructions and docking instructions in the experience data to the safety management module.

[0011] The experience replay cache unit sets a maximum storage capacity. When the storage capacity exceeds the maximum storage capacity, the experience replay cache unit deletes the earliest experience data and stores new data, always maintaining the maximum storage capacity. If the obstacle avoidance instruction and / or the docking instruction is a valid value, the experience replay cache unit suspends storing the new data.

[0012] Furthermore, the training module has a built-in automatic reset control unit, which stores the driving action instruction signal sequence sent by the control signal generating module from the starting position, and receives the obstacle avoidance instruction and the docking instruction sent by the excitation module. The obstacle avoidance instruction and the docking instruction trigger the automatic reset control unit to send a reset action instruction to the control signal generating module in the reverse order of the cached action instruction signal sequence, and then send it to the driving control module via the control signal generating module. The reset action instruction includes a motion action instruction and a rotation action instruction. The motion action instruction is the negative value of the original instruction, and the rotation action instruction is the original instruction, which drives the mine car back to the starting position.

[0013] Furthermore, the excitation module has a built-in target detection network, which calculates the coordinates of the upper left and lower right corners of the following and stopping target object and the obstacle detection box inside and outside the path relative to the upper left corner of the output RGB image data, and sends them to the excitation value calculation unit; The incentive value calculation unit is used to perform incentive value calculation and obstacle avoidance / docking value calculation to obtain incentive values, obstacle avoidance instructions and docking instructions; The incentive value calculation includes: The coordinates of the upper left and lower right corners of the RGB image detection box of the target object being followed and docked are expressed as and , then the minimum distance between the RGB image detection frame and the D depth map is ,but: ; in, Denotes the depth map D The distance value of the coordinate point; The incentive value is calculated by the following formula: ; in is the incentive value, is the empirical coefficient, arctg is the inverse tangent function; The obstacle avoidance / docking value calculation includes: In the RGBD image output by the vehicle vision module, establish the forward path of the mine car. x q Indicates the displacement of the left edge of the forward path relative to the left side of the image, x t Represents the forward path width.

[0014] The coordinates of the upper left and lower right corners of the RGB image detection box of the following target object inspection are expressed as and , the coordinates of the upper left and lower right corners of the obstacle RGB image detection box are expressed as and ; Calculate D depth map The corresponding path coverage decision value : ; in, Indicates a preset maximum value; According to the coordinates of the upper left and lower right corners of the detection frame of the following docked target and , calculate the minimum distance value of the target object: ; According to the coordinates of the upper left and lower right corners of the obstacle detection frame and , calculate the minimum obstacle distance value: ; Assume the output value of obstacle avoidance is b , the docking output value is ,but: ; in, is the minimum distance threshold between the minecart and the obstacle, It is the minimum distance threshold to the docking target; when the minimum distance between the mine car and the obstacle is less than the minimum threshold, the obstacle avoidance command is output; when the minimum distance between the mine car and the target is less than the minimum threshold, the docking command is output.

[0015] Furthermore, in the training phase, the deep Q network consisting of the target detection network in the excitation module, the prediction supervision network in the safety management module, the main prediction network in the control signal generation module, and the target network in the training module is trained respectively; The target detection network is pre-trained; during training, the mine car is manually controlled to run, and the on-board vision module shoots forward to follow the docked target object and nearby obstacles inside and outside the path, outputs RGBD data, divides and establishes training data sets and verification data sets, and carries out the target detection network training. After completing the target detection network training and performance verification test, the target detection network outputs the incentive value and obstacle avoidance instructions and docking instructions, which are applied to the prediction supervision network and the deep Q network training; The predictive supervision network is a classification network and is trained independently of the deep Q network. During training, the safety management module randomly obtains experience data in batches from the experience playback cache unit. The experience data includes RGBD data and its corresponding obstacle avoidance instructions and docking instructions. The obstacle avoidance instructions and docking instructions are used as RGBD data classification labels, and training data sets and verification data sets are divided to carry out the predictive supervision network. After completing the training and performance verification test of the predictive supervision network, the predictive supervision network outputs obstacle avoidance / docking instructions, which are used for obstacle avoidance or docking in the autonomous driving stage. The deep Q network training phase is divided into an observation mode and a training mode; the observation mode and the training mode are performed alternately; In the observation mode, the control signal generation module obtains RGBD data from the vehicle vision module, the main prediction network calculates and generates the probability of the mine car driving action command, the input action command generation strategy unit calculates and generates the action command, and sends it to the drive control module to drive the mine car to drive automatically, and synchronously sends the RGBD data and the corresponding action command data to the experience playback cache unit. In the observation mode, the main prediction network does not perform network weight parameter training; In training mode, the deep Q network is trained as follows: Initialize the network weight parameters of the main prediction network and the target network respectively; The main prediction network calculates and outputs the probability of action instructions: the control signal generation module randomly obtains experience data in batches from the experience playback cache unit. The experience data includes the current frame RGBD data and the action instruction data corresponding to the current frame; the current frame RGBD data is input into the main prediction network to calculate the probability of the mine car driving action instruction; The index strategy unit calculates and outputs the action instruction index value; The target network calculates the output cumulative incentive; Calculate the loss function according to the action instruction probability index value and the accumulated incentive; the loss function includes running speed loss and steering loss; Main prediction network weight update: Use speed loss, steering loss and Adam gradient descent optimization algorithm to back-propagate the main prediction network and update the network weight parameters; Target network weight update: Set the target network weight parameter update cycle. After the main prediction network is trained for a certain number of cycles, the main prediction network weight parameters are input into the target network, and the target network weight parameters are updated to be the same as the main prediction network.

[0016] The second invention is to provide a mine car automatic driving visually guided following and docking method, based on the mine car automatic driving visually guided following and docking system as described above, the method comprises: Obtain RGBD data through the vehicle vision module; Establish an incentive module, and build a target detection network and an incentive value calculation unit in the incentive module. The target detection network is used to pre-train and identify weight parameters of the following and stopping target objects and the adjacent obstacles inside and outside the path according to the input RGBD data, perform target detection on the following and stopping target objects and the adjacent obstacles inside and outside the path, and send the detection result data to the incentive value calculation unit; the incentive value calculation unit is used to calculate and generate distance data of the mine car deviating from the following and stopping target object according to the detection result data, and then calculate the incentive value and the stopping instruction of the approaching stopping target object, and calculate the relative distance data between the mine car and the adjacent obstacles inside and outside the path, generate obstacle avoidance instructions according to the strategy, and transmit them to the training module in real time; Establish a control signal generation module, and embed a main prediction network, an action instruction generation strategy unit, an index strategy unit and a mode control unit in the control signal generation module to generate control instructions, send them to the drive control module, and control the action instruction generation mode and control the switching of the observation mode and the training mode through the mode control unit; Establish a training module, and embed a target network, an experience replay cache unit, a remote monitoring unit, and an automatic reset control unit in the training module. The experience replay cache unit is used to store RGBD data, excitation values, obstacle avoidance instructions, and docking instructions. The target network is used to generate an expected value of the probability of an action instruction based on the data stored in the experience replay cache unit. The remote monitoring unit is used to send and receive manual remote control instructions. The automatic reset control unit is used to generate a reset action instruction based on the driving action instruction, the obstacle avoidance instruction, and the docking instruction. Establishing a driving control module, the driving control module is used to receive control instructions and control the operation of the mining vehicle based on the control instructions; A security management module is established, and a prediction supervision network is built into the security management module. The prediction supervision network is trained independently of the main prediction network and the target network. During training, the prediction supervision network obtains training data from the experience replay cache unit.

[0017] Furthermore, the main prediction network generates a rotation action instruction probability, and the action instruction generation strategy unit generates a rotation action instruction. The action instruction generation strategy is as follows: Assume that the probability of the main prediction network generating a rotation action command is , the action instruction generation strategy outputs the steering instruction as , the custom random steering instruction is expressed as , the custom random value is , the steering command random adjustment threshold is , ,but: ; After the control signal generation module receives the obstacle avoidance instruction sent by the safety management module, the obstacle avoidance method is as follows: Randomly adjust the threshold of action instructions Increases, the randomness of the output steering command increases, and the mine car is driven to reselect the driving direction that effectively avoids obstacles.

[0018] Furthermore, the autonomous driving subsystem is constructed by the following method: Obtain RGBD data through the vehicle vision module; Establish a control signal generation module, and build a main prediction network and an action instruction generation strategy unit in the control signal generation module, generate control instructions according to the RGBD data sent by the vehicle vision module, and send them to the drive control module; and receive obstacle avoidance instructions or parking instructions from the safety management module; Establishing a driving control module, the driving control module is used to receive control instructions and control the operation of the mining vehicle based on the control instructions; A safety management module is established, and a prediction supervision network is built into the safety management module. According to the RGBD data sent by the vehicle-mounted vision module, an obstacle avoidance instruction or a parking instruction is generated and sent to the control signal generation module.

[0019] The beneficial effects of the present invention are: Based on theoretical research and experimental verification in a simulated real environment, the present invention establishes a DQN network training environment and an autonomous driving navigation and obstacle avoidance control system in a real open-pit mine. Aiming at the unmanned transportation operation scenario in an open-pit mine, the present invention effectively makes up for the lack of mobility of traditional navigation technology and path planning strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It shows an operation framework diagram of a training subsystem of a mine car automatic driving vision-guided following docking system according to an embodiment of the present invention; Figure 2 A structural diagram of a vehicle-mounted vision module according to an embodiment of the present invention is shown; Figure 3 A schematic diagram of an RGBD data flow according to an embodiment of the present invention is shown; Figure 4 A structural diagram of a control signal generating module according to an embodiment of the present invention is shown; Figure 5 A structural diagram of a training module according to an embodiment of the present invention is shown; Figure 6 A schematic diagram of an experience data transmission path according to an embodiment of the present invention is shown; Figure 7 It shows a structural diagram of an excitation module according to an embodiment of the present invention; Figure 8A running framework diagram of a mine car automatic driving vision-guided following docking system in an automatic driving subsystem according to an embodiment of the present invention is shown; Fig. 9 A deep Q network model architecture diagram according to an embodiment of the present invention is shown; Fig.10 A structural diagram of a preprocessing module according to an embodiment of the present invention is shown; Fig.11 A residual module structure diagram according to an embodiment of the present invention is shown; Fig.12 A residual unit structure diagram according to an embodiment of the present invention is shown; Fig.13 A schematic diagram of data flow during a deep Q network training process according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0021] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0022] The specific implementation of the present invention is further described in detail below in conjunction with the drawings and examples.

[0023] Embodiment 1:

[0024] Deep Q-Network (DQN) is an algorithm based on deep learning and reinforcement learning, which is used to solve the Markov decision process (MDP) problem in discrete action space. DQN uses two neural networks, the main network (Q network) and the target network (Target Q network). The main network is used to select actions, and the target network is used to calculate the target Q value. At present, deep Q networks have achieved excellent performance in intelligent control of Atari games and intelligent decision-making in chess games. In terms of autonomous driving, research has been carried out on following the driving trajectory of real cars on open streets, automatic lane changes, and autonomous driving in simulated urban road scenes. The training of deep Q networks requires a large amount of experience data, and an environment that allows potentially unsafe operations to be performed is required to collect experience data, which poses a huge challenge to providing a large amount of experience data in the real world.

[0025] Based on this, an embodiment of the present invention provides a mine car automatic driving vision-guided following and docking system, which is based on a DQN reinforcement learning network and is divided into two subsystems, namely a training subsystem in the training phase and an automatic driving subsystem in the automatic driving phase. The training subsystem is used to train the network model or module required for training in the system, and the trained weight parameters are applied to the corresponding network model of the automatic driving subsystem to realize the automatic following and docking of the mine car based on vision guidance.

[0026] like Figure 1 The figure shows the operation framework diagram of the training subsystem of the vision-guided follow-up docking system for automatic driving of a mine car. The subsystem includes an on-board vision module, a drive control module, a control signal generation module, a training module, an excitation module and a safety management module. The connection relationship of each module is as follows: the signal output end of the vision module is connected to the signal input end of the excitation module, the signal output end of the excitation module is connected to the signal input end of the training module and the control signal generation module, the signal output end of the training module is connected to the signal input end of the safety management module and the control signal generation module, and the signal output end of the control signal generation module is connected to the signal input end of the training module and the drive control module. The vehicle vision module obtains RGBD data and feeds it to the excitation module. The excitation module generates excitation values ​​and obstacle avoidance / docking instructions based on the RGBD data. The RGBD data and the generated excitation values ​​and obstacle avoidance / docking instructions are fed to the control signal generation module. The excitation values ​​and obstacle avoidance / docking instructions are fed to the training module. The training module can also obtain manual remote control instructions, which are fed to the drive control module through the control signal generation module. At the same time, the training module also interacts with the control signal generation module. The interactive data includes RGBD data, action instructions and network weights. The training module also generates a reset action instruction, which is directly fed to the drive control module. The training module also feeds RGBD data, excitation values, and obstacle avoidance (docking) instructions to the safety management module. The control instructions generated by the control signal generation module, such as motion action instructions, are fed to the drive control module. The drive control module controls the mine car based on the corresponding instructions obtained from the control signal generation module. The mode control unit in the control signal generation module controls the action instruction generation mode, including the main prediction network generation, reset instruction or manual remote control instruction, and controls the switching of observation mode and training mode. The manual remote control instruction has the highest execution priority, followed by the reset action instruction, and the main prediction network generation action instruction has the lowest priority.

[0027] like Figure 2The figure shows the structure of the vehicle-mounted vision module. The vehicle-mounted vision module has a built-in depth camera, an RGB camera and a video preprocessing unit. The depth camera and the RGB camera respectively shoot the depth video and RGB video in front of the mine car, and input them into the video preprocessing unit synchronously. The video preprocessing unit extracts the frame images from the depth video stream and the RGB video stream at a predetermined frame interval, and preprocesses the depth video frame images using the nearest neighbor interpolation method and the filtering method to fill in and repair the missing pixel values ​​of the depth map caused by black objects, smooth surfaces, transparent objects, parallax effects, etc.; the depth map and RGB map alignment algorithm are used to unify the coordinate system and merge them into a 4-channel RGBD map, as shown in Figure 3 shown.

[0028] The drive control module receives the driving action instructions generated by the control signal generation module, including steering instructions and speed instructions, and converts the steering instructions and speed instructions into electrical signals respectively and sends them to the steering control system and power control system to control the driving state of the mine car.

[0029] like Figure 4 The control signal generation module is shown in FIG. 1 . The control signal generation module has a built-in main prediction network, an action instruction generation strategy unit, an index strategy unit, and a mode control unit. In the training phase, the control signal generation module is divided into an observation mode and a training mode.

[0030] In the observation mode, the control signal generation module obtains RGBD data, incentive value, and obstacle avoidance / docking instruction data from the excitation module. The RGBD data is input into the main prediction network to calculate the probability of the mine car driving action instruction, and then input into the action instruction generation strategy unit to calculate and generate the mine car driving action instruction data, which is transmitted to the drive control module, and the RGBD data and its corresponding action instructions, incentive value, and obstacle avoidance / docking instruction data are sent to the experience replay cache unit; at the same time, the obstacle avoidance / docking instruction of the excitation module is received, which triggers the stop observation to generate action instructions, stops driving the mine car, and stops transmitting experience data to the experience replay cache unit.

[0031] In the training mode, the main prediction network obtains RGBD data and action command data from the experience replay cache unit, uses the RGBD data to calculate the probability of the mine car driving action command, combines the action command data of the experience replay cache unit, and calculates and generates the action command index value through the index strategy unit to carry out the main prediction network training.

[0032] During the training phase, the control signal generation module can receive manual remote control instructions, trigger the stop observation and generate action instructions, stop driving the mine car, and stop transmitting experience data to the experience playback cache unit.

[0033] like Figure 5The figure shows the result of the training module. The training module has a built-in target network, an experience playback cache unit, a remote monitoring unit, and an automatic reset control unit.

[0034] The target network and the main prediction network form a deep Q network, and the target network and the main prediction network have the same network structure. The target network receives the weight parameters of the main prediction network at a predetermined period and updates itself. The target network provides the expected value of the action instruction probability for the main prediction network training.

[0035] The experience playback buffer unit stores the RGBD data, operation action instructions, stimulus values, and obstacle avoidance / docking instruction data transmitted by the control signal generation module. Figure 6 As shown, the experience replay cache unit randomly selects experience data in batches, sends the current frame RGBD data and the corresponding running action instruction data in the experience data to the main prediction network of the control signal generation module; and sends the corresponding excitation value of the current frame and the next frame RGBD data in the experience data to the target network. In addition, the experience replay cache unit randomly selects experience data in batches, sends the current frame RGBD data and the current corresponding obstacle avoidance / docking instruction (3-bit vector) in the experience data to the safety management module.

[0036] The experience replay cache unit sets the maximum storage capacity. If the maximum storage capacity is exceeded, the experience replay cache unit will delete the earliest experience data and store the new data to always maintain the maximum storage capacity. If the obstacle avoidance / docking instruction is a valid value, the experience replay cache unit will stop storing the new data.

[0037] The remote monitoring unit obtains manual control instructions for the operation of the mine car and sends them to the control signal generation module and the drive control module in turn, so as to intervene in the operation control during the training phase and avoid dangerous accidents.

[0038] The automatic reset control unit stores the driving action command signal sequence sent by the mine car control signal generation module from the starting position, and receives the obstacle avoidance / docking command sent by the excitation module. The obstacle avoidance / docking command can trigger the automatic reset control unit to send a driving action command to the drive control module in the reverse order of the cached action command signal sequence, and the motion action command is the negative value of the original command, and the rotation action command is the original command, driving the mine car back to the starting position.

[0039] like Figure 7 As shown in the figure, it is the structure diagram of the excitation module. The excitation module has a built-in target detection network and an excitation value calculation unit. Figure 7As shown. The excitation module obtains RGBD data from the on-board vision module and transmits it to the target detection network. The target detection network pre-trains the weight parameters for identifying the following and stopping targets and obstacles inside and outside the path, realizes target detection of the following and stopping targets and obstacles inside and outside the path, and sends the detection result data to the excitation value calculation unit. The excitation value calculation unit calculates and generates the distance data of the mine car deviating from the stopping target according to the detection results of the target detection network, and then calculates the excitation value and the docking instruction of the approaching stopping target, and calculates the relative distance data between the mine car and the obstacles inside and outside the path, generates the obstacle avoidance instruction according to the strategy, and transmits it to the training module in real time.

[0040] The security management module has a built-in prediction supervision network, which is a classification network. The prediction supervision network is trained independently of the main prediction network and the target network. The prediction supervision network obtains training data from the experience replay cache unit.

[0041] like Figure 8 The figure shows the operation framework diagram of the automatic driving subsystem of a mine car automatic driving visual guidance following docking system. The subsystem includes an on-board vision module (including RGB camera, depth camera, video preprocessing unit), a drive control module, a control signal generation module (including main prediction network, action instruction generation strategy unit, mode control unit), and a safety management module (prediction supervision network). The network weight parameters of the automatic driving subsystem are provided by the corresponding network module of the training subsystem. During the automatic driving stage, the training subsystem does not participate in the control of the mine car movement, and the training subsystem can continuously optimize the weight parameters of each network. The depth camera and RGB camera of the on-board vision module synchronously shoot the depth video and RGB video in front of the mine car. After the video preprocessing unit, the 4-channel RGBD time series data stream is synthesized and output and sent to the control signal generation module and the safety management module. The control signal generation module calculates and generates the probability of the mine car running action instruction, and the action instruction generation strategy unit calculates and generates the mine car driving action instruction. The drive control module receives the running action instruction generated by the control signal generation module to control the driving state of the mine car. The safety management module has a built-in prediction supervision network, which generates obstacle avoidance / docking instructions based on RGBD data and sends the docking instructions to the control signal generation module. The built-in mode control unit in the control signal generation module adjusts the control strategy for the operation of the mine car and drives the mine car to park; it sends obstacle avoidance instructions to the control signal generation module to readjust the driving direction.

[0042] The following is a detailed description of the specific architecture and mechanism of the neural network model involved in the mine car autonomous driving visual guidance following docking system. Among them, the neural network model architecture used in the system includes a deep Q network architecture, a target detection network, and a prediction supervision network.

[0043] The deep Q network includes a main prediction network and a target network. The main prediction network and the target network have the same network structure. The target network receives the weight parameters of the main prediction network and updates itself according to a predetermined period.

[0044] The main prediction network uses the ResNet residual network as the baseline network, redesigns the feature classification head, and adopts a two-way feature classification output, such as Fig. 9 The specific design method is as follows: the main prediction network inputs RGBD 4-channel data, which is converted into Data, such as Fig.10 As shown. Input 4 residual network modules (ResidualModel) in sequence and output The data is then passed through the average pooling layer and output The data is then passed through two fully connected layers and a SoftMax layer, outputting two The probability data represent the speed probability and steering probability of the mine car action command, and finally input the action command generation strategy to generate the speed command and steering command. Among them, the residual network module contains n residual units connected in series, such as Fig.11 and Fig.12 shown.

[0045] The action instruction generation strategy unit generates action instructions based on the output of the main prediction network. Let the state-action value output of the main prediction network be expressed as , , including the speed value and turn value , where the speed value The format is shown in Table 1.

[0046] Table 1 Speed ​​value Format No. 0 No. 1 No. 2 Speed+ Speed ​​maintenance speed- Turning to value The format is shown in Table 2.

[0047] Table 2 Turning value Format No. 0 No. 1 No. 2 Turn left + Steering hold Turn right + Assume that the action instruction generation strategy outputs the action instruction represented as , the custom random action instruction is expressed as , the custom random value is , the action command randomly adjusts the threshold to , ,but: ; in, Indicates the bit index value of the maximum value in the output action instruction. Indicates a custom random speed command. Indicates a custom random turn instruction; During the training phase, as the number of training rounds increases, the , so that Continuously tending towards action-value dependence The action command randomly adjusts the threshold The reduction strategy is expressed as: ; in, is the initial value, is the pre-set minimum value, is the reduction value for each training round, epoches is the number of training rounds.

[0048] During the autonomous driving phase, if the control signal generation module receives an obstacle avoidance command from the safety management module, the steering command random adjustment threshold is increased. , the probability of random adjustment of steering instructions increases, mainly by adjusting the steering and reselecting the effective driving direction to avoid obstacles in the driving path.

[0049] The target detection network can adopt a single-stage target detection YOLO series network, for example, YOLOv5, and add an improved spatial channel hybrid attention mechanism network module after the SPPF network layer in the backbone network.

[0050] The target detection network outputs the coordinates of the upper left and lower right corners of the target detection frame and the obstacle detection frame inside and outside the path in the RGBD data stream RGB frame image relative to the upper left corner of the image, and sends them to the stimulus value calculation unit. The stimulus value calculation unit performs stimulus value calculation and obstacle avoidance / docking value calculation.

[0051] The specific process of the excitation value calculation is as follows: Assume that the coordinates of the upper left and lower right corners of the RGB image detection box of the target object being followed and docked are expressed as and , then the minimum distance between the RGB image detection frame and the D depth map is ,but: ; in, Denotes the depth map D The distance value of the coordinate point; The incentive value is: ; in is the incentive value, is the empirical coefficient, arctg is the inverse tangent function.

[0052] Empirical coefficient Make It can be distributed more evenly in the interval [0,1], rather than biased towards 0 or 1.

[0053] According to the characteristics of the inverse tangent function, it can be analyzed that as the distance between the mine cart and the docking target increases, the r value decreases. At the same time, when the distance increases to a certain extent, the rate of decrease of the r value decreases, indicating that even if the distance between the mine cart and the docking target is far, the incentive value will not be too small. On the contrary, as the distance between the mine cart and the docking target decreases, the r value increases. At the same time, when the mine cart and the docking target are close, the r value increases rapidly, indicating that when the mine cart is close to the docking target, it obtains a larger positive incentive (reward).

[0054] The specific process of obstacle avoidance / docking value calculation is as follows: When the distance between the minecart and the obstacle is lower than a certain threshold, the obstacle avoidance action is triggered; when the distance between the minecart and the docking target is lower than a certain threshold, the docking stop is triggered.

[0055] In the RGBD image output by the vehicle vision module, establish the forward path of the mine car. x q Indicates the displacement of the left edge of the forward path relative to the left side of the image, x t Represents the forward path width.

[0056] The coordinates of the upper left and lower right corners of the RGB image detection box of the following target object inspection are expressed as and , the coordinates of the upper left and lower right corners of the obstacle RGB image detection box are expressed as and .

[0057] Calculate D depth map The corresponding path coverage decision value : ; in, Indicates a preset maximum value; According to the coordinates of the upper left and lower right corners of the detection frame of the following docked target and , calculate the minimum distance value of the target object: ; According to the coordinates of the upper left and lower right corners of the obstacle detection frame and , calculate the minimum obstacle distance value: ; Assume the output value of obstacle avoidance is b, the docking output value is ,but: ; in, is the minimum distance threshold between the minecart and the obstacle, It is the minimum distance threshold to the docking target; when the minimum distance between the mine car and the obstacle is less than the minimum threshold, the obstacle avoidance command is output; when the minimum distance between the mine car and the target is less than the minimum threshold, the docking command is output.

[0058] The prediction supervision network uses the resnet50 classification network, and adds an improved spatial channel hybrid attention mechanism network module between the last residual module and the average pooling layer. The prediction supervision network outputs a 3-bit vector, which represents the normal driving, obstacle avoidance instruction bit and parking instruction bit respectively.

[0059] On the premise of knowing the specific structures of the deep Q network architecture, target detection network and prediction detection network, the following will introduce in detail the training process of the deep Q network architecture, target detection network and prediction detection network. When the system is in the training stage, the target detection network, prediction supervision network and deep Q network are trained respectively.

[0060] Object detection network training: The target detection network is trained in advance. During the training phase, the mine car is manually controlled, and the built-in depth camera and RGB camera in the on-board vision module synchronously shoot forward to follow the docked target objects and obstacles inside and outside the path, establish training data sets and verification data sets, and carry out target detection network training to achieve accurate detection of follow-up docked targets and obstacles inside and outside the path.

[0061] After completing the target detection network training and performance verification test, the target detection network outputs the stimulus value and obstacle avoidance / docking instructions, which are used for the prediction supervision network and deep Q network training.

[0062] Predictive supervision network training: The prediction supervision network is a classification network and is trained independently of the deep Q network (main prediction network, target network). During the training phase, the safety management module obtains batches of random experience data from the experience replay cache unit. The experience data includes RGBD data and its corresponding obstacle avoidance / docking instructions. The obstacle avoidance / docking instructions are used as data classification labels, the RGBD data and its corresponding data classification labels are constructed into a data dictionary, and the batch data dictionary is constructed into a data list. The training process includes the following steps: Step 10: Equalization processing of sample experience data.

[0063] The classification label format is set as shown in Table 3.

[0064] Table 3 Classification label format

[0065] Step 10 can be used to perform sample experience data equalization processing through the following steps: Step 11, retrieve the data dictionary with data classification label [0,0,1] from the empirical data dictionary list, and calculate the number of data dictionaries; Step 12: Randomly extract the same number of data dictionaries with data classification labels of [1,0,0] and [0,1,0] from the empirical data dictionary list; Step 13, performing data enhancement operation on the RGBD data in the data dictionary extracted in steps 11 and 12, keeping the data classification label value unchanged, and expanding the number of elements in the data dictionary list to a predetermined number; Step 14: sort and shuffle the data dictionary list after data amplification.

[0066] Step 20: Divide the data dictionary list into a training set and a validation set according to a predetermined ratio.

[0067] Step 30: Conduct prediction supervision network training.

[0068] After completing the prediction supervision network training and performance verification test, the prediction supervision network outputs obstacle avoidance / docking instructions, which are applied to the automatic driving stage of the mine car.

[0069] Deep Q network training. Fig.13 As shown in Figure 1, the deep Q network training phase is divided into observation mode and training mode. The observation mode and the training mode are performed alternately.

[0070] In the observation mode, the control signal generation module obtains RGBD data from the vehicle vision module, the main prediction network calculates and generates the probability of the mine car driving action command, the input action command generation strategy unit calculates and generates the action command, and sends it to the drive control module to drive the mine car to drive automatically, and simultaneously sends the RGBD data and the corresponding action command data to the experience playback cache unit, such as Fig.13 shown.

[0071] In observation mode, the main prediction network does not perform network weight parameter training.

[0072] In training mode, the deep Q network is trained by the following steps: Step 1: The main prediction network and the target network are first initialized with network weight parameters.

[0073] Step 2: The main prediction network calculates the output action instruction probability.

[0074] The control signal generation module obtains batch random experience data from the experience playback buffer unit, such as Fig.13As shown in Figure 1, the experience data includes the current frame RGBD data and the action instruction data corresponding to the current frame. The current frame RGBD data is input into the main prediction network to calculate the probability of the mine car driving action instruction.

[0075] Step 3: The index strategy unit calculates and outputs the action instruction index value.

[0076] The action command probability and the action command data corresponding to the current frame are input into the index strategy unit together, and the driving action command speed probability index value and the steering probability index value are calculated and output respectively. The index algorithm is as follows: Assume that the action instruction data corresponding to the current frame is expressed as , the main prediction network generates the action instruction speed probability expressed as , the turning probability is expressed as , concatenated row by row , The output action command speed probability index value is represented by , the turning probability index value is expressed as ,but: ; Step 4: The target network calculates and outputs the cumulative incentive.

[0077] The target network obtains batches of random experience data from the experience replay cache unit, such as Fig.13 As shown, the experience data includes the corresponding incentive value of the current frame and the RGBD data of the next frame. The RGBD data of the next frame is input into the target network to calculate the probability of the mine car driving action instruction, including the speed probability of the action instruction and the steering probability. The action instruction probability and the incentive value corresponding to the current frame are input into the cumulative incentive strategy unit to calculate the speed target cumulative incentive value and the steering target cumulative incentive value of the mine car driving action instruction respectively. The cumulative incentive is as follows: Assume that the corresponding stimulus value of the current frame is expressed as , ,in, represents speed incentive, represents the steering incentive, and the speed probability of the mine car driving action instruction is expressed as , the turning probability is expressed as , then the speed target cumulative incentive value is expressed as: ; The cumulative incentive value of the steering target is expressed as: ; in, is the empirical coefficient.

[0078] Step 5: Calculate the loss function based on the action instruction probability index value and the cumulative incentive.

[0079] The speed loss function is: ; The steering loss function is: ; in, is the number of training samples in each round.

[0080] Step 6: Update the weights of the main prediction network.

[0081] Use speed loss, steering loss and Adam gradient descent optimization algorithm to backpropagate the main prediction network and update the network weight parameters.

[0082] Step 7: Update the target network weights.

[0083] Set the target network weight parameter update cycle. After the main prediction network is trained for a certain number of cycles, input the main prediction network weight parameters into the target network, and update the target network weight parameters to be the same as the main prediction network.

[0084] The above steps 2 to 7 are executed cyclically to continuously reduce the predicted motion speed command loss and the steering command loss.

[0085] Embodiment 2: An embodiment of the present invention further provides a method for automatic driving of a mine car under visual guidance and following and docking, the method comprising: Step S1, obtaining RGBD data stream through the vehicle-mounted vision module; Step S2, establishing an incentive module, and in the incentive module, a target detection network and an incentive value calculation unit are built-in, the target detection network is used to pre-train and identify the weight parameters of the following and stopping target and the obstacles inside and outside the path according to the input RGBD data stream, realize the target detection of the following and stopping target and the obstacles inside and outside the path, and send the detection result data to the incentive value calculation unit; the incentive value calculation unit is used to calculate and generate the distance data of the mine car deviating from the following and stopping target according to the detection result data, and then calculate the incentive value and the docking instruction of the approaching stopping target, and calculate the relative distance data between the mine car and the obstacles inside and outside the path, generate the obstacle avoidance instruction according to the strategy, and transmit it to the training module in real time; Step S3, establishing a control signal generation module, and in the control signal generation module, a main prediction network, an action instruction generation strategy unit, an index strategy unit and a mode control unit are built in, so as to generate a control instruction and send it to the drive control module; Step S4, establishing a training module, and in the training module, a target network, an experience replay cache unit, a remote monitoring unit and an automatic reset control unit are built in, the experience replay cache unit is used to store RGBD data streams, excitation values, obstacle avoidance instructions and docking instructions, the target network is used to generate an expected value of the probability of an action instruction according to the data stored in the experience replay cache unit, the remote monitoring unit is used to send and receive manual remote control instructions, and the automatic reset control unit is used to generate a reset action instruction according to the driving action instruction, the obstacle avoidance instruction and the docking instruction; Step S5, establishing a drive control module, wherein the drive control module is used to receive control instructions and control the operation of the mining vehicle based on the control instructions; Step S6, establish a safety management module, and embed a prediction supervision network in the safety management module. The prediction supervision network is trained independently of the main prediction network and the target network. During training, the prediction supervision network obtains training data from the experience replay cache unit. In the automatic driving stage, the prediction supervision network generates obstacle avoidance / docking instructions based on RGBD data, sends the docking instructions to the control signal generation module, adjusts the motion signal generation strategy, and drives the mine car to park; sends obstacle avoidance instructions to the control signal generation module to readjust the driving direction.

[0086] It should be noted that the mine car automatic driving vision-guided following docking method belongs to the same technical concept as the previously described system, and has the same technical principles and beneficial effects, so it will not be repeated here.

[0087] The above implementation modes are only used to illustrate the present invention, but not to limit the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention. The patent protection scope of the present invention should be defined by the claims.

Claims

1. A mine car automatic driving vision-guided following docking system, characterized in that: The system includes a training subsystem and an automatic driving subsystem, wherein the training subsystem includes an on-board vision module, a drive control module, a control signal generating module, a training module, an excitation module and a safety management module; wherein the signal output end of the on-board vision module is connected to the signal input end of the excitation module, the signal output end of the excitation module is connected to the signal input end of the training module and the control signal generating module, the signal output end of the training module is connected to the signal input end of the safety management module and the control signal generating module, and the signal output end of the control signal generating module is connected to the signal input end of the training module and the drive control module; The automatic driving subsystem includes an on-board vision module, a drive control module, a control signal generating module and a safety management module; wherein, the signal output end of the on-board vision module is connected to the signal input end of the control signal generating module and the safety management module, the signal output end of the control signal generating module is connected to the signal input end of the drive control module, and the signal output end of the safety management module is connected to the signal input end of the control signal generating module.

2. The mine car automatic driving visual guidance following docking system according to claim 1, characterized in that: In the training subsystem, the vehicle-mounted vision module has a built-in RGB camera, a depth camera and a video preprocessing unit, which are used to obtain RGBD data and send the RGBD data to the excitation module; The excitation module has a built-in target detection network and an excitation value calculation unit. The target detection network is used to pre-train and identify weight parameters of the following and stopping target objects and obstacles inside and outside the path according to the input RGBD data, realize target detection of the following and stopping target objects and obstacles inside and outside the path, and send the detection result data to the excitation value calculation unit; the excitation value calculation unit is used to calculate and generate distance data of the mine car deviating from the following and stopping target object according to the detection result data, and then calculate the excitation value and the docking instruction of the approaching following and stopping target object, and calculate the relative distance data between the mine car and the obstacles inside and outside the path, generate obstacle avoidance instructions according to the strategy, and transmit them to the training module in real time; The control signal generation module has a built-in main prediction network, an action instruction generation strategy unit, an index strategy unit and a mode control unit, and is used to generate control instructions and send them to the drive control module.

3. The mine car automatic driving visual guidance following docking system as claimed in claim 2, characterized in that: The training module has a built-in target network, an experience replay cache unit, a remote monitoring unit and an automatic reset control unit. The experience replay cache unit is used to store RGBD data, excitation values, obstacle avoidance instructions and docking instructions. The target network is used to generate an expected value of the probability of an action instruction based on the data stored in the experience replay cache unit. The remote monitoring unit is used to send and receive manual remote control instructions. The automatic reset control unit is used to generate a reset action instruction based on the driving action instruction, obstacle avoidance instruction and docking instruction. The driving control module is used to receive a driving action instruction and control the mining vehicle based on the driving action instruction, wherein the driving action instruction includes a steering instruction and a speed instruction; The security management module has a built-in prediction supervision network, which obtains training data from the built-in experience replay cache unit of the training module and is trained independently of the main prediction network and the target network; In the autonomous driving subsystem, the on-board vision module has a built-in RGB camera, a depth camera and a video preprocessing unit, which are used to obtain RGBD data and send the RGBD data to the control signal generation module and the safety management module; The control signal generation module calculates and generates the probability of the mine car running action instruction, and the action instruction generation strategy unit calculates and generates the mine car driving action instruction; The driving control module is used to receive a driving action instruction and control the mining vehicle based on the driving action instruction, wherein the driving action instruction includes a steering instruction and a speed instruction; The safety management module has a built-in prediction supervision network, which generates obstacle avoidance / docking instructions based on RGBD data and sends them to the control signal generation module. The control signal generation module readjusts the driving direction action instructions based on the obstacle avoidance instructions, or generates a docking operation action instruction based on the docking instruction.

4. The mine car automatic driving visual guidance following docking system as claimed in claim 3, characterized in that: The control signal generating module alternately executes the observation mode and the training mode during the training phase; When the observation mode is executed, RGBD data, incentive value, obstacle avoidance instruction and docking instruction are obtained from the incentive module, the RGBD data is input into the main prediction network, the probability of the mine car driving action instruction is calculated, and the action instruction generation strategy unit is input to calculate and generate the mine car driving action instruction, which is transmitted to the drive control module, and the RGBD data and its corresponding action instruction, incentive value and obstacle avoidance instruction, docking instruction data are sent to the experience replay cache unit; at the same time, based on the obstacle avoidance instruction and the docking instruction, the stop observation action instruction is triggered to drive the mine car to stop running, and stop transmitting the experience data to the experience replay cache unit; When executing the training mode, the main prediction network obtains RGBD data and action instruction data from the experience replay cache unit, calculates the probability of the mine car driving action instruction using the RGBD data, combines the action instruction data in the experience replay cache unit, calculates and generates the action instruction index value through the index strategy unit, and carries out the main prediction network training; The control signal generation module has a built-in mode control unit to control the action instruction generation mode and control the switching of the observation mode and the training mode. The action instruction generation mode includes the main prediction network generating action instructions, obtaining and sending manual remote control instructions, obtaining and sending obstacle avoidance instructions and docking instructions.

5. The mine car automatic driving visual guidance following docking system as claimed in claim 3, characterized in that: The control signal generation module has a built-in action instruction generation strategy unit, which generates action instructions in the following way: Get the state-action value output by the main prediction network , ,in is the speed value, To turn to value; The action instruction generation strategy unit outputs the action instruction as , the custom random action instruction is expressed as , the custom random value is , the action command randomly adjusts the threshold to , , the speed command is calculated by the following formula a v and steering instructions a r : ; in, Indicates the bit index value of the maximum value in the output action instruction. Indicates a custom random speed command. Indicates a custom random turn instruction; During the training phase, as the number of training rounds increases, the , so that Constantly tending to the dependent state-action value, the action instruction randomly adjusts the threshold The reduction strategy is expressed as: ; in, is the initial value, is the pre-set minimum value, is the reduction value for each training round, epoches is the number of training rounds.

6. The mine car automatic driving visual guidance following docking system as claimed in claim 2, characterized in that: The training module has a built-in target network that forms a deep Q network with the main prediction network. The target network has the same network structure as the main prediction network. The target network receives weight parameters of the main prediction network at a predetermined period and updates itself. The target network provides expected values ​​of action instruction probabilities for the main prediction network training.

7. The mine car automatic driving vision-guided following docking system according to claim 1, characterized in that: The training module has a built-in experience replay cache unit to store the RGBD data, running action instructions, incentive values, obstacle avoidance instructions and docking instructions transmitted by the control signal generation module; in the training stage, the experience replay cache unit randomly selects experience data in batches, and sends the current frame RGBD data and the running action instruction data corresponding to the current frame in the experience data to the main prediction network in the control signal generation module; sends the current frame corresponding incentive value and the next frame RGBD data in the experience data to the target network; the experience replay cache unit randomly selects experience data in batches, and sends the current frame RGBD data, the current corresponding obstacle avoidance instructions and the docking instructions in the experience data to the safety management module; The experience replay cache unit sets a maximum storage capacity. When the storage capacity exceeds the maximum storage capacity, the experience replay cache unit deletes the earliest experience data and stores new data, always maintaining the maximum storage capacity. If the obstacle avoidance instruction and / or the docking instruction is a valid value, the experience replay cache unit suspends storing the new data.

8. The mine car automatic driving vision-guided following docking system according to claim 1, characterized in that: The training module has a built-in automatic reset control unit, which stores the driving action instruction signal sequence sent by the control signal generating module from the starting position, and receives the obstacle avoidance instruction and the docking instruction sent by the excitation module. The obstacle avoidance instruction and the docking instruction trigger the automatic reset control unit to send a reset action instruction to the control signal generating module in the reverse order of the cached action instruction signal sequence, and then send it to the driving control module through the control signal generating module. The reset action instruction includes a motion action instruction and a rotation action instruction. The motion action instruction is the negative value of the original instruction, and the rotation action instruction is the original instruction, which drives the mine car to return to the starting position.

9. The mine car automatic driving vision-guided following docking system according to claim 1, characterized in that: The excitation module has a built-in target detection network, which calculates and outputs the coordinates of the upper left and lower right corners of the following and stopping target object and the obstacle detection box inside and outside the path relative to the upper left corner of the RGB image data, and sends them to the excitation value calculation unit; The incentive value calculation unit is used to perform incentive value calculation and obstacle avoidance / docking value calculation to obtain incentive values, obstacle avoidance instructions and docking instructions; The incentive value calculation includes: The coordinates of the upper left and lower right corners of the RGB image detection box of the target object being followed and docked are expressed as and , then the minimum distance between the RGB image detection frame and the D depth map is ,but: ; in, Denotes the depth map D The distance value of the coordinate point; The incentive value is calculated by the following formula: ; in is the incentive value, is the empirical coefficient, arctg is the inverse tangent function; The obstacle avoidance / docking value calculation includes: In the RGBD image output by the vehicle vision module, establish the forward path of the mine car. x q Indicates the displacement of the left edge of the forward path relative to the left side of the image, x t represents the forward path width; The coordinates of the upper left and lower right corners of the RGB image detection box of the following target object inspection are expressed as and , the coordinates of the upper left and lower right corners of the obstacle RGB image detection box are expressed as and ; Calculate D depth map The corresponding path coverage decision value : ; in, Indicates a preset maximum value; According to the coordinates of the upper left and lower right corners of the detection frame of the following docked target and , calculate the minimum distance value of the target object: ; According to the coordinates of the upper left and lower right corners of the obstacle detection frame and , calculate the minimum distance to the obstacle: ; Assume the obstacle avoidance output value is b , the docking output value is ,but: ; in, is the minimum distance threshold between the minecart and the obstacle, It is the minimum distance threshold to the docking target; when the minimum distance between the mine car and the obstacle is less than the minimum threshold, the obstacle avoidance command is output; when the minimum distance between the mine car and the target is less than the minimum threshold, the docking command is output.

10. The mine car automatic driving vision-guided following docking system according to claim 3, characterized in that: In the training phase, the deep Q network consisting of the target detection network in the excitation module, the prediction supervision network in the safety management module, the main prediction network in the control signal generation module, and the target network in the training module is trained respectively; The target detection network is trained in advance; during training, the mine car is manually controlled to run, the on-board vision module forward shoots and follows the docking target object, obstacles inside and outside the path, outputs RGBD data, divides and establishes a training data set and a verification data set, and carries out the target detection network training. After completing the target detection network training and performance verification test, the target detection network outputs an incentive value and an obstacle avoidance instruction and a docking instruction, which are applied to the prediction supervision network and the deep Q network training; The predictive supervision network belongs to a classification network and is trained independently of the deep Q network. During training, the safety management module randomly obtains experience data in batches from the experience playback cache unit. The experience data includes RGBD data and its corresponding obstacle avoidance instructions and docking instructions. The obstacle avoidance instructions and docking instructions are used as RGBD data classification labels, and training data sets and verification data sets are divided to carry out the predictive supervision network. After completing the training and performance verification test of the predictive supervision network, the predictive supervision network outputs obstacle avoidance / docking instructions, which are applied to obstacle avoidance or docking in the autonomous driving stage. The deep Q network training phase is divided into an observation mode and a training mode; the observation mode and the training mode are performed alternately; In the observation mode, the control signal generation module obtains RGBD data from the vehicle vision module, the main prediction network calculates and generates the probability of the mine car driving action command, the input action command generation strategy unit calculates and generates the action command, and sends it to the drive control module to drive the mine car to drive automatically, and synchronously sends the RGBD data and the corresponding action command data to the experience playback cache unit. In the observation mode, the main prediction network does not perform network weight parameter training; In training mode, the Deep Q Network is trained as follows: Initialize the network weight parameters of the main prediction network and the target network respectively; The main prediction network calculates and outputs the probability of action instructions: the control signal generation module randomly obtains experience data in batches from the experience playback cache unit. The experience data includes the current frame RGBD data and the action instruction data corresponding to the current frame; the current frame RGBD data is input into the main prediction network to calculate the probability of the mine car driving action instruction; The index strategy unit calculates and outputs the action instruction index value; The target network calculates the output cumulative incentive; Calculate the loss function according to the action instruction probability index value and the accumulated incentive; the loss function includes running speed loss and steering loss; Main prediction network weight update: Use speed loss, steering loss and Adam gradient descent optimization algorithm to back-propagate the main prediction network and update the network weight parameters; Target network weight update: Set the target network weight parameter update cycle. After the main prediction network is trained for a certain number of cycles, the main prediction network weight parameters are input into the target network, and the target network weight parameters are updated to be the same as the main prediction network.

11. A mine car automatic driving vision-guided following docking method, characterized in that: Based on the mine car automatic driving vision-guided following and docking system according to any one of claims 1 to 10, the method comprises: Obtain RGBD data through the vehicle vision module; Establish an incentive module, and build a target detection network and an incentive value calculation unit in the incentive module. The target detection network is used to pre-train and identify weight parameters of the following and stopping target object and the adjacent obstacles inside and outside the path according to the input RGBD data, perform target detection on the following and stopping target object and the adjacent obstacles inside and outside the path, and send the detection result data to the incentive value calculation unit; the incentive value calculation unit is used to calculate and generate distance data of the mine car deviating from the following and stopping target object according to the detection result data, and then calculate the incentive value and the stopping instruction of the approaching stopping target object, and calculate the relative distance data between the mine car and the obstacles inside and outside the path, generate obstacle avoidance instructions according to the strategy, and transmit them to the training module in real time; Establish a control signal generation module, and embed a main prediction network, an action instruction generation strategy unit, an index strategy unit and a mode control unit in the control signal generation module to generate control instructions, send them to the drive control module, and control the action instruction generation mode and control the switching of the observation mode and the training mode through the mode control unit; Establish a training module, and embed a target network, an experience replay cache unit, a remote monitoring unit, and an automatic reset control unit in the training module. The experience replay cache unit is used to store RGBD data, excitation values, obstacle avoidance instructions, and docking instructions. The target network is used to generate an expected value of the probability of an action instruction based on the data stored in the experience replay cache unit. The remote monitoring unit is used to send and receive manual remote control instructions. The automatic reset control unit is used to generate a reset action instruction based on the driving action instruction, the obstacle avoidance instruction, and the docking instruction. Establishing a driving control module, the driving control module is used to receive control instructions and control the operation of the mining vehicle based on the control instructions; A security management module is established, and a prediction supervision network is built into the security management module. The prediction supervision network is trained independently of the main prediction network and the target network. During training, the prediction supervision network obtains training data from the experience replay cache unit.

12. The mine car automatic driving vision-guided following docking method according to claim 11, characterized in that: The main prediction network generates the probability of the rotation action instruction, and the action instruction generation strategy unit generates the rotation action instruction. The action instruction generation strategy is as follows: Assume that the probability of the main prediction network generating a rotation action command is , the action instruction generation strategy outputs the steering instruction as The custom random turn instruction is expressed as , the custom random value is , the steering command random adjustment threshold is , ,but: ; After the control signal generation module receives the obstacle avoidance instruction sent by the safety management module, the obstacle avoidance method is as follows: Randomly adjust the threshold of action instructions Increases, the randomness of the output steering command increases, and the mine car is driven to reselect the driving direction that effectively avoids obstacles.

13. The mine car automatic driving vision-guided following docking method according to claim 11, characterized in that: The autonomous driving subsystem is constructed as follows: Obtain RGBD data through the vehicle vision module; Establish a control signal generation module, and embed a main prediction network and an action instruction generation strategy unit in the control signal generation module, generate a control instruction according to the RGBD data sent by the vehicle vision module, and send it to the drive control module; And receive obstacle avoidance instructions or docking instructions from the safety management module; Establishing a driving control module, the driving control module is used to receive control instructions and control the operation of the mining vehicle based on the control instructions; A safety management module is established, and a prediction supervision network is built into the safety management module. According to the RGBD data sent by the vehicle-mounted vision module, an obstacle avoidance instruction or a parking instruction is generated and sent to the control signal generation module.

Citation Information

Patent Citations

  • Automatic driving method, training method and related device

    CN109901572A

  • Target following and dynamic obstacle avoidance control method for speed difference slip steering vehicle

    CN110989576A

  • Pure vision automatic driving control system and method based on improved RTFNet, and medium

    CN114708568A

  • Automatic driving obstacle avoidance system and method based on machine vision

    CN117037115A

  • Automatic driving control model, training method and training system thereof, equipment and medium

    CN117406720A