Automatic unloading method and system for mine tipping bucket type tramcar based on visual positioning

Through the method of visual positioning and robotic arm collaboration, the problems of low efficiency and poor safety of automatic unloading of bucket mine cars have been solved, efficient and safe automatic unloading has been achieved, and labor costs have been reduced.

CN120747572AActive Publication Date: 2025-10-03CHINA MINMETALS CHANGSHA MINING RES INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511242784.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-03
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

The existing automatic unloading technology for dump trucks is difficult to meet the needs of trucks with different structures, and has low efficiency, poor safety, and high labor costs.

Method used

A visual positioning method is adopted, through the YOLOv8s recognition model and the traditional machine vision ORB feature point detection algorithm, combined with a robotic arm to locate, fix, flip and unload the mine car, use electromagnets to achieve automatic unloading, and use a stereo vision positioning system to achieve precise positioning and safety pin identification.

Benefits of technology

It improves the unloading efficiency of mine cars, reduces labor costs, improves safety performance, and meets the automatic unloading needs of mine cars with different structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747572A_ABST
    Figure CN120747572A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic unloading method and system for a mine tipping bucket type tramcar based on visual positioning. The method comprises the following steps: acquiring historical high-angle images of the tipping bucket mine car and a mechanical arm; labeling the historical high-angle shot image to obtain a first training data set; based on YOLOv8s, a recognition model of the initial mine car and the mechanical arm is constructed; training by using the first training data set to obtain an identification model of the mine car and the mechanical arm; constructing a mine car safety pin identification model based on a traditional machine vision ORB feature point detection algorithm; a stereoscopic vision positioning system is established by utilizing a mine car and mechanical arm identification model and a mine car safety pin identification model, and automatic unloading of the mine tilting cart is achieved. According to the method, artificial intelligence recognition is introduced, and a self-adaptive control strategy is matched, so that the unloading efficiency and safety performance of the mine tipping bucket type tramcar are improved, and the labor cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of automatic control, and in particular relates to a method and system for automatically unloading a mine dump truck based on visual positioning. Background Art

[0002] Dump trucks are key equipment used in mining to transport ore, gangue, and other materials. Automatic unloading devices are designed to address the low efficiency and poor safety of traditional manual or semi-manual unloading methods, and are a crucial component of mining machinery and equipment. Modern mining operations tend toward large-scale, high-efficiency operations, while dump trucks mostly utilize traditional manual or semi-manual unloading methods (such as manual flipping and track tilting). Traditional unloading methods are all manual, resulting in low parking positioning accuracy and an inability to achieve automatic unloading, making it difficult to meet the production capacity requirements of continuous transportation lines. Previous manual unloading methods required close proximity operation, posing safety risks such as overturning of the truck and slipping of materials. Furthermore, this method was inefficient and had high labor costs.

[0003] The existing automatic unloading technology of bucket-type mine cars has major defects. For example, CN 220410552 designed a new type of bucket-type mine car, which modified the main structure of the existing bucket-type mine car and realized automatic unloading under external force loading. However, it cannot automatically unload the mine cars used in existing mines, and it is difficult to meet the needs of different mines for mine cars with different structures. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, one of the purposes of the present invention is to provide an automatic unloading method for a mine bucket truck based on visual positioning, so as to achieve high unloading efficiency, low labor cost and high safety performance.

[0005] The second object of the present invention is to provide a system for realizing the automatic unloading method of a mine bucket truck based on visual positioning.

[0006] The present invention provides a method for automatically unloading a mine bucket-type mine car based on visual positioning, comprising the following steps:

[0007] S1. Obtain historical overhead images of the dump truck and robotic arm;

[0008] S2. Annotate the historical overhead images to obtain a first training dataset;

[0009] S3. Build an initial mine cart and robotic arm recognition model based on YOLOv8s; and train it using the first training dataset to obtain a mine cart and robotic arm recognition model.

[0010] S4. Build a mine car safety pin recognition model based on the traditional machine vision ORB feature point detection algorithm;

[0011] S5. Using the mine car and robotic arm recognition models obtained in step S3 and the mine car safety pin recognition model obtained in step S4, a stereoscopic vision positioning system is established to achieve automatic unloading of the mine dump truck.

[0012] Furthermore, the present invention provides two robotic arms and a slide rail for the robotic arms on the side of the bucket-type mine car track, and the robotic arms are used to position, fix, flip, unload, and straighten the mine car; the two robotic arms are placed on different bases, which are located on the slide rail for the robotic arms, and an electromagnet is provided at the end of each robotic arm; the first robotic arm is responsible for fixing the mine car and flipping the mine car for unloading; the second robotic arm is responsible for opening the locking mechanism; the first robotic arm is also responsible for pushing the mine car down and straightening the mine car through the adsorption of the electromagnet after unloading.

[0013] The second mechanical arm is responsible for opening the locking mechanism. Specifically, the second mechanical arm is responsible for moving to the safety pin position, energizing the electromagnet to adsorb the safety pin, and then rotating the safety pin to open the safety. After the mine car returns to the center position, the power is turned on again to adsorb the safety pin and then rotate it into the safety slot.

[0014] A hand-eye integrated camera is installed at the end of the second robotic arm to identify the position of the safety pin;

[0015] A top camera is installed on the top of the working area of ​​the dump truck to identify the relative position of the truck and the robotic arm.

[0016] In step S3, the initial mine car and robotic arm recognition model includes a backbone network, a first metal anti-reflective attention module, a second metal anti-reflective attention module, a third metal anti-reflective attention module, a feature fusion module, a secondary attention enhancement module, a first detection head module, a second detection head module, a third detection head module, and a result output module;

[0017] The input image is first processed by the backbone network module to obtain the first to third feature maps of three different scales: high resolution, medium resolution, and low resolution; the three feature maps are respectively input into the first to third metal anti-reflective attention modules for processing, and the processing results are input into the feature fusion module; the feature fusion module performs feature fusion on the received feature maps and outputs the fourth to sixth feature maps of three different scales: high resolution, medium resolution, and low resolution; the fourth feature map is input into the secondary attention enhancement module for processing, and the processing results are input into the first detection head module; the fifth feature map is directly input into the second detection head module; the sixth feature map is directly input into the third detection head module; the first to third detection head modules respectively perform mine car bounding box positioning and robotic arm bounding box positioning according to the input feature maps, and input the positioning results into the result output module; the result output module obtains the dump truck positioning map and the robotic arm positioning map as the output of the model based on the input data.

[0018] The recognition model of the mining cart and the robotic arm is as follows: based on YOLOv8s, the backbone network is replaced by ShuffleNetV2, and the metal anti-reflective attention module is embedded after the three feature layers output by it; the metal anti-reflective attention module includes a channel attention module, a spatial attention module, and a reflection suppression module connected in series; the metal anti-reflective attention module strengthens the metal texture feature channel through the channel attention module, the spatial attention module focuses on the contact area of ​​the robotic arm, and the reflection suppression module actively reduces high light interference; the feature fusion module adopts an improved FPN+PAN structure, and applies a 1x1 convolution kernel to each feature layer to be fused output by the three metal anti-reflective attention modules, unifying the channels to 256 dimensions (all input feature maps are uniformly compressed / expanded to 256 through 1×1 convolution). channel), multi-scale metal features are fused respectively through the FPN and PAN dual-path transmission mechanism (in the feature fusion process, FPN and PAN are retained and the feature streams of the two paths are transmitted), and finally the fourth to ninth feature maps of three different scales are output; among them, the high-resolution fourth feature map is obtained by FPN fusion; the high-resolution fifth feature map is obtained by PAN fusion; the medium-resolution sixth feature map is obtained by FPN fusion; the medium-resolution seventh feature map is obtained by PAN fusion; the low-resolution eighth feature map is obtained by FPN fusion; the low-resolution ninth feature map is obtained by PAN fusion; secondary attention enhancement is applied to the highest-resolution feature map before the detection head, and finally the three-scale detection head is used to achieve precise positioning of the dump truck, and the final result is output through the result output module.

[0019] The channel attention module is specifically: given a feature map ,Generate channel descriptors through global average pooling: ,in is the cth channel of the input channel, with a dimension of (height × width); is the spatial resolution of the feature map, corresponding to the feature size of the neural network output; is the spatial coordinate index, traversing the entire two-dimensional feature map plane; is the global feature descriptor of the c-th channel, obtained by spatial dimension compression; the weights are calculated using a two-layer MLP (multi-layer perceptron): ,in is the Sigmoid function; is the ReLU activation function; The first layer fully connected weight matrix has the dimension ( is the compression ratio), is the input dimension, is the output dimension; is the second layer fully connected weight matrix ; is the final generated channel attention weight vector; by learning the contribution of each channel, the expression of key feature channels such as the metal texture of the minecart is enhanced. The final output is ,in .

[0020] The spatial attention module is specifically: the input is ,against Apply max pooling and average pooling along the channel axis to generate , fused through convolutional layers: ,in is a 7×7 convolution kernel; It is a feature map generated by the maximum pooling operation of the channel dimension, with a dimension of ; is the feature map generated by average pooling of the channel dimension, with dimension ;[;] is a tensor concatenation operation that merges two feature maps along the channel dimension; this allows the network to focus on the contact area between the end of the robotic arm and the safety pin, suppressing background interference. The output is .

[0021] The reflection suppression module is specifically as follows: given an input feature map and ,in is the number of channels, and are the height and width of the feature map respectively. First, the statistical characteristics of the feature map are calculated. The highlight mask is generated by analyzing the statistical characteristics of the feature map. The specific steps include:

[0022] Calculate the channel average feature map using the following formula: ,in is the number of channels of the neural network feature map, represents the average feature intensity at each spatial location, For the Feature maps of channels;

[0023] Calculate the channel contrast feature map using the following formula: ,in Represents the feature variance at each spatial location;

[0024] The mean map and variance map are concatenated along the channel dimension and fused through the convolution layer, using the following formula: ,in: Represents a tensor concatenation operation, generating dimensions of Feature map of for Convolution kernel, output single channel feature map , used to encode high-light responses;

[0025] Apply the Sigmoid function to map the highlight response to the range [0, 1] to generate a highlight mask, which is calculated using the following formula: in is the highlight mask, a value close to 1 indicates a highlight area, and a value close to 0 indicates a non-highlight area; Sigmoid function Ensure the mask is smooth and differentiable;

[0026] Apply the highlight mask to the original feature map to reduce the feature response in the highlight area, calculated using the following formula: in is the feature map output by the spatial attention module, represents element-wise multiplication, To generate suppression weights; the weight of the highlight area is close to 0, and the feature is suppressed; the weight of the non-highlight area is close to 1, and the feature is retained.

[0027] The mine car safety pin recognition model described in step S5 includes the following steps: based on the mechanical structure knowledge of the safety pin (key areas such as the thread junction and the edge of the positioning hole), a region of interest (ROI) is preset to guide the algorithm to focus on the effective area with rich and stable texture to reduce background interference; within the preset ROI, the FAST algorithm is used to detect candidate corner points, and the feature points with the highest response value are screened through non-maximum suppression to retain the corner points with strong geometric stability;

[0028] A Gaussian difference pyramid is established to handle the detection requirements of safety pins of different sizes. Each layer of the pyramid is generated by a downsampling rate of 1.2 to achieve multi-scale feature coverage. According to the distance between the camera on the second robotic arm and the target depth And dynamic calculation of safety pin nominal size: ;in is the camera sensor pixel size, is the minimum recognizable physical size, is the safety factor;

[0029] Calculate the image moment within a 15px radius neighborhood of the feature point and determine the main direction angle Direction correction range is ±15°, to adapt to the deflection angle of the safety pin during assembly; is the main direction angle of the feature point (range 0-360°); is the image moment, representing the center of gravity of the grayscale distribution of the pixel neighborhood; for Gray value at ; , and is the order of the moment; the feature description is carried out in sequence using the steps of domain selection, Gaussian smoothing, and binary descriptor construction;

[0030] The neighborhood selection covers the entire safety pin and matches the physical size of the safety pin; the Gaussian smoothing suppresses surface reflection noise of the metal; the binary descriptor is set to adopt a 256-bit Steered BRIEF mode to support rotation errors;

[0031] The binary descriptor is constructed using the following formula: ,in is the pixel coordinate pair after rotation; is a binary index (ORB uses a 256-bit descriptor, i = 0 ~ 255); is the grayscale query function; is the rotation matrix;

[0032] The DBSCAN clustering algorithm is applied to the ORB feature points extracted from the whole image, and the clustering parameters are adaptive; wherein the key parameters are dynamically set; the key parameters include the area radius and minimum points The dynamic setting is specifically: based on the camera internal parameters ( , , main point , ) and current distance , calculate the minimum physical diameter of the safety pin The minimum pixel size corresponding to the current viewing angle set up ( is the empirical coefficient); set Based on the statistical lower bound of the effective texture area of ​​the safety pin;

[0033] The candidate clusters obtained by DBSCAN were geometrically verified to ensure that they conformed to the mechanical structure of the safety pin. Based on the safety pin's characteristic of being an elongated cylinder with a typical aspect ratio of ≥4:1 and the mechanical constraints of the safety pin, the covariance matrix of the points within the feature clusters was calculated and principal component analysis was performed. Clusters with an aspect ratio greater than 3:1 and a second principal component variance less than 15% were retained.

[0034] The cluster that passes the above clustering and verification is identified as the target safety pin. If there are multiple candidate clusters, the cluster with the most expected size and the highest confidence is selected;

[0035] After selecting the target feature cluster, extract its convex hull center point as the target position reference and calculate the depth distance to the target using the following formula: , the moving goal of the robot arm 2 is to minimize value; where Determined through calibration experiments, is the scaling factor, is the offset, is the distance to the target depth.

[0036] Step S6 specifically includes the following steps:

[0037] S61. After the mine car stops, the top camera captures the relative position image of the mine car and the robotic arm, determines the relative position of the robotic arm and the mine car, and controls the slide rail to move the robotic arm so that the horizontal position of the second robotic arm and the safety pin are aligned based on the recognition model of the mine car and the robotic arm.

[0038] S62. Robotic Arm 1 extends, causing the electromagnet at its top to contact the mine cart. When powered on, the electromagnet adheres to the mine cart, maintaining the relative position of the mine cart and the robotic arm. A camera on Robotic Arm 2 captures an image of the safety pin to determine whether Robotic Arm 2 and the mine cart's safety pin are vertically aligned. An object detection algorithm is used to locate the safety pin, and based on the result, Robotic Arm 2 is controlled to align with the mine cart's safety pin.

[0039] S63. After both cameras have determined that they have reached their positions, the electromagnet is powered on and the latch is searched for by shaking using a random perturbation search algorithm. A determination is made as to whether the latch is correctly attached to the latch. If not, the electromagnet is turned off and the robotic arm is fine-tuned by reusing the random perturbation search algorithm. If the robotic arm has moved to the appropriate position, the second robotic arm extends, causing the electromagnet at its top to contact the locking mechanism. After the locking mechanism is attracted, the second robotic arm rotates to unlock the locking mechanism. After unlocking the locking mechanism, the second robotic arm maintains the electromagnet's attraction to the latch and operates synchronously with the first robotic arm.

[0040] S64. Robotic arm 1 continues to extend, pushing the mine cart, causing it to flip over and complete unloading; after unloading is completed, robotic arm 1 contracts to return the mine cart to the center position, and then robotic arm 2 rotates to close the locking mechanism; finally, the power is cut off, causing the electromagnet to lose its magnetic force, and the two robotic arms return to their initial positions.

[0041] Step S61 specifically includes: using the top camera and the mine car relative position recognition model to detect the mine car, obtain the coordinates of the mine car bounding box, and calculate the distance between the mine car position and the standard dumping position using the following formula: ;in For the target location ; is the top of the minecart's bounding box, The bottom edge of the minecart's bounding box. For minecarts in The distance between the axis and the standard unloading position; by moving the robot arm 1 and the robot arm 2 synchronously along the track, the robot arm 1 Axis coordinates and minecart Axis coordinates are aligned, and a robotic arm is used to absorb the dump truck through an electromagnet and then minimized. , place the minecart in Move to the specified target area in coordinates.

[0042] The random perturbation search algorithm uses a sine wave superimposed on random noise to generate a shaking trajectory: Where A is the amplitude, which is less than the safety pin hole tolerance; is the frequency, avoiding the resonant frequency of the robotic arm; The zero-mean Gaussian noise enhances robustness; when the contact force at the end of the robot arm Exceeding the threshold Stop shaking and judge the position of the robotic arm and the pin.

[0043] Furthermore, in actual use, it also includes the use of a multi-sensor collaborative power-off protection mechanism, specifically: real-time detection of motor operating current , when exceeding the rated value The protection is triggered when the resistance torque reaches 120%; the abnormal resistance torque is detected by the joint torque sensor. The threshold is set to 150% of the rated torque, forming a double check with the current monitoring to avoid false triggering.

[0044] The present invention also provides a system for realizing the automatic unloading method of a mine dump truck based on visual positioning, comprising an image acquisition module, an image annotation module, a model construction and training module, a safety pin recognition model construction module, and a mine dump truck automatic unloading module;

[0045] The image acquisition module acquires historical overhead images of the dump truck and the robotic arm, and uploads the data to the image annotation module;

[0046] The image annotation module annotates the historical overhead images based on the received data to obtain a training dataset and uploads the data to the model building and training module;

[0047] The model building and training module builds an initial mine car and robotic arm recognition model based on YOLOv8s according to the received data, and uses the training data set for training to obtain the mine car and robotic arm recognition model, and uploads the data to the mine dump truck automatic unloading module;

[0048] The safety pin recognition model building module builds a mine car safety pin recognition model based on the traditional machine vision ORB feature point detection algorithm and uploads the data to the mine dump truck automatic unloading module;

[0049] The automatic unloading module of the mine dump truck establishes a three-dimensional vision positioning system based on the received data, using the recognition model of the mine truck and the robotic arm, as well as the recognition model of the mine truck safety pin, to realize the automatic unloading of the mine dump truck.

[0050] The present invention discloses an automatic unloading method and system for a mine dump truck based on visual positioning. By introducing artificial intelligence recognition and coordinating with an adaptive control strategy, the unloading efficiency and safety performance of the mine dump truck are improved, and labor costs are reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 Schematic diagram of the process of the present invention;

[0052] Figure 2 Schematic diagram of the robotic arm and mining car device in the method of the present invention; wherein 1 is robotic arm 1; 2 is robotic arm 2; 3 is the top camera;

[0053] Figure 3 Schematic diagram of the structure of the identification model of the mining car and the robotic arm in the present invention;

[0054] Figure 4 Schematic diagram of the operation flow of the reflection suppression module in the identification model of the mining car and the robotic arm of the present invention;

[0055] Figure 5 Schematic diagram of the structure of the system of the present invention. DETAILED DESCRIPTION

[0056] The present invention provides a method for automatically unloading a mine dump truck based on visual positioning, the flow diagram of which is shown in FIG. Figure 1 As shown, the following steps are included:

[0057] S1. Obtain an overhead image of the dump truck and the robotic arm.

[0058] The present invention is provided with two mechanical arms and slide rails for the mechanical arms on the side of the bucket mine car track, and the mechanical arms are used to position, fix, flip, unload and straighten the mine car. The schematic diagram of the mechanical arm and the mine car device is shown in FIG. Figure 2 As shown; the two robotic arms are placed on different bases, which are located on the slide rails for the robotic arms. An electromagnet is set at the end of each robotic arm; robotic arm one is responsible for fixing the mine cart and turning the mine cart for unloading; robotic arm two is responsible for opening the locking mechanism; robotic arm one is also responsible for pushing the mine cart down and straightening the mine cart through the adsorption of the electromagnet after unloading.

[0059] The second mechanical arm is responsible for opening the locking mechanism. Specifically, the second mechanical arm is responsible for moving to the safety pin position, energizing the electromagnet to adsorb the safety pin, and then rotating the safety pin to open the safety. After the mine car returns to the center position, the power is turned on again to adsorb the safety pin and then rotate it into the safety slot.

[0060] A hand-eye integrated camera is installed at the end of the second robotic arm to identify the position of the safety pin;

[0061] A top camera is installed on the top of the working area of ​​the dump truck to identify the relative position of the truck and the robotic arm.

[0062] S2. Annotate the historical overhead images to obtain a first training dataset;

[0063] S3. Build an initial mine cart and robotic arm recognition model based on YOLOv8s; and train it using the first training dataset to obtain a mine cart and robotic arm recognition model.

[0064] In step S3, the structural diagram of the identification model of the initial mine car and the robotic arm is as follows: Figure 3 As shown, it includes a backbone network, a first metal anti-reflective attention module, a second metal anti-reflective attention module, a third metal anti-reflective attention module, a feature fusion module, a secondary attention enhancement module, a first detection head module, a second detection head module, a third detection head module, and a result output module;

[0065] The input image is first processed by the backbone network module to obtain the first to third feature maps of three different scales: high resolution, medium resolution, and low resolution; the three feature maps are respectively input into the first to third metal anti-reflective attention modules for processing, and the processing results are input into the feature fusion module; the feature fusion module performs feature fusion on the received feature maps and outputs the fourth to sixth feature maps of three different scales: high resolution, medium resolution, and low resolution; the fourth feature map is input into the secondary attention enhancement module for processing, and the processing results are input into the first detection head module; the fifth feature map is directly input into the second detection head module; the sixth feature map is directly input into the third detection head module; the first to third detection head modules respectively perform mine car bounding box positioning and robotic arm bounding box positioning according to the input feature maps, and input the positioning results into the result output module; the result output module obtains the dump truck positioning map and the robotic arm positioning map as the output of the model based on the input data.

[0066] The recognition model of the mining cart and the robotic arm is as follows: based on YOLOv8s, the backbone network is replaced by ShuffleNetV2, and the metal anti-reflective attention module is embedded after the three feature layers output by it; the metal anti-reflective attention module includes a channel attention module, a spatial attention module, and a reflection suppression module connected in series; the metal anti-reflective attention module strengthens the metal texture feature channel through the channel attention module, the spatial attention module focuses on the contact area of ​​the robotic arm, and the reflection suppression module actively reduces high light interference; the feature fusion module adopts an improved FPN+PAN structure, and applies a 1x1 convolution kernel to each feature layer to be fused output by the three metal anti-reflective attention modules, unifying the channels to 256 dimensions (all input feature maps are uniformly compressed / expanded to 256 through 1×1 convolution). Channel), multi-scale metal features are fused respectively through the FPN and PAN dual-path transmission mechanism (in the feature fusion process, FPN and PAN are retained and the feature streams of the two paths are transmitted), and the fourth to ninth feature maps of three different scales are finally output. Among them, the high-resolution fourth feature map is obtained by FPN path fusion; the high-resolution fifth feature map is obtained by PAN path fusion; the medium-resolution sixth feature map is obtained by FPN path fusion; the medium-resolution seventh feature map is obtained by PAN path fusion; the low-resolution eighth feature map is obtained by FPN path fusion; the low-resolution ninth feature map is obtained by PAN path fusion; secondary attention enhancement is applied to the highest-resolution feature map before the detection head, and finally the three-scale detection head is used to achieve accurate positioning of the dump truck, and the final result is output through the result output module.

[0067] The channel attention module is specifically: given a feature map ,Generate channel descriptors through global average pooling: ,in is the cth channel of the input channel, with a dimension of (height × width); is the spatial resolution of the feature map, corresponding to the feature size of the neural network output; is the spatial coordinate index, traversing the entire two-dimensional feature map plane; is the global feature descriptor of the c-th channel, obtained by spatial dimension compression; the weights are calculated using a two-layer MLP (multi-layer perceptron): ,in is the Sigmoid function; is the ReLU activation function; The first layer fully connected weight matrix has the dimension ( is the compression ratio), is the input dimension, is the output dimension; is the second layer fully connected weight matrix ; is the final generated channel attention weight vector; by learning the contribution of each channel, the expression of key feature channels such as the metal texture of the minecart is enhanced. The final output is ,in .

[0068] The spatial attention module is specifically: the input is ,against Apply max pooling and average pooling along the channel axis to generate , fused through convolutional layers: ,in is a 7×7 convolution kernel; It is a feature map generated by the maximum pooling operation of the channel dimension, with a dimension of ; is the feature map generated by average pooling of the channel dimension, with dimension ;[;] is a tensor concatenation operation that merges two feature maps along the channel dimension; this allows the network to focus on the contact area between the end of the robotic arm and the safety pin, suppressing background interference. The output is .

[0069] The reflection suppression module is specifically as follows: given an input feature map and ,in is the number of channels, and are the height and width of the feature map respectively. First, the statistical characteristics of the feature map are calculated. The highlight mask is generated by analyzing the statistical characteristics of the feature map. The processing flow diagram is as follows Figure 4 As shown, the specific steps include:

[0070] Calculate the channel average feature map using the following formula: ,in is the number of channels of the neural network feature map, represents the average feature intensity at each spatial location, For the Feature maps of channels;

[0071] Calculate the channel contrast feature map using the following formula: ,in Represents the feature variance at each spatial location;

[0072] The mean map and variance map are concatenated along the channel dimension and fused through the convolution layer, using the following formula: ,in: Represents a tensor concatenation operation, generating dimensions of Feature map of for Convolution kernel, output single channel feature map , used to encode high-light responses;

[0073] Apply the Sigmoid function to map the highlight response to the range [0, 1] to generate a highlight mask, which is calculated using the following formula: in is the highlight mask, a value close to 1 indicates a highlight area, and a value close to 0 indicates a non-highlight area; Sigmoid function Ensure the mask is smooth and differentiable;

[0074] Apply the highlight mask to the original feature map to reduce the feature response in the highlight area, calculated using the following formula: in is the feature map output by the spatial attention module, represents element-wise multiplication, To generate suppression weights; the weight of the highlight area is close to 0, and the feature is suppressed; the weight of the non-highlight area is close to 1, and the feature is retained.

[0075] S4. Build a mine car safety pin recognition model based on the traditional machine vision ORB feature point detection algorithm;

[0076] The mine car safety pin recognition model described in step S4 includes the following steps: based on the mechanical structure knowledge of the safety pin (key areas such as the thread junction and the edge of the positioning hole), a region of interest (ROI) is preset to guide the algorithm to focus on the effective area with rich and stable texture to reduce background interference; within the preset ROI, the FAST algorithm is used to detect candidate corner points, and the feature points with the highest response value are screened through non-maximum suppression to retain the corner points with strong geometric stability;

[0077] A Gaussian difference pyramid is established to handle the detection requirements of safety pins of different sizes. Each layer of the pyramid is generated by a downsampling rate of 1.2 to achieve multi-scale feature coverage. According to the distance between the camera on the second robotic arm and the target depth And dynamic calculation of safety pin nominal size: ;in is the camera sensor pixel size, is the minimum recognizable physical size, is the safety factor;

[0078] Calculate the image moment within a 15px radius neighborhood of the feature point and determine the main direction angle Direction correction range is ±15°, to adapt to the deflection angle of the safety pin during assembly; is the main direction angle of the feature point (range 0-360°); is the image moment, representing the center of gravity of the grayscale distribution of the pixel neighborhood; for Gray value at ; , and is the order of the moment; the feature description is carried out in sequence using the steps of domain selection, Gaussian smoothing, and binary descriptor construction;

[0079] The neighborhood selection covers the entire safety pin and matches the physical size of the safety pin; the Gaussian smoothing suppresses surface reflection noise of the metal; the binary descriptor is set to adopt a 256-bit Steered BRIEF mode to support rotation errors;

[0080] The binary descriptor is constructed using the following formula: ,in is the pixel coordinate pair after rotation; is a binary index (ORB uses a 256-bit descriptor, i = 0 ~ 255); is the grayscale query function; is the rotation matrix;

[0081] The DBSCAN clustering algorithm is applied to the ORB feature points extracted from the whole image, and the clustering parameters are adaptive; wherein the key parameters are dynamically set; the key parameters include the area radius and minimum points The dynamic setting is specifically: based on the camera internal parameters ( , , main point , ) and current distance , calculate the minimum physical diameter of the safety pin The minimum pixel size corresponding to the current viewing angle set up ( is the empirical coefficient); set Based on the statistical lower bound of the effective texture area of ​​the safety pin;

[0082] The candidate clusters obtained by DBSCAN were geometrically verified to ensure that they conformed to the mechanical structure of the safety pin. Based on the safety pin's characteristic of being an elongated cylinder with a typical aspect ratio of ≥4:1 and the mechanical constraints of the safety pin, the covariance matrix of the points within the feature clusters was calculated and principal component analysis was performed. Clusters with an aspect ratio greater than 3:1 and a second principal component variance less than 15% were retained.

[0083] The cluster that passes the above clustering and verification is identified as the target safety pin. If there are multiple candidate clusters, the cluster with the most expected size and the highest confidence is selected;

[0084] After selecting the target feature cluster, extract its convex hull center point as the target position reference and calculate the depth distance to the target using the following formula: , the moving goal of the robot arm 2 is to minimize value; where Determined through calibration experiments, is the scaling factor, is the offset, is the distance to the target depth.

[0085] S5. Using the mine car and robotic arm recognition models obtained in step S3 and the mine car safety pin recognition model obtained in step S4, a stereoscopic vision positioning system is established to achieve automatic unloading of the mine dump truck.

[0086] Step S5 specifically includes the following steps:

[0087] S51. After the mine car stops, the top camera captures the image of the relative position of the mine car and the robotic arm, determines the relative position of the robotic arm and the mine car as a whole, and controls the slide rail to move the robotic arm so that the horizontal position of the second robotic arm and the safety pin are aligned through the mine car and robotic arm recognition model.

[0088] S52. Arm 1 extends, causing the electromagnet at its top to contact the mine cart. When powered on, the electromagnet adheres to the cart, maintaining the relative position of the cart and the arm. A camera on Arm 2 captures an image of the safety pin to determine whether the vertical alignment of Arm 2 and the safety pin is established. An object detection algorithm is used to locate the safety pin, and based on the result, Arm 2 is controlled to align with the safety pin.

[0089] S53. After both cameras determine that they have reached the desired position, the electromagnet is powered on and the latch is searched for by shaking the latch using a random perturbation search algorithm. A determination is made as to whether the latch is properly attached to the latch. If not, the electromagnet is turned off and the robotic arm is fine-tuned by reusing the random perturbation search algorithm. If the robotic arm has moved to the proper position, the second robotic arm extends, causing the electromagnet at the top to contact the locking mechanism. After the locking mechanism is engaged, the second robotic arm rotates to unlock the locking mechanism. After unlocking the locking mechanism, the second robotic arm maintains the electromagnet's attachment to the latch and operates synchronously with the first robotic arm.

[0090] S54. Robotic arm 1 continues to extend, pushing the mine cart, causing it to flip over and complete unloading; after unloading is completed, robotic arm 1 contracts to return the mine cart to the center position, and then robotic arm 2 rotates to close the locking mechanism; finally, the power is cut off, causing the electromagnet to lose its magnetic force, and the two robotic arms return to their initial positions.

[0091] Step S51 specifically includes: using the top camera and the mine car relative position recognition model to perform target detection on the mine car, obtaining the coordinates of the mine car bounding box, and calculating the distance between the mine car position and the standard dumping position using the following formula: ;in For the target location ; is the top of the minecart's bounding box, The bottom edge of the minecart's bounding box. For minecarts in The distance between the axis and the standard unloading position; by moving the robot arm 1 and the robot arm 2 synchronously along the track, the robot arm 1 Axis coordinates and minecart Axis coordinates are aligned, and a robotic arm is used to absorb the dump truck through an electromagnet and then minimized. , place the minecart in Move to the specified target area in coordinates.

[0092] The random perturbation search algorithm uses a sine wave superimposed on random noise to generate a shaking trajectory: Where A is the amplitude, which is less than the safety pin hole tolerance; is the frequency, avoiding the resonant frequency of the robotic arm; The zero-mean Gaussian noise enhances robustness; when the contact force at the end of the robot arm Exceeding the threshold Stop shaking and judge the position of the robotic arm and the pin.

[0093] Furthermore, in actual use, it also includes the use of a multi-sensor collaborative power-off protection mechanism, specifically: real-time detection of motor operating current , when exceeding the rated value The protection is triggered when the resistance torque reaches 120%; the abnormal resistance torque is detected by the joint torque sensor. The threshold is set to 150% of the rated torque, forming a double check with the current monitoring to avoid false triggering.

[0094] The present invention also provides a system for realizing the automatic unloading method of a mine dump truck based on visual positioning, and its structural diagram is shown as follows: Figure 5 As shown, it includes image acquisition module, image annotation module, model construction and training module, safety pin recognition model construction module, and mine dump truck automatic unloading module;

[0095] The image acquisition module acquires historical overhead images of the dump truck and the robotic arm, and uploads the data to the image annotation module;

[0096] The image annotation module annotates the historical overhead images based on the received data to obtain a training dataset and uploads the data to the model building and training module;

[0097] The model building and training module builds an initial mine car and robotic arm recognition model based on YOLOv8s according to the received data, and uses the training data set for training to obtain the mine car and robotic arm recognition model, and uploads the data to the mine dump truck automatic unloading module;

[0098] The safety pin recognition model building module builds a mine car safety pin recognition model based on the traditional machine vision ORB feature point detection algorithm and uploads the data to the mine dump truck automatic unloading module;

[0099] The automatic unloading module of the mine dump truck establishes a three-dimensional vision positioning system based on the received data, using the recognition model of the mine truck and the robotic arm, as well as the recognition model of the mine truck safety pin, to realize the automatic unloading of the mine dump truck.

Claims

1. A method for automatic unloading of a mine dump truck based on visual positioning, characterized in that: The following steps are involved: S1. Obtain historical overhead images of the dump truck and robotic arm; S2. Annotate the historical overhead images to obtain a first training dataset; S3. Build an initial mine cart and robotic arm recognition model based on YOLOv8s; and train it using the first training dataset to obtain a mine cart and robotic arm recognition model. S4. Build a mine car safety pin recognition model based on the traditional machine vision ORB feature point detection algorithm; S5. Using the mine car and robotic arm recognition models obtained in step S3 and the mine car safety pin recognition model obtained in step S4, a stereoscopic vision positioning system is established to achieve automatic unloading of the mine dump truck.

2. The automatic unloading method of a mine dump truck based on visual positioning according to claim 1 is characterized in that: Two robotic arms and slide rails are installed on the side of the bucket mine car track. The robotic arms are used to position, fix, flip, unload, and straighten the mine car. The two robotic arms are placed on different bases, which are located on the slide rails. An electromagnet is installed at the end of each robotic arm. Robotic arm one is responsible for fixing the mine car and flipping the mine car for unloading; robotic arm two is responsible for opening the locking mechanism. Robotic arm one is also responsible for pushing the mine car down and straightening the mine car after unloading through the adsorption of the electromagnet. The second mechanical arm is responsible for opening the locking mechanism. Specifically, the second mechanical arm is responsible for moving to the safety pin position, energizing the electromagnet to adsorb the safety pin, and then rotating the safety pin to open the safety. After the mine car returns to the center position, the power is turned on again to adsorb the safety pin and then rotate it into the safety slot. A hand-eye integrated camera is installed at the end of the second robotic arm to identify the position of the safety pin; A top camera is installed on the top of the working area of ​​the dump truck to identify the relative position of the truck and the robotic arm.

3. The automatic unloading method of a mine dump truck based on visual positioning according to claim 2 is characterized in that: In step S3, the initial mine car and robotic arm recognition model includes a backbone network, a first metal anti-reflective attention module, a second metal anti-reflective attention module, a third metal anti-reflective attention module, a feature fusion module, a secondary attention enhancement module, a first detection head module, a second detection head module, a third detection head module, and a result output module; The input image is first processed by the backbone network module to obtain the first to third feature maps of three different scales: high resolution, medium resolution, and low resolution. These three feature maps are then fed into the first to third metal anti-reflective attention modules for processing, and the results are fed into the feature fusion module. The feature fusion module performs feature fusion on the received feature maps and outputs three feature maps of different scales: the fourth and fifth feature maps with high resolution, the sixth and seventh feature maps with medium resolution, and the eighth and ninth feature maps with low resolution; The fourth and fifth feature maps are input into the secondary attention enhancement module for processing, and the processing results are input into the first detection head module; The sixth and seventh feature maps are directly input into the second detection head module; The eighth and ninth feature maps are directly input into the third detection head module. The first to third detection head modules respectively perform bounding box positioning of the mine car and the robotic arm based on the input feature maps, and input the positioning results into the result output module. The result output module obtains the positioning map of the dump truck and the robotic arm as the output of the model based on the input data. The mining cart and robotic arm recognition model is specifically based on YOLOv8s, replacing the backbone network with ShuffleNetV2. A metal anti-reflective attention module is embedded after each of the three feature layers output by ShuffleNetV2. The metal anti-reflective attention module includes a channel attention module, a spatial attention module, and a reflection suppression module connected in series. The metal anti-reflective attention module uses the channel attention module to enhance the metal texture feature channel, the spatial attention module focuses on the robotic arm contact area, and the reflection suppression module actively reduces high light interference. The feature fusion module uses an improved FPN+PAN structure. It applies a 1x1 convolution kernel to each feature layer to be fused, output by the three metal anti-reflective attention modules, unifies the channels to 256 dimensions, and fuses multi-scale metal features through the FPN and PAN dual-path transmission mechanism. The final output is the fourth to ninth feature maps of three different scales. Among them, the high-resolution fourth feature map is obtained by fusion through the FPN path, and the high-resolution fifth feature map is obtained by fusion through the PAN path. The sixth feature map of medium resolution is obtained by fusion of FPN pathway; the seventh feature map of medium resolution is obtained by fusion of PAN pathway; the eighth feature map of low resolution is obtained by fusion of FPN pathway; the ninth feature map of low resolution is obtained by fusion of PAN pathway; secondary attention enhancement is applied to the high-resolution feature map before the detection head, and finally the precise positioning of the dump truck is achieved through the three-scale detection head, and the final result is output through the result output module.

4. The automatic unloading method of a mine dump truck based on visual positioning according to claim 3 is characterized in that: The channel attention module is specifically: given a feature map ,Generate channel descriptors through global average pooling: ,in is the cth channel of the input channel, with a dimension of ; is the spatial resolution of the feature map, corresponding to the feature size of the neural network output; is the spatial coordinate index, traversing the entire two-dimensional feature map plane; is the global feature descriptor of the c-th channel, obtained by spatial dimension compression; the weights are calculated using a two-layer MLP: ,in is the Sigmoid function; is the ReLU activation function; The first layer fully connected weight matrix has the dimension , is the compression ratio, is the input dimension, is the output dimension; is the second layer fully connected weight matrix ; is the final generated channel attention weight vector; by learning the contribution of each channel, the expression of key feature channels such as the minecart metal texture is enhanced; the final output is ,in .

5. The automatic unloading method of a mine dump truck based on visual positioning according to claim 3 is characterized in that: The spatial attention module is specifically: the input is ,against Apply max pooling and average pooling along the channel axis to generate , fused through convolutional layers: ,in is a 7×7 convolution kernel; It is a feature map generated by the maximum pooling operation of the channel dimension, with a dimension of ; is the feature map generated by average pooling of the channel dimension, with dimension ; [;] is a tensor concatenation operation, merging two feature maps along the channel dimension; the output of the spatial attention module is .

6. The automatic unloading method of a mine dump truck based on visual positioning according to claim 3 is characterized in that: The reflection suppression module is specifically as follows: given an input feature map and ,in is the number of channels, and are the height and width of the feature map respectively. First, the statistical characteristics of the feature map are calculated. The highlight mask is generated by analyzing the statistical characteristics of the feature map. The specific steps include: Calculate the channel average feature map using the following formula: ,in is the number of channels of the neural network feature map, represents the average feature intensity at each spatial location, For the Feature maps of channels; Calculate the channel contrast feature map using the following formula: ,in Represents the feature variance at each spatial location; The mean map and variance map are concatenated along the channel dimension and fused through the convolution layer, using the following formula: ,in: Represents a tensor concatenation operation, generating dimensions of Feature map of for Convolution kernel, output single channel feature map , used to encode high-light responses; Apply the Sigmoid function to map the highlight response to the range [0, 1] to generate a highlight mask, which is calculated using the following formula: in is the highlight mask, a value close to 1 indicates a highlight area, and a value close to 0 indicates a non-highlight area; Sigmoid function Ensure the mask is smooth and differentiable; Apply the highlight mask to the original feature map to reduce the feature response in the highlight area, calculated using the following formula: in is the feature map output by the spatial attention module, represents element-wise multiplication, To generate suppression weights; the weight of the highlight area is close to 0, and the feature is suppressed; the weight of the non-highlight area is close to 1, and the feature is retained.

7. According to the method for automatic unloading of a mine dump truck based on visual positioning according to claim 2, the mine truck safety pin recognition model in step S4 comprises the following steps: Based on the mechanical structure of the safety pin, the region of interest is preset to guide the algorithm to focus on the effective area with rich and stable texture, reducing background interference; In the preset ROI, the FAST algorithm is used to detect candidate corner points, and the feature points with the highest response value are screened through non-maximum suppression to retain the corner points with strong geometric stability; A Gaussian difference pyramid is established to handle the detection requirements of safety pins of different sizes. Each layer of the pyramid is generated by a downsampling rate of 1.2 to achieve multi-scale feature coverage. According to the distance between the camera on the second robotic arm and the target depth And dynamic calculation of safety pin nominal size: ;in is the camera sensor pixel size, is the minimum recognizable physical size, is the safety factor; Calculate the image moment within a 15px radius neighborhood of the feature point and determine the main direction angle Direction correction range is ±15°, to adapt to the deflection angle of the safety pin during assembly; is the main direction angle of the feature point; is the image moment, representing the center of gravity of the grayscale distribution of the pixel neighborhood; for Gray value at ; , and is the order of the moment; the feature description is carried out in sequence using the steps of domain selection, Gaussian smoothing, and binary descriptor construction; The neighborhood selection covers the entire safety pin and matches the physical size of the safety pin; the Gaussian smoothing suppresses surface reflection noise of the metal; the binary descriptor is set to adopt a 256-bit Steered BRIEF mode to support rotation errors; The binary descriptor is constructed using the following formula: ,in is the pixel coordinate pair after rotation; is a binary index; is the grayscale query function; is the rotation matrix; Apply DBSCAN clustering algorithm to ORB feature points extracted from the whole image and adapt clustering parameters; The key parameters are dynamically set; the key parameters include the field radius and minimum points The dynamic setting is specifically: based on the camera internal parameters ( , , main point , ) and current distance , calculate the minimum physical diameter of the safety pin The minimum pixel size corresponding to the current viewing angle set up , is the empirical coefficient; set Based on the statistical lower bound of the effective texture area of ​​the safety pin; Verify the geometric characteristics of the candidate clusters obtained by DBSCAN to ensure that they conform to the mechanical structure of the safety pin; Based on the fact that the safety pin is a slender cylinder with a typical aspect ratio of ≥4:1, and the mechanical structural constraints of the safety pin, the covariance matrix of the points within the feature cluster is calculated and principal component analysis is performed. Clusters with an aspect ratio greater than 3:1 and a second principal component variance less than 15% are retained. The cluster that passes the above clustering and verification is identified as the target safety pin; if there are multiple candidate clusters, the cluster with the most expected size and the highest confidence is selected; After selecting the target feature cluster, extract its convex hull center point as the target position reference and calculate the depth distance to the target using the following formula: , the moving goal of the robot arm 2 is to minimize value; where Determined through calibration experiments, is the scaling factor, is the offset, is the distance to the target depth.

8. The method for automatically unloading a mine dump truck based on visual positioning according to claim 2, wherein step S5 specifically comprises the following steps: S51. After the mine car stops, the top camera captures the image of the relative position of the mine car and the robotic arm, determines the relative position of the robotic arm and the mine car as a whole, and controls the slide rail to move the robotic arm so that the horizontal position of the second robotic arm and the safety pin are aligned through the mine car and robotic arm recognition model. S52. Arm 1 extends, causing the electromagnet at its top to contact the mine cart. When powered on, the electromagnet adheres to the cart, maintaining the relative position of the cart and the arm. A camera on Arm 2 captures an image of the safety pin to determine whether the vertical alignment of Arm 2 and the safety pin is established. An object detection algorithm is used to locate the safety pin, and based on the result, Arm 2 is controlled to align with the safety pin. S53. After both cameras determine that they have reached the desired position, the electromagnet is powered on and the latch is searched for by shaking the latch using a random perturbation search algorithm. A determination is made as to whether the latch is properly attached to the latch. If not, the electromagnet is turned off and the robotic arm is fine-tuned by reusing the random perturbation search algorithm. If the robotic arm has moved to the proper position, the second robotic arm extends, causing the electromagnet at the top to contact the locking mechanism. After the locking mechanism is engaged, the second robotic arm rotates to unlock the locking mechanism. After unlocking the locking mechanism, the second robotic arm maintains the electromagnet's attachment to the latch and operates synchronously with the first robotic arm. S54. Arm 1 continues to extend, pushing the mine cart, causing it to flip and complete unloading. After unloading is complete, arm 1 retracts, returning the mine cart to its normal position. Arm 2 then rotates, closing the locking mechanism. Finally, power is turned off, causing the electromagnet to lose its magnetic force, and both arms return to their initial positions. Step S51 specifically includes: using the top camera and the mine car relative position recognition model to perform target detection on the mine car, obtaining the coordinates of the mine car bounding box, and calculating the distance between the mine car position and the standard dumping position using the following formula: ;in For the target location ; is the top of the minecart's bounding box, The bottom edge of the minecart's bounding box. For minecarts in The distance between the axis and the standard unloading position; by moving the robot arm 1 and the robot arm 2 synchronously along the track, the robot arm 1 Axis coordinates and minecart Axis coordinates are aligned, and a robotic arm is used to absorb the dump truck through an electromagnet and then minimized. , place the minecart in Move to the designated target area on the coordinates; The random perturbation search algorithm uses a sine wave superimposed on random noise to generate a shaking trajectory: Where A is the amplitude, which is less than the safety pin hole tolerance; is the frequency, avoiding the resonant frequency of the robotic arm; The zero-mean Gaussian noise enhances robustness; when the contact force at the end of the robot arm Exceeding the threshold Stop shaking and judge the position of the robotic arm and the pin.

9. The automatic unloading method of a mine dump truck based on visual positioning according to claim 8, characterized in that: Step S5 also includes a protection mechanism that uses multiple sensors to trigger power failure, specifically: real-time detection of the motor operating current , when exceeding the rated value Protection is triggered when the value reaches 120%; Detecting abnormal resistance torque through joint torque sensors The threshold is set to 150% of the rated torque, forming a double check with the current monitoring to avoid false triggering.

10. A system for realizing the automatic unloading method of a mine dump truck based on visual positioning according to any one of claims 1 to 9, characterized in that: It includes image acquisition module, image annotation module, model construction and training module, safety pin recognition model construction module, and mine dump truck automatic unloading module; The image acquisition module acquires historical overhead images of the dump truck and the robotic arm, and uploads the data to the image annotation module; The image annotation module annotates the historical overhead images based on the received data to obtain a training dataset and uploads the data to the model building and training module; The model building and training module builds an initial mine car and robotic arm recognition model based on YOLOv8s according to the received data, and uses the training data set for training to obtain the mine car and robotic arm recognition model, and uploads the data to the mine dump truck automatic unloading module; The safety pin recognition model building module builds a mine car safety pin recognition model based on the traditional machine vision ORB feature point detection algorithm and uploads the data to the mine dump truck automatic unloading module; The automatic unloading module of the mine dump truck establishes a three-dimensional vision positioning system based on the received data, using the recognition model of the mine truck and the robotic arm, as well as the recognition model of the mine truck safety pin, to realize the automatic unloading of the mine dump truck.

Citation Information

Patent Citations

  • Marine target identification and positioning method based on stereoscopic vision

    CN117351345A

  • LNG unloading arm target identification method based on improved YOLO-V5s model

    CN117649589A

  • Railway bulk cargo discharge hopper identification and positioning method based on machine vision

    CN120220066A

  • Control method and system for autonomous unloading of carry-scraper

    CN120428708A

  • General target detection method for adaptive attention guidance mechanism

    WO2021139069A1