A grasping control method for a coal mine underground manipulator based on visual positioning

By applying the MAE-GRU neural network model underground in coal mines, combined with the RGBD binocular camera and acceleration sensor, the problems of precise positioning and accumulation of robotic arm errors in complex environments are solved, high-precision positioning of the drill rod and long-term and efficient operation of the robotic arm are achieved, and the safety and efficiency of drilling pressure relief operations in coal mines are improved.

CN115674192BActive Publication Date: 2025-05-27CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211225768.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-05-27
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

Under coal mines, traditional visual positioning technology is difficult to achieve accurate positioning of drill rods in complex environments such as fog, dust, etc., and the robotic arm will accumulate positioning errors after long-term operation, affecting the safety and efficiency of drilling pressure relief operations.

Method used

The MAE-GRU neural network model is adopted, combined with the RGBD binocular camera and acceleration sensor to achieve accurate positioning of the drill rod and real-time correction of the position of the robot arm. The MAE network enhances image characteristics through a random mask mechanism, while the GRU network realizes self-calibration of the robotic arm through error accumulation and threshold adjustment.

Benefits of technology

In a complex downhole environment, high-precision positioning of the drill pipe and long-term and efficient operation of the robot arm are achieved, improving the safety and efficiency of drilling pressure relief operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115674192B_ABST
    Figure CN115674192B_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical arm grasping control method for underground coal mines based on visual positioning, comprising a drill rod loading and unloading mechanical arm, an acceleration sensor, a drill rod library, an RGBD binocular camera and a data processing center; the invention imports the collected visual image into a MAE-based neural network, and through the MAE automatic masking mechanism, the entire neural network pays more attention to the overall information, thereby achieving the effect of image enhancement to realize the accurate recognition and positioning of the drill rod in a complex background including floating fog and the like; the mechanical arm grasping information output by the acceleration sensor on the mechanical arm and the position information output by the MAE neural network are input into a GRU-based neural network, and the displacement and posture are adaptively adjusted according to the situation of each mechanical arm grasping, so that the mechanical arm can always maintain good accuracy in long-term work and improve work efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a robotic arm grasping and positioning system, specifically a method for controlling the grasping of a robotic arm underground in a coal mine based on visual positioning, belonging to the technical field of coal mining. Background Technique

[0002] Coal is one of the main energy sources used by humans. China has a large coal reserve. In China's energy consumption structure, coal consumption still occupies a dominant position. Rock burst is a special form of mine pressure manifestation. During the mining process, many hazards will occur due to rock burst. Moreover, with the increase of the mining depth and intensity of coal mines in China, the occurrence frequency and damage intensity of rock bursts are also increasing continuously, seriously threatening the safety production of coal mines. Borehole pressure relief is an effective method for preventing rock bursts. However, the current pressure relief work requires human participation, with high labor intensity and high risk. Achieving unmanned operation of borehole pressure relief has increasingly become an important measure to deal with rock burst disasters. As a key process in borehole pressure relief operations, the unmanned operation of drill pipe handling has increasingly become an important measure to deal with rock burst disasters.

[0003] Previously, traditional drill pipe handling usually chose to place the drill pipe in a fixed position and required manual operation of the drill pipe handling process, resulting in low drill pipe handling efficiency and certain safety hazards. The accurate positioning of the drill pipe and the movement accuracy of the robotic arm will directly determine the efficiency of automatic drill pipe handling.

[0004] Traditional visual positioning has been applied in many other industrial fields. However, its application scenarios are often on industrial assembly lines, with a stable working environment and a simple working background, making it easy to achieve good industrial effects. However, due to special working environments such as fog and floating dust underground, it is very difficult for the visual system to achieve ideal accuracy underground. At the same time, based on the fact that positioning errors will continuously accumulate during the movement of the robotic arm, it may affect the automatic drill pipe handling process after long-term operation. The continuous superposition of these two errors will directly affect the safety and working efficiency of borehole pressure relief operations. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for controlling the grasping of a robotic arm underground in a coal mine based on visual positioning to solve at least one of the above technical problems. The present invention proposes a MAE-GRU neural network model, which uses the implementation of the MAE network model to achieve precise positioning of the drill pipe in special working environments such as fog and floating dust underground. At the same time, the position information output by the MAE network model and the position information output by the acceleration sensor are imported into the GRU network model, and the pose of the robotic arm is corrected in real time by judging the size of the error value and the pre-set threshold in real time, so as to achieve long-term high-precision operation of the drill pipe handling process.

[0006] The present invention realizes the above object through the following technical solutions: A method for grasping and controlling a manipulator in a coal mine based on visual positioning, including

[0007] A drill pipe loading and unloading manipulator, which is a six-axis manipulator, and the rear end of the drill pipe loading and unloading manipulator is connected to an impact prevention drilling robot through a rotating base, and a gripper for loading and unloading drill pipes is provided at the front end of the drill pipe loading and unloading manipulator;

[0008] An acceleration sensor, which calculates the actual position of the manipulator by obtaining the attitude parameters of the roll angle, pitch angle, angular velocity, and acceleration of the manipulator, and provides data for the error calibration of the manipulator. A number of acceleration sensors are provided and are respectively installed on both sides of each section of the drill pipe loading and unloading manipulator. The two acceleration sensors on the manipulator of each section of the drill pipe loading and unloading manipulator cancel out the influence of errors on each other to improve the signal-to-noise ratio of the output signal;

[0009] A drill pipe library, which is fixed on the impact prevention drilling robot for storing drill pipes, and the drill pipe library is located on the front side of the drill pipe loading and unloading manipulator;

[0010] An RGBD binocular camera, which is used to collect real-time video data and is set independently of the drill pipe loading and unloading manipulator. The RGBD binocular camera is installed above the drill pipe library, and the best viewing angle of the RGBD binocular camera completely covers the drill pipe library;

[0011] A data processing center, which realizes the pipeline processing and parallel computing of video stream data based on the arithmetic unit of FPGA, and obtains the spatial position information of the drill pipe by receiving the real-time video data collected by the RGBD binocular camera. The data processing center includes the pre-training of the MAE and GRU neural networks and the calculation during actual work. The MAE neural network of the data processing center specifically includes an encoder, a decoder, and a positioning algorithm. The GRU neural network of the data processing center specifically includes an input, a reset gate, an update gate, a candidate memory, and an output;

[0012] The specific adjustment method of the data processing center includes the following steps:

[0013] Step S1: Build a neural network based on MAE-GRU: Couple the Transformer and GRU modules together;

[0014] Step S2: Pre-train the neural network based on MAE: Collect the downhole video data of the RGBD binocular camera at a certain moment multiple times. The default number of video frames is 30, and the resolution is 1080×720. Make the collected videos into a sample set and import it into the neural network based on MAE to complete the model pre-training;

[0015] Step S3: Transmit the collected RGBD video stream to the data processing center, output the spatial position information of the drill pipe for visual positioning through the neural network based on MAE and make a sample set, and obtain a GRU neural network by inputting the sample data;

[0016] Step S4: Import the sample set containing the spatial position information of the drill pipe into the built GRU neural network and set the error threshold T;

[0017] Step S5: At time t, input the real-time collected RGBD into the neural network based on MAE, complete image enhancement and output the three-dimensional coordinate information S of the drill pipe 1 , and the robotic arm performs a grasping operation according to the three-dimensional coordinate information S of the drill pipe 1 ;

[0018] Step S6: Inversely calculate the three-dimensional coordinate information S of the robotic arm when it completes the grasping task according to the digital signal fed back by the acceleration sensor 2 ;

[0019] Step S7: Subtract the three-dimensional coordinate information S of the drill pipe for visual positioning 1 from the three-dimensional coordinate information S of the robotic arm during grasping 2 to obtain the error value T at this time t , and at the same time set the error output X of the GRU network at this time t as the error value T t , first compare the error output X t with the set error threshold T. If the error value is less than the set error threshold, continue to collect the error data T at the next grasping moment t + 1 t+1 , and accumulate the error between T t+1 and the previously generated error to obtain the cumulative error output X t+1 , and compare the obtained error output X t+1 with the set error threshold T;

[0020] Step S8: Repeat Step S4 until the generated cumulative error output X t+n is greater than the set error threshold T, then end the loop, adjust the robotic arm control system according to the azimuth relationship between the error value and the actual value to offset the influence of the error. The error generated by the robotic arm during operation will continue to accumulate, and by setting the error threshold, the error is always kept within the range that does not affect the drill pipe grasping operation.

[0021] As a further solution of the present invention: The encoder structure of the MAE neural network constituting the data processing center specifically includes:

[0022] ① Convert the downhole video data collected by the RGBD binocular camera into pictures and make a sample set. Divide the input pictures into blocks of 60*60, and set the masking rate to 75%;

[0023] ② After adding the corresponding position information to each image block to form a vector, use a linear transformation matrix to perform a linear mapping on the vector;

[0024] ③ After randomly shuffling the vector sequence, remove the last element. At this time, take the first 25% of the elements of the vector. The first 25% of the elements retain both their corresponding position information and feature information, and the last 75% of the elements only retain the corresponding position information and then perform a masking operation;

[0025] ④ After obtaining the vector information, pass it into the Transformer Encoder for feature extraction.

[0026] As a further solution of the present invention: The decoder structure of the MAE neural network constituting the data processing center specifically includes:

[0027] ① Through linear projection, convert the encoder output dimension to the decoder input dimension;

[0028] ② Through Reshuffle, restore the previously randomly shuffled picture vectors to their original order and update the vectors of their position information;

[0029] ③ After obtaining the vector information, pass it into the Transformer decoder for feature extraction. At this time, the amount of information that the decoder needs to process is one-fourth of that of the encoder;

[0030] ④ Project the vectors output by the Transformer blocks into the pixel space through MLP to complete the image classification task, and calculate the loss function between the output data and the data of the image normal distribution;

[0031] ⑤ Identify the drill pipe in the image according to the classification information, determine the X-Y axis information of the target object according to the position of the drill pipe library in the image, and at the same time superimpose the depth information of the depth camera to obtain the Z-axis information, and obtain the true three-dimensional coordinates S of the drill pipe to be grasped 1 。

[0032] As a further solution of the present invention: The input of the GRU neural network constituting the data processing center specifically includes:

[0033] Take the difference between the three-dimensional coordinate information S of the drill pipe for visual positioning 1 and the three-dimensional coordinate information S during the grasping of the robotic arm 2 to obtain the error value T t+1 as the input X at the current moment t+1 。

[0034] As a further solution of the present invention: The reset gate constituting the GRU neural network of the data processing center specifically includes:

[0035] Determine the error output X at the previous moment t And the error value T at the current moment t+1 Feed into the activation function for output

[0036] The gating signal value within the range of [0, 1]. The larger the signal value, the greater the weight of the error output X at the previous moment t .

[0037] Wherein:

[0038] r t+1 = σ(W r · [X t , T t+1 + b r )

[0039] r t+1 Represents the gating signal value at the current moment, σ represents the sigmoid function, W r Represents the weight of the gating signal value, X t Represents the error output at the previous moment, T t+1 Represents the error value at the current moment, b r Represents the bias of the gating signal value.

[0040] As a further solution of the present invention: The new gate constituting the GRU neural network of the data processing center specifically includes:

[0041] Input information X t , T t+1 Pass through the activation function to obtain the gating threshold, thereby dividing the threshold into z t+1 And 1 - z t+1 Two parts to achieve the input of new information and the abandonment of historical information.

[0042] Wherein:

[0043] z t+1 = σ(W r · [X t , T t+1 + b z )

[0044] z t+1 Represents the gating threshold at the current moment, σ represents the sigmoid function, W z Represents the weight of the gating threshold, X t Represents the error output at the previous moment, T t+1 Represents the error value at the current moment, b z Represents the bias of the gating threshold.

[0045] As a further solution of the present invention: The candidate memories that make up the GRU neural network of the data processing center specifically include:

[0046] The candidate memory consists of two parts. One part is the past error output X determined by the reset gate value signal t , and the other part is the current input error value T t+1 . When the gate value r t+1 = 0, it means that the past information is completely abandoned, then it only contains the current information.

[0047] Where:

[0048]

[0049] represents the candidate memory at the current moment, tanh represents the hyperbolic tangent function, represents the weight of the candidate memory value, z t+1 represents the gate threshold value at the current moment, X t represents the error output at the previous moment, T t+1 represents the error value at the current moment, b z represents the bias of the gate signal value.

[0050] As a further solution of the present invention: The output that makes up the GRU neural network of the data processing center specifically includes:

[0051] The final output is controlled by the update gate. One part determines the degree of forgetting information from the error output X t at the previous moment, and the other part determines the degree to which the current candidate memory is added to the output. The control coefficients of the two are respectively 1 - z t+1 , and z t+1 controls, which makes the two in a balanced conversion state.

[0052] Where:

[0053]

[0054] X t+1 represents the error output at the current moment, z t+1 represents the gate threshold value at the current moment, X t represents the error output at the previous moment, represents the candidate memory at the current moment.

[0055] The beneficial effects of the present invention are as follows: The collected visual images are imported into the MAE neural network. Through the MAE automatic masking mechanism, the entire neural network pays more attention to the overall information, thereby achieving the effect of image enhancement to accurately identify and position the drill pipe in complex backgrounds such as floating dust and fog. By inputting the manipulator grasping information output by the acceleration sensor on the manipulator and the position information output by the MAE neural network into the GRU-based neural network, the GRU network model weights and accumulates the current manipulator pose error input and the error output at a certain previous moment to obtain the error output at the current moment. This error value is compared with a preset error threshold. When it is greater than the threshold, the operating parameters of the manipulator are adjusted to achieve real-time error correction. When it is less than the threshold, the current operation continues, and the displacement and posture are adaptively adjusted according to the situation of each manipulator grasp, so that the manipulator can always maintain good accuracy during long-term operation and improve work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic diagram of the drill pipe loading and unloading structure of the present invention;

[0057] Figure 2 It is a schematic diagram of the MAE-based neural network of the present invention;

[0058] Figure 3 It is a schematic diagram of the GRU-based neural network of the present invention;

[0059] Figure 4 It is a schematic diagram of the automatic adjustment process of the drill pipe loading and unloading manipulator of the present invention.

[0060] In the figure: 1. Drill pipe loading and unloading manipulator, 2. Acceleration sensor, 3. RGBD binocular camera, 4. Drill pipe library. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0062] Embodiment 1

[0063] As Figures 1 to 3 shown, a method for controlling the grasping of a manipulator in a coal mine underground based on visual positioning includes a drill pipe loading and unloading manipulator 1, which is a six-axis manipulator, and the rear end of the drill pipe loading and unloading manipulator 1 is connected to an anti-collision drilling robot through a rotating base. The front end of the drill pipe loading and unloading manipulator 1 is provided with a gripper for loading and unloading drill pipes;

[0064] An acceleration sensor 2, which calculates the actual position of the robotic arm by obtaining the attitude parameters of the roll angle, pitch angle, angular velocity, and acceleration of the robotic arm, and provides data for the error calibration of the robotic arm. A plurality of acceleration sensors 2 are provided and are respectively installed on both sides of each section of the drill pipe handling robotic arm 1. The two acceleration sensors 2 on the robotic arm of each section of the drill pipe handling robotic arm 1 cancel out the influence of errors on each other to improve the signal-to-noise ratio of the output signal;

[0065] A drill pipe library 4, which is fixed on the anti-impulse drilling robot for storing drill pipes, and the drill pipe library 4 is located on the front side of the drill pipe handling robotic arm 1;

[0066] An RGBD binocular camera 3, which is used to collect real-time video data and is arranged independently of the drill pipe handling robotic arm 1. The RGBD binocular camera 3 is installed above the drill pipe library 4, and the best viewing angle of the RGBD binocular camera 3 completely covers the drill pipe library 4;

[0067] A data processing center, which realizes the pipeline processing and parallel computing of video stream data based on the arithmetic unit of FPGA, and obtains the drill pipe spatial position information by receiving the real-time video data collected by the RGBD binocular camera 3. The data processing center includes the pre-training of the MAE and GRU neural networks and the calculation during actual work. The MAE neural network of the data processing center specifically includes an encoder, a decoder, and a positioning algorithm. The GRU neural network of the data processing center specifically includes an input, a reset gate, an update gate, a candidate memory, and an output;

[0068] The specific adjustment method of the data processing center includes the following steps:

[0069] Step S1: Build a neural network based on MAE-GRU: Couple the two main modules of Transformer and GRU. The MAE module is used because it removes redundant information in the image through a random masking mechanism, enabling the algorithm to pay more attention to global information to improve the image recognition and classification ability, while the GRU layer can directly learn multiple parallel sequences of input data to improve the self-calibration accuracy of the robotic arm;

[0070] Step S2: Pre-train the neural network based on MAE: Collect the downhole video data of the RGBD binocular camera 3 at a certain moment multiple times. The default number of video frames is 30, and the resolution is 1080×720. Make the collected videos into a sample set and import it into the neural network based on MAE to complete the model pre-training;

[0071] Step S3: Transmit the collected RGBD video stream to the data processing center, output the spatial position information of the drill pipe for visual positioning through the neural network based on MAE and create a sample set, and obtain a trained GRU neural network by inputting a large amount of sample data;

[0072] Step S4: Import the sample set containing the spatial position information of the drill pipe into the established GRU neural network and set the error threshold T;

[0073] Step S5: At time t, input the real-time collected RGBD into the neural network based on MAE, complete image enhancement and output the three-dimensional coordinate information S of the drill pipe 1 , and the robotic arm performs a grasping operation according to the three-dimensional coordinate information S of the drill pipe 1 ;

[0074] Step S6: According to the digital signal fed back by the acceleration sensor 2, inversely calculate the three-dimensional coordinate information S when the robotic arm completes the grasping task 2 ;

[0075] Step S7: Take the difference between the three-dimensional coordinate information S of the drill pipe obtained by visual positioning 1 and the three-dimensional coordinate information S when the robotic arm grasps, obtain the error value T at this time 2 , and at the same time set the error output X of the GRU network at this time t as the error value T t , first compare the error output X t with the set error threshold T. If the error value is less than the set error threshold, continue to collect the error data T at the next grasping moment t + 1 t , accumulate the error between T t+1 and the previously generated error to obtain the cumulative error output X t+1 , and compare the obtained error output X t+1 with the set error threshold T; t+1

[0076] Step S8: Repeat Step S4 until the generated cumulative error output X t+n is greater than the set error threshold T, then end the loop, adjust the robotic arm control system according to the azimuth relationship between the error value and the actual value to offset the influence of the error. The error generated during the operation of the robotic arm will continuously accumulate, and by setting the error threshold, the error is always kept within the range that does not affect the drill pipe grasping operation.

[0077] In the embodiment of the present invention, the encoder structure of the MAE neural network constituting the data processing center specifically includes:

[0078] ① Convert the downhole video data collected by the RGBD binocular camera 2 into pictures and create a sample set,

[0079] Divide the input image into blocks of 60*60 and set the masking rate to 75%;

[0080] ② After adding the corresponding position information to each image block to form a vector, use a linear transformation matrix to perform a linear mapping on the vector;

[0081] ③ After randomly shuffling the vector sequence, remove the last element. At this time, take the first 25% of the elements of the vector. The first 25% of the elements retain both their corresponding position information and feature information, and the last 75% of the elements only retain the corresponding position information and then perform a masking operation;

[0082] ④ After obtaining the vector information, pass it into the Transformer Encoder for feature extraction.

[0083] In the embodiment of the present invention, the decoder structure constituting the data processing center MAE neural network specifically includes:

[0084] ① Through linear projection, convert the encoder output dimension to the decoder input dimension;

[0085] ② Restore the previously randomly shuffled picture vectors to their original order through Reshuffle and update the vectors of their position information;

[0086] ③ After obtaining the vector information, pass it into the Transformer decoder for feature extraction. At this time, the amount of information that the decoder needs to process is one-fourth of that of the encoder;

[0087] ④ Project the vectors output by the Transformer blocks into the pixel space through MLP to complete the image classification task, and calculate the loss function between the output data and the data of the image normal distribution;

[0088] ⑤ Identify the drill pipe in the image according to the classification information, determine the X-Y axis information of the target object according to the position of the drill pipe library in the image, and at the same time superimpose the depth information of the depth camera to obtain the Z axis information, and obtain the true three-dimensional coordinates S of the drill pipe to be grasped 1 。

[0089] In the embodiment of the present invention, the input constituting the data processing center GRU neural network specifically includes:

[0090] Take the difference between the three-dimensional coordinate information S of the drill pipe for visual positioning 1 and the three-dimensional coordinate information S during the grasping of the robotic arm 2 to obtain the error value T t+1 as the input X at the current moment t+1 。

[0091] In the embodiment of the present invention, the reset gate constituting the GRU neural network of the data processing center specifically includes:

[0092] Determine the error output X at the previous moment t And the error value T at the current moment t+1 Are sent into the activation function to output a gating signal value within the range of [0, 1]. The larger the signal value, the greater the weight of the error output X at the previous moment t Is.

[0093] Where:

[0094] r t+1 = σ(W r ·[X t , T t+1 + b r )

[0095] r t+1 Represents the gating signal value at the current moment, σ represents the sigmoid function, W r Represents the weight of the gating signal value, X t Represents the error output at the previous moment, T t+1 Represents the error value at the current moment, b r Represents the bias of the gating signal value.

[0096] In the embodiment of the present invention, the new gate constituting the GRU neural network of the data processing center specifically includes:

[0097] Input information X t , T t+1 Pass through the activation function to obtain a gating threshold, so as to divide the threshold into z t+1 And 1 - z t+1 Two parts to realize the input of new information and the abandonment of historical information.

[0098] Where:

[0099] z t+1 = σ(W r ·[X t , T t+1 + b z )

[0100] z t+1 Represents the gating threshold at the current moment, σ represents the sigmoid function, W z Represents the weight of the gating threshold, X t Represents the error output at the previous moment, T t+1 Represents the error value at the current moment, b z Represents the bias of the gating threshold.

[0101] In the embodiments of the present invention, the candidate memories that make up the GRU neural network of the data processing center specifically include:

[0102] The candidate memory consists of two parts. One part is the past error output X determined by the reset gate value signal t , and the other part is the current input error value T t+1 . When the gate value r t+1 = 0, it means that the past information is completely abandoned, then it only contains the current information.

[0103] Where:

[0104]

[0105] represents the candidate memory at the current moment, tanh represents the hyperbolic tangent function, represents the weight of the candidate memory value, z t+1 represents the gate threshold value at the current moment, X t represents the error output at the previous moment, T t+1 represents the error value at the current moment, b z represents the bias of the gate signal value.

[0106] In the embodiments of the present invention, the output that makes up the GRU neural network of the data processing center specifically includes:

[0107] The final output is controlled by the update gate. One part determines the degree of forgetting information from the error output X at the previous moment t , and the other part determines the degree to which the current candidate memory is added to the output. The control coefficients of the two are respectively controlled by 1 - z t+1 , and z t+1 , which makes the two in a balanced conversion state.

[0108] Where:

[0109]

[0110] X t+1 represents the error output at the current moment, z t+1 represents the gate threshold value at the current moment, X t represents the error output at the previous moment, represents the candidate memory at the current moment.

[0111] Embodiment 2

[0112] As Figure 4 shown, a method for controlling the grasping of a coal mine underground manipulator based on visual positioning, and its specific adjustment method includes the following steps:

[0113] Step S1: Build a MAE-GRU neural network: couple the two main modules, Transformer and GRU, together. The MAE module is used because it removes redundant information in the image through a random mask mechanism, allowing the algorithm to pay more attention to global information to improve the image recognition and classification capabilities, while the GRU layer can directly learn multiple parallel sequences of input data to improve the self-calibration accuracy of the robotic arm;

[0114] Step S2: MAE-based neural network pre-training: The downhole video data of the RGBD binocular camera 2 at a certain moment is collected multiple times. The video frame number is 30 frames by default, and the resolution is 1080×720. The collected videos are made into a sample set and imported into the MAE-based neural network to complete the model pre-training;

[0115] Step S3: The collected RGBD video stream is transmitted to the data processing center, and the spatial position information of the drill rod located by visual positioning is output through the MAE-based neural network and a sample set is prepared. A trained GRU neural network is obtained by inputting a large amount of sample data;

[0116] Step S4: import the sample set containing the spatial position information of the drill rod into the constructed GRU neural network, and set the error threshold T;

[0117] Step S5: At time t, the real-time collected RGBD is input into the MAE-based neural network to complete image enhancement and output the three-dimensional coordinate information S of the drill rod. 1 , the robot arm is based on the three-dimensional coordinate information S of the drill pipe 1 Number of crawl operations;

[0118] Step S6: Reversely calculate the three-dimensional coordinate information S when the robot arm completes the grasping task based on the digital signal fed back by the acceleration sensor 2 2 ;

[0119] Step S7: The three-dimensional coordinate information S of the visually positioned drill rod 1 and the three-dimensional coordinate information S when the robot grasps 2 Subtract and get the error value T at this time t , and at the same time output the error of the GRU network at this time X t Set to error value T t , first output the error X t Compare it with the set error threshold T. If the error value is less than the set error threshold, continue to collect the error data T at the next capture time t+1 t+1 , T t+1 The error is accumulated with the previous error to obtain the cumulative error output X t+1 , and output the obtained error to X t+1 Compare with the set error threshold T;

[0120] Step S8: Repeat Step S4 until the generated cumulative error output X t+n is greater than the set error threshold T, then end the loop, and adjust the manipulator control system according to the azimuth relationship between the error value and the actual value to offset the influence of the error. The error generated during the operation of the manipulator will accumulate continuously. By setting the error threshold, the error is always kept within the range that does not affect the drill pipe grasping operation.

[0121] In the embodiment of the present invention, the collected data is imported into the MAE-GRU network architecture for training to obtain a prediction model. This can make the trained model more reliable. By using the three-dimensional coordinate information S of the drill pipe obtained by visual positioning 1 and the three-dimensional coordinate information S when the manipulator grasps 2 are compared and real-time error compensation is performed, so as to realize the automatic adjustment of the displacement and posture of the drill pipe grasping manipulator, improve the drill pipe loading and unloading efficiency, and improve the unmanned level of borehole pressure relief.

[0122] Working principle: Based on the MAE-GRU network architecture, while obtaining more accurate visual positioning, through the manipulator automatic calibration system based on the gated recurrent unit, the positioning error of the manipulator is calibrated in time, so that the positioning cumulative error of the manipulator is always kept within a reasonable range that does not affect normal operation.

[0123] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0124] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A grasping control method for a robotic arm in a coal mine based on visual positioning, characterized in that: It includes: A drill pipe loading and unloading robotic arm (1), which is a six-axis robotic arm, and the rear end of the drill pipe loading and unloading robotic arm (1) is connected to an outburst prevention drilling robot through a rotating base. A gripper for loading and unloading drill pipes is provided at the front end of the drill pipe loading and unloading robotic arm (1); Acceleration sensors (2), which calculate the actual position of the robotic arm by obtaining the attitude parameters of the robotic arm such as roll angle, pitch angle, angular velocity, and acceleration, and provide data for the error calibration of the robotic arm. A number of acceleration sensors (2) are provided and are respectively installed on both sides of each section of the drill pipe loading and unloading robotic arm (1). The two acceleration sensors (2) on each section of the drill pipe loading and unloading robotic arm (1) cancel out the influence of errors to improve the signal-to-noise ratio of the output signal; A drill pipe library (4), which is fixed on the outburst prevention drilling robot for storing drill pipes, and the drill pipe library (4) is located on the front side of the drill pipe loading and unloading robotic arm (1); An RGBD binocular camera (3), which is used to collect real-time video data and is set independently of the drill pipe loading and unloading robotic arm (1). The RGBD binocular camera (3) is installed above the drill pipe library (4), and the best viewing angle of the RGBD binocular camera (3) completely covers the drill pipe library (4); A data processing center, which realizes the pipeline processing and parallel computing of video stream data based on the operation unit of FPGA, and obtains the spatial position information of the drill pipe by receiving the real-time video data collected by the RGBD binocular camera (3). The data processing center includes the pre-training and calculation during actual work of a neural network based on MAE and a GRU neural network. The neural network of MAE in the data processing center specifically includes an encoder, a decoder, and a positioning algorithm. The GRU neural network in the data processing center specifically includes an input, a reset gate, an update gate, a candidate memory, and an output; The method for adjusting the data processing center includes the following steps: Step S1: Build a neural network based on MAE-GRU: Couple the Transformer and GRU modules together; Step S2: Pre-train the neural network based on MAE: Collect the underground video data of the RGBD binocular camera (3) at a certain moment multiple times. The default number of video frames is 30 frames, and the resolution is 1080×720. Make the collected videos into a sample set and import it into the neural network based on MAE to complete the model pre-training; Step S3: Transmit the collected RGBD video stream to the data processing center, output the spatial position information of the visually positioned drill pipe through the neural network based on MAE and make a sample set, and obtain a GRU neural network by inputting the sample data; Step S4: Import the sample set containing the spatial position information of the drill pipe into the built GRU neural network and set the error threshold T; Step S5: At time t, input the real-time collected RGBD video stream into the MAE neural network to complete image enhancement and output the three-dimensional coordinate information S of the drill pipe 1 , and the robotic arm performs a grasping operation according to the three-dimensional coordinate information S of the drill pipe 1 ; Step S6: Reverse-calculate the three-dimensional coordinate information S of the robotic arm when it completes the grasping task according to the digital signal fed back by the acceleration sensor (2). 2 ; Step S7: Take the difference between the three-dimensional coordinate information S of the drill pipe for visual positioning 1 and the three-dimensional coordinate information S 2 to obtain the error value T at this time t . At the same time, set the error output X of the GRU neural network at this time t as the error value T t . First, compare the error output X t with the set error threshold T. If the error output is less than the set error threshold, continue to collect the error data T at the next grasping moment t + 1 t+1 . Add T t+1 to the previous error output for error accumulation to obtain the cumulative error output X t+1 . Compare the obtained error output X t+1 with the set error threshold T; Step S8: Repeat step S4 until the generated cumulative error output X t+n is greater than the set error threshold T, then end, and adjust the manipulator control system according to the azimuth relationship between the error value and the actual value to offset the influence of the error. The error generated during the operation of the manipulator will accumulate continuously. By setting the error threshold, the error is always kept within the range that does not affect the drill pipe grasping operation.

2. According to the grasping control method for a robotic arm in a coal mine based on visual positioning described in claim 1, characterized in that: The encoder structure of the neural network of MAE that constitutes the data processing center specifically includes: ① Convert the downhole video data collected by the RGBD binocular camera (3) into pictures and make a sample set. Divide the input pictures into blocks of 60*60, and set the masking rate to 75%; ② Add the corresponding position information to each image block to form a vector, and perform a linear mapping on the vector using a linear transformation matrix; ③ After randomly shuffling the vector sequence, remove the last element. At this time, take the first 25% of the elements of the vector sequence. The first 25% of the elements retain both their corresponding position information and feature information, and the last 75% of the elements only retain the corresponding position information and then perform a masking operation; ④ After obtaining the vector information, pass it into the Transformer Encoder for feature extraction.

3. A method for controlling the grasping of a coal mine underground manipulator based on visual positioning according to claim 1, characterized in that: The decoder structure of the neural network of the MAE constituting the data processing center specifically includes: ① Through linear projection, convert the encoder output dimension into the decoder input dimension; ② Restore the previously randomly shuffled pictures to their original order through Reshuffle, and update the vector of the picture position information; ③ After obtaining the vector of the position information, pass it into the Transformer decoder for feature extraction. At this time, the amount of information that the decoder needs to process is one-fourth of that of the encoder; ④ Project the vector output by the Transformer blocks into the pixel space through the MLP to complete the image classification task, and calculate the loss function through the output data and the data of the image normal distribution; ⑤Identify the drill pipe in the image according to the classification information, determine the X-Y axis information of the target object based on the position of the drill pipe in the image, and obtain the Z-axis information by superimposing the depth information of the depth camera to get the three-dimensional coordinate information S of the drill pipe to be grasped 1 。 4. A method for controlling the grasping of a coal mine underground manipulator based on visual positioning according to claim 1, characterized in that: The reset gate of the GRU neural network constituting the data processing center specifically includes: Determine the error output X at the current moment t , and use the error output X at the current moment t and the error data T at the next moment t+1 as the input to the activation function, and output the gating signal value within the range of [0, 1]. The larger the signal value, the greater the weight of the error output X t at the current moment Wherein: r t+1 = σ(W r · [X t ,T t+1 + b r ) r t+1 represents the value of the gating signal at the next moment, σ represents the sigmoid function, W r represents the weight of the gating signal value, b r represents the bias of the gating signal value.

5. A method for controlling the grasping of a coal mine underground manipulator based on visual positioning according to claim 1, characterized in that: The update gate of the GRU neural network constituting the data processing center specifically includes: Input information X t , T t+1 Pass through the activation function to obtain the gating threshold, thereby dividing the gating threshold into z t+1 and 1 - z t+1 Two parts are used to implement the input of new information and the abandonment of historical information; Wherein: z t+1 = σ(W z · [X t ,T t+1 + b z ) z t+1 represents the next moment's gating threshold, σ represents the sigmoid function, W z represents the weight of the gating threshold, X t represents the error output at the current moment, T t+1 represents the error data at the next moment, b z represents the bias of the gating threshold.

6. A method for controlling the grasping of a coal mine underground manipulator based on visual positioning according to claim 1, characterized in that: The output of the GRU neural network constituting the data processing center specifically includes: The final output is controlled by the update gate. One part determines the error output X from the current moment. t The degree of forgetting information, and the other part determines the degree to which the current candidate memory is added to the output. The control coefficients of the two are respectively determined by 1 - z t+1 and z t+1 controlled, which makes the two in a balanced conversion state. Wherein: X t+1 represents the error output at the next moment, z t+1 represents the gating threshold at the next moment, X t represents the error output at the current moment represents the candidate memory at the next moment

Citation Information

Patent Citations

  • Full-automatic control method of mining hydraulic drilling rig

    CN104196517A

  • Multiple target type oriented mechanical arm self-adaptive grabbing method

    CN109986560A