A Fast Monocular Vision Three-Dimensional Target Detection Method for Underground Coal Mines

The method addresses the challenges of three-dimensional target detection in coal mines by leveraging mine-specific features and camera calibration to enhance detection accuracy and reduce computational demands, achieving efficient and cost-effective real-time performance.

CN115984766BActive Publication Date: 2025-07-15CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211571246.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-07-15
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

The existing three-dimensional object detection methods of monocular cameras have problems such as narrow perception range, insufficient information, large calculation amount, high energy consumption, time-consuming and labor-consuming data acquisition and labeling in the underground environment of coal mines, making it difficult to achieve fast and accurate three-dimensional object detection.

Method used

Build a unique two-dimensional object detection data set under the coal mine, train a neural network model, combine camera internal and external parameter calibration and lidar assistance to calculate the target three-dimensional coordinates and enclosing frames, and use TensorRT to accelerate inference to design a fast three-dimensional object detection framework suitable for the underground environment of coal mines.

Benefits of technology

It realizes efficient real-time three-dimensional object detection in coal mines, reduces calculation requirements, improves identification accuracy and efficiency, reduces data collection and labeling costs, and is suitable for underground equipment deployment of coal mines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984766B_ABST
    Figure CN115984766B_ABST
Patent Text Reader

Abstract

The present invention discloses a rapid monocular vision three-dimensional target detection method for underground coal mines, which includes the following steps: S1. First, use the cameras installed underground to collect image data containing specified targets and perform two-dimensional box annotation to construct a target detection data set; S2. Construct a two-dimensional target detection network model and train it on the constructed target detection data set to obtain a trained two-dimensional target detection model; S3. Define the world coordinate system, calibrate the camera internal parameters using the checkerboard calibration method, and use lidar to assist in calibrating the external parameters between the camera and the world coordinate system; S4. Input the image into the two-dimensional target detection model to obtain the target two-dimensional bounding box, and calculate the target three-dimensional coordinates in combination with the internal and external camera parameters; S5. Calculate the three-dimensional bounding box of the target in the camera coordinate system using the three-dimensional coordinates of the target, the heading angle of the target along the roadway direction, and the length, width, and height of the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for rapid monocular vision three-dimensional target detection in coal mines. Background Art

[0002] In recent years, algorithms based on lidar have greatly improved the accuracy of three-dimensional target detection. However, the disadvantages of lidar are: high cost and susceptibility to the surrounding environment. When the lidar module fails, three-dimensional detection based on monocular camera data can improve the robustness of the detection system. Therefore, how to achieve reliable and accurate three-dimensional detection based on camera data is particularly important. Compared with lidar-based methods, methods for estimating three-dimensional bounding boxes only from images face greater challenges because recovering three-dimensional information from two-dimensional input data is an ill-posed problem. However, despite this inherent difficulty, in the past few years, image-based three-dimensional target detection methods have been widely used in the computer vision community.

[0003] However, current research on monocular three-dimensional target detection methods focuses on the field of autonomous driving, where the cameras for perception are installed on vehicles. Well-known datasets include KITTI, Waymo, NuScenes, etc. There is little research in this field for fixed cameras and coal mine underground environments. With the development of technology and the need for safety, various automated devices have emerged in coal mines, such as inspection robots and autonomous navigation transport vehicles. For these devices, accurate positioning information is essential. Therefore, using a camera installed on the top of the roadway to perform three-dimensional target detection on autonomous vehicles inside the roadway is not only beneficial for assisting the vehicle in positioning, but also beneficial for managers to monitor it. For the complex environment in coal mines, if one wants to use a monocular camera installed on the top of the roadway for three-dimensional target detection, the existing methods have the following problems:

[0004] (1) The camera installed on the top of the roadway has a narrow and long sensing range, and there is little information available for extraction. The coal mine underground has a long and narrow structure. The camera on the top of the roadway observes a farther distance than the camera installed on the vehicle, and the proportion of the target in the image changes greatly. Moreover, due to the relatively dim light underground, both the depth features and visual features available for neural network extraction become fewer, and it is difficult for existing methods to accurately infer the three-dimensional information of the target.

[0005] (2) Existing methods rely on a large amount of labeled data for training, and the processes of data collection and annotation are time-consuming and laborious.

[0006] (3) Although existing methods have achieved a high inference speed, they rely on high-performance graphics cards, still have a large amount of calculations, high energy consumption, and are difficult to be transplanted into underground equipment. Summary of the Invention

[0007] Objective of the Invention: The objective of the present invention is to overcome the deficiencies of the above-mentioned existing technologies and realize a method for rapid three-dimensional target detection using a monocular camera installed underground in a coal mine. Compared with the existing monocular three-dimensional target detection methods, the present invention makes full use of the unique environmental characteristics of the mine, has lower requirements for the computing power of the equipment, better real-time performance, and is more suitable for the environment underground in a coal mine.

[0008] Technical Solution: To achieve the objective of the present invention, the present invention proposes a method for rapid monocular vision three-dimensional target detection underground in a coal mine. The method includes the following steps:

[0009] S1. Construct a two-dimensional target detection data set: Use a monocular camera installed underground in a coal mine to collect a data set with specified targets, determine the number of target categories in the data set, label the two-dimensional boxes and categories of the targets in each image sample, and randomly shuffle the obtained images and corresponding labels and divide them into a training set and a test set;

[0010] S2. Train a neural network model: Train on the training set and test on the test set, and obtain the optimal neural network model by adjusting parameters;

[0011] S3. Calibrate the internal and external parameters of the camera: Define the world coordinate system, use the checkerboard calibration method to calibrate the internal parameters of the camera to obtain the camera projection matrix P, and use lidar assistance to complete the calibration of the external parameters of the camera and the world coordinate system to obtain the transformation matrix T from the world coordinate system to the camera coordinate system W2C ;

[0012] S4. Calculate the three-dimensional coordinates of the target: Input the image into the two-dimensional target detection model obtained in step S2 to obtain the two-dimensional border of the target, calculate the two-dimensional coordinates C of the center of the bottom edge of the border 2d , and calculate the three-dimensional coordinates C of the target in combination with the internal and external parameters of the camera;

[0013] S5. Calculate the three-dimensional bounding box of the target: Use the three-dimensional coordinates C of the target obtained in step S4, the heading angle α of the target along the roadway direction, and the prior information of the length, width, and height [l, w, h] of the target to calculate the three-dimensional bounding box of the target in the camera coordinate system.

[0014] Furthermore, in step S1, the specific method for constructing the two-dimensional target detection data set is as follows:

[0015] S11. Use a monocular camera installed underground in a coal mine to collect target image data of different types of targets in different environments. Different environments include different lighting conditions and different postures of the targets;

[0016] S12. Determine the number of target categories to be detected. Use the labelImg tool to perform two-dimensional bounding box annotation on the collected images and export the labels. Each image corresponds to one label, which contains the categories of all targets in the image and the two-dimensional bounding boxes. Divide the annotated dataset into a training set and a test set, including the images and their corresponding label files.

[0017] Further, in step S2, the specific method for training the two-dimensional object detection model is as follows:

[0018] S21. Use the target images in the training set as the input, and the target categories and two-dimensional bounding boxes as the output to train the neural network model, and perform tests and parameter tuning on the test set to obtain a trained neural network model. This neural network model can use the target image as the input, and the output is the target categories in the target image and the two-dimensional bounding boxes in the image;

[0019] S22. Use TensorRT to further improve the inference speed of the trained neural network model to obtain an accelerated neural network model.

[0020] Further, in step S3, the camera internal and external parameter calibration methods are as follows:

[0021] S31. Use the checkerboard calibration method to calibrate the camera internal parameters to obtain the camera projection matrix P;

[0022] S32. Use the lidar to assist in calibrating the external parameter matrix. The external parameter refers to the transformation matrix from the world coordinate system to the camera coordinate system. The lidar and the camera have the same installation height. Use the MATLAB tool to calibrate the transformation matrix between the camera and the radar to obtain the transformation matrix T from the lidar to the camera L2C ;

[0023] S33. Define the world coordinate system: The origin is the projection point of the radar origin on the ground, the xOy plane is the ground, the y-axis points forward along the roadway, the x-axis points to the right, and the z-axis points upward, satisfying the right-hand coordinate system. Measure the installation height h of the lidar. According to the definition of the world coordinate system, the translation vector t between the radar coordinate system and the world coordinate system L2W =(0, 0, h);

[0024] S34. Obtain the coordinates of the two points A L (x1, y1, z1), B L (x2, y2, z2) on the roadway edge in the point cloud that are farthest apart. Then and is parallel to the y-axis of the world coordinate system. The subscript L indicates that this coordinate is in the lidar coordinate system. From the dot product of vectors, the cosine value of the angle between the direction vector of the y-axis of the to the lidar coordinate system is:

[0025]

[0026] Among them, normalize means converting the vector into a unit vector. to The rotation axis is obtained by cross - multiplying the vectors:

[0027]

[0028] Denote According to the Rodriguez rotation formula, the to rotation matrix R1 is obtained:

[0029]

[0030] S35. Determine the ground normal vector, thereby determining the rotation between the other two axes of the two coordinate systems. In the acquired point cloud, use the plane fitting method to fit the ground part of the point cloud to obtain the normal vector of the ground in the radar coordinate system.

[0031] Same as in S34 to For the calculation steps of the rotation matrix, calculate the vector to the normal vector of the xOy plane in the radar coordinate system The rotation matrix is R2, then the rotation matrix R L2W from the radar coordinate system to the world coordinate system is:

[0032] R L2W = R1R2

[0033] Calculate the transformation matrix from the radar coordinate system to the world coordinate system as:

[0034]

[0035] The transformation matrix from the world coordinate system to the radar coordinate system is the inverse matrix T W2L = T' L2W ;

[0036] S36. Calculate the transformation matrix T L2C from the world coordinate system to the camera coordinate system by T W2L and T W2C :

[0037] T W2C = T W2L ·T L2C .

[0038] Furthermore, the specific method for calculating the target three - dimensional coordinates in step S4 is as follows:

[0039] S41. Utilize T W2C to transform the origin O W =(0, 0, 0) and the ground normal vector in the world coordinate system into the camera coordinate system:

[0040]

[0041] Denote O C =(x o , y o , z o ) and

[0042] S42. Obtain the ground equation G(x, y, z) in the camera coordinate system from the point - normal equation of the plane:

[0043] G(x, y, z): n1(x - x o ) + n2(y - y o ) + n3(z - z o ) = 0

[0044] S43. Input the target image into the neural network model obtained in step S2 to obtain the categories and 2D bounding boxes of all targets, and calculate the bottom - center coordinates C 2d =(x 2d , y 2d ) of the target 2D bounding box;

[0045] S44. Use the pinhole camera projection model for back - projection to calculate the 3D coordinate representation of the 2D coordinate C 2d :

[0046] C 3d =(z 3d (x 2d - c x ) / f x , z 3d (y 2d - c y ) / f y , z 3d )

[0047] where f x and f y respectively represent the focal lengths in the x - axis and y - axis directions of the camera, obtained from the camera projection matrix P, that is: f x = P[0][0], f y = P[1][1], z 3d ∈(0, ∞] represents the depth;

[0048] S45. Substitute C 3d into the ground equation to calculate C 3dthe unknown z in 3d :

[0049] z 3d =(n1x o +n2y o +n3z o ) / (n1(x 2d -c x ) / f x +n2(y 2d -c y ) / f y +n3)

[0050] S46. Use the obtained three-dimensional coordinate C 3d , calculate the center coordinate C of the target bottom surface, and calculate C 3d The coordinate in the world coordinate system is:

[0051] C W =T C2W C 3d =(x W , y W , z W )

[0052] where, T C2W =T′ W2C , T′ W2C is the inverse matrix of T W2C . The calculated C W and the target bottom center C′ W have an offset, and this offset changes with the heading angle α of the target. Take the target heading angle along the roadway direction, then α is equal to the angle between the vector corresponding to the projection of the x-axis of the projection camera coordinate system on the ground and the y-axis of the world coordinate system. According to the geometric relationship, the magnitude of the offset offset is:

[0053] offset=(|l·sinα|+|w·cosα|) / 2

[0054] where, l and w are the length and width of the target respectively. According to the definition of the world coordinate system, calculate the bottom center coordinate C′ W of the target in the world coordinate system:

[0055] C′ W =(x W -|cosα·offset|, y W +|sinα·offset|, z W )

[0056] The coordinate calculation of the target bottom center in the camera coordinate system is:

[0057] C=T W2C C′W 。

[0058] Furthermore, the specific method for calculating the target three-dimensional bounding box in step S5 is as follows:

[0059] S51. Calculate the coordinates of the 8 vertices of the three-dimensional bounding box in the target coordinate system. Define the target coordinate system: the xOz plane coincides with the target bottom surface, that is, coincides with the ground. The coordinate origin is located at the center point of the target bottom surface. The y-axis is perpendicular to the ground and points downward. The x-axis is along the target forward direction, and the z-axis points to the left of the target forward direction, conforming to the right-hand coordinate system. Let the origin of the target coordinate system be O V =(0, 0, 0), then the coordinates of the 8 vertices of the three-dimensional bounding box are calculated as follows:

[0060]

[0061] where l, w, and h respectively represent the length, width, and height of the target;

[0062] S52. Rotate the eight vertices to the camera coordinate system to obtain the three-dimensional bounding box in the camera coordinate system. Use the ground normal vector and the target heading angle α obtained in step 4 to perform two rotations on the target, and then perform a translation to obtain the coordinates of the eight vertices in the camera coordinate system. Using the Rodriguez formula, similar to the calculation steps of the rotation matrix in step 3 to calculate the rotation matrix from the direction vector (0, -1, 0) in the negative y-axis direction of the camera to as R n , calculate the rotation matrix along the target heading angle α direction:

[0063]

[0064] Rotate the eight vertices to get:

[0065] V′ = R α R n V

[0066] Let the coordinates of the center of the target bottom surface in the camera coordinate system calculated in step 4 be C = (x c , y c , z c ), then:

[0067]

[0068] The obtained V″ is the coordinates of the eight vertices in the camera coordinate system;

[0069] S53. Project the coordinates of the 8 vertices in the camera coordinate system onto the image coordinate system to obtain the two-dimensional coordinates of the three-dimensional bounding box on the image. First, project using the camera projection matrix P

[0070] V‴ = PV″

[0071] The x and y values of each point are divided by the pixel depth to obtain its coordinates in the image coordinate system:

[0072]

[0073] where x n , y n , z n respectively represent the coordinate values of the nth vertex in V‴.

[0074] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:

[0075] First, the present invention designs an efficient and real-time three-dimensional object detection framework for the unique environment and camera installation method in underground coal mines, greatly improving the speed of three-dimensional object detection. TensorRT is used to accelerate the detection model during deployment, and an inference speed of 50 frames per second is achieved on an NVIDIA 3060 graphics card.

[0076] Second, since the present invention uses a two-dimensional object detection model, the accuracy and efficiency of identifying objects are greatly improved. By constructing a reasonable data set, the object detection model can achieve a high accuracy rate, reducing the occurrence of misdetection. In addition, the two-dimensional object detection model can be flexibly adjusted and replaced, and the coupling degree between modules is low.

[0077] Third, since the present invention only needs to label the two-dimensional information of the object for model training, the cost of data collection and data annotation is greatly reduced, saving manpower and material resources, and being convenient for transplantation. Description of the Drawings

[0078] Figure 1 is the overall flowchart of the present invention.

[0079] Figure 2 is the schematic diagram of the calculation of the three-dimensional coordinates of the object of the present invention.

[0080] Figure 3 is the schematic diagram of the calculation of the offset of the center of the bottom of the object of the present invention. Detailed Embodiments

[0081] The following further describes the present invention with reference to the drawings.

[0082] The present invention discloses a method for rapid monocular vision three-dimensional object detection in underground coal mines. The overall flowchart is as shown in the appendix Figure 1 and the method includes the following steps:

[0083] S1. Construct a 2D object detection dataset: Use a monocular camera installed in the coal mine to collect a dataset with specified objects, determine the number of object categories in the dataset, annotate the 2D bounding boxes and categories of the objects in each image sample, and randomly shuffle the obtained images and corresponding labels, and divide them into a training set and a test set;

[0084] S2. Train a neural network model: Train on the training set and test on the test set, and obtain the optimal neural network model by tuning parameters;

[0085] S3. Calibrate the camera internal and external parameters: Define the world coordinate system, use the checkerboard calibration method to calibrate the camera internal parameters to obtain the camera projection matrix P, and use the lidar to assist in calibrating the external parameters of the camera and the world coordinate system to obtain the transformation matrix T from the world coordinate system to the camera coordinate system W2C ;

[0086] S4. Calculate the 3D coordinates of the object: Input the image into the 2D object detection model obtained in step S2 to obtain the 2D bounding box of the object, and calculate the 2D coordinates C of the center of the bottom edge of the bounding box 2d , and calculate the 3D coordinates C of the object in combination with the camera internal and external parameters;

[0087] S5. Calculate the 3D bounding box of the object: Use the 3D coordinates C of the object obtained in step S4, the heading angle α of the object along the roadway direction, and the prior information of the object length, width, and height [l, w, h] to calculate the 3D bounding box of the object in the camera coordinate system.

[0088] Furthermore, in step S1, the specific method for constructing the 2D object detection dataset is as follows:

[0089] S11. Use a monocular camera installed in the coal mine to collect object image data of different types of objects in different environments. Different environments include different lighting conditions and different poses of the objects;

[0090] S12. Determine the number of object categories to be detected, use the labellmg tool to perform 2D bounding box annotation on the collected images and export the labels. Each picture corresponds to a label, which contains the categories and 2D bounding boxes of all the objects in the picture. Divide the annotated dataset into a training set and a test set, including the pictures and the corresponding label files.

[0091] Furthermore, in step S2, the specific method for training the 2D object detection model is as follows:

[0092] S21. Use the object images in the training set as the input, the object categories and 2D bounding boxes as the output to train the neural network model, and perform testing and parameter tuning on the test set to obtain the trained neural network model. This neural network model can use the object image as the input, and the output is the object categories in the object image and the 2D bounding boxes in the image;

[0093] S22. The trained neural network model uses TensorRT to further improve the inference speed, obtaining an accelerated neural network model.

[0094] Furthermore, in step S3, the calibration method for the camera's internal and external parameters is as follows:

[0095] S31. Use the checkerboard calibration method to calibrate the camera's internal parameters to obtain the camera projection matrix P;

[0096] S32. Use lidar for auxiliary calibration of the external parameter matrix. The external parameter refers to the transformation matrix from the world coordinate system to the camera coordinate system. The lidar and the camera have the same installation height. Use the MATLAB tool to calibrate the transformation matrix between the camera and the lidar to obtain the transformation matrix T from the lidar to the camera L2C ;

[0097] S33. Define the world coordinate system: The origin is the projection point of the lidar origin on the ground. The xOy plane is the ground. The y-axis points forward along the roadway, the x-axis points to the right, and the z-axis points upward, satisfying the right-hand coordinate system. Measure the installation height h of the lidar. From the definition of the world coordinate system, the translation vector t between the lidar coordinate system and the world coordinate system L2W =(0, 0, h);

[0098] S34. Obtain the coordinates A L (x1, y1, z1), B L (x2, y2, z2) of the two points on the roadway edge that are farthest apart in the point cloud. Then is parallel to the y-axis of the world coordinate system. The subscript L indicates that this coordinate is in the lidar coordinate system. From the dot product of vectors, to the direction vector of the y-axis of the lidar coordinate system The cosine value of the angle between them is:

[0099]

[0100] where normalize means converting the vector to a unit vector. from to

[0101]

[0102] Denote According to the Rodriguez rotation formula, the rotation matrix R1 from to is obtained:

[0103]

[0104] S35. Determine the ground normal vector, thereby determining the rotation between the other two axes of the two coordinate systems. In the acquired point cloud, use the plane fitting method to fit the ground part in the point cloud to obtain the normal vector of the ground in the radar coordinate system.

[0105] Same as in S34 to Calculation steps of the rotation matrix, calculate the vector to the normal vector of the xOy plane in the radar coordinate system The rotation matrix to is R2, then the rotation matrix R from the radar coordinate system to the world coordinate system L2W is:

[0106] R L2W = R1R2

[0107] Calculate the transformation matrix from the radar coordinate system to the world coordinate system as:

[0108]

[0109] The transformation matrix from the world coordinate system to the radar coordinate system is its inverse matrix T W2L = T' L2W ;

[0110] S36. From T L2C and T W2L calculate the transformation matrix T from the world coordinate system to the camera coordinate system W2C :

[0111] T W2C = T W2L ·T L2C .

[0112] Furthermore, the specific method for calculating the target three-dimensional coordinates in step S4 is as follows:

[0113] S41. Use T W2C to transform the origin O W =(0, 0, 0) and the ground normal vector in the world coordinate system to the camera coordinate system:

[0114]

[0115] Denote O C =(x o , y o , z o ) and

[0116] S42. Obtain the ground equation G(x, y, z) in the camera coordinate system from the point-normal equation of the plane:

[0117] G(x, y, z): n1(x - x o ) + n2(y - y o ) + n3(z - z o ) = 0

[0118] S43. Input the target image into the neural network model obtained in step S2 to obtain the categories and 2D bounding boxes of all targets, and calculate the bottom center coordinate C of the target 2D bounding box 2d =(x 2d , y 2d )

[0119] S44. Use the pinhole camera projection model for back-projection to calculate the 3D coordinate representation of the 2D coordinate C 2d :

[0120] C 3d =(z 3d (x 2d - c x ) / f x , z 3d (y 2d - c y ) / f y , z 3d )

[0121] where f x and f y respectively represent the focal lengths in the x-axis and y-axis directions of the camera, obtained from the camera projection matrix P, i.e., f x = P[0][0], f y = P[1][1], z 3d ∈ (0, ∞] represents the depth

[0122] S45. Substitute C 3d into the ground equation to calculate the unknown z 3d in C 3d :

[0123] z 3d =(n1x o + n2y o + n3z o ) / (n1(x 2d - c x ) / f x + n2(y 2d - c y ) / f y + n3)

[0124] S46. Use the obtained 3D coordinate C 3d to calculate the bottom center coordinate C of the target, and calculate C 3dThe coordinates in the world coordinate system are:

[0125] C W = T C2W C 3d = (x W , y W , z W )

[0126] Wherein, T C2W = T′ W2C , T′ W2C is the inverse matrix of T W2C , and the calculated C W has an offset from the center C′ of the bottom of the target W . This offset changes with the heading angle α of the target. When the heading angle of the target is along the roadway direction, α is equal to the angle between the vector corresponding to the projection of the x-axis of the projection camera coordinate system on the ground and the y-axis of the world coordinate system. According to the geometric relationship, the magnitude of the offset offset is:

[0127] offset = (|l·sinα| + |w·cosα|) / 2

[0128] Wherein, l and w are the length and width of the target respectively. According to the definition of the world coordinate system, the coordinates of the center C′ of the bottom of the target in the world coordinate system are calculated W :

[0129] C′ W = (x W - |cosα·offset|, y W + |sinα·offset|, z W )

[0130] The coordinates of the center of the bottom of the target in the camera coordinate system are calculated as:

[0131] C = T W2C C′ W .

[0132] Furthermore, the specific method for calculating the three-dimensional bounding box of step S5 is as follows:

[0133] S51. Calculate the coordinates of the 8 vertices of the three-dimensional bounding box in the target coordinate system. Define the target coordinate system: the xOz plane coincides with the bottom surface of the target, that is, coincides with the ground, the coordinate origin is located at the center point of the bottom surface of the target, the y-axis is perpendicular to the ground and points downward, the x-axis is along the forward direction of the target, and the z-axis points to the left of the forward direction of the target, conforming to the right-hand coordinate system. Let the origin of the target coordinate system be O V = (0, 0, 0), then the coordinates of the 8 vertices of the three-dimensional bounding box are calculated as follows:

[0134]

[0135] Among them, l, w, and h respectively represent the length, width, and height of the target;

[0136] S52. Rotate the eight vertices into the camera coordinate system to obtain the three-dimensional bounding box in the camera coordinate system, and use the ground normal vector in the camera coordinate system obtained in step 4 and the target heading angle α, perform two rotations on the target, and then perform a translation to obtain the eight vertex coordinates in the camera coordinate system. Using the Rodriguez formula, similar to the to calculation steps of the rotation matrix in step 3, calculate the rotation matrix from the direction vector (0, -1, 0) in the negative y-axis direction of the camera to as R n , and calculate the rotation matrix along the target heading angle α direction:

[0137]

[0138] Rotate the eight vertices to get:

[0139] V′ = R α R n V

[0140] Let the center coordinates of the target bottom surface in the camera coordinate system calculated in step 4 be C = (x c , y c , z c ), then:

[0141]

[0142] The obtained V″ is the eight vertex coordinates in the camera coordinate system;

[0143] S53. Project the 8 vertex coordinates in the camera coordinate system onto the image coordinate system to obtain the two-dimensional coordinates of the three-dimensional bounding box on the image. First, project using the camera projection matrix P:

[0144] V″′ = PV″

[0145] The x and y values of each point are divided by the pixel depth to obtain its coordinates in the image coordinate system:

[0146]

[0147] Among them, x n , y n , z n respectively represent the coordinate values of the nth vertex in V″′.

[0148] The above has introduced in detail a method for rapid monocular vision three-dimensional target detection in coal mines provided by the embodiments of the present invention. For those of ordinary skill in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A rapid monocular vision three-dimensional target detection method for underground coal mines, characterized in that, The method includes the following steps: S1. Construct a two-dimensional object detection dataset: Use a monocular camera installed in a coal mine to collect a dataset with specified objects, determine the number of object categories in the dataset, label the two-dimensional bounding boxes and categories of the objects in each image sample, and randomly shuffle the obtained images and corresponding labels, and divide them into a training set and a test set; S2. Train a neural network model: Train on the training set and test on the test set, and obtain the optimal neural network model by adjusting parameters; S3. Camera intrinsic and extrinsic parameter calibration: Define the world coordinate system, calibrate the camera intrinsics using the checkerboard calibration method to obtain the camera projection matrix P, and use lidar assistance to complete the calibration of the extrinsic parameters between the camera and the world coordinate system to obtain the transformation matrix T from the world coordinate system to the camera coordinate system W2C ; S4. Target three-dimensional coordinate calculation: Input the image into the two-dimensional target detection model obtained in step S2 to obtain the target two-dimensional bounding box, and calculate the two-dimensional coordinates C of the center of the bottom edge of the bounding box; 2d and calculate the target three-dimensional coordinates C in combination with the internal and external camera parameters; S5. Calculate the three-dimensional bounding box of the object: Use the three-dimensional coordinates C of the object obtained in step S4, the heading angle α of the object along the roadway direction, and the prior information of the object's length, width, and height [l, w, h] to calculate the three-dimensional bounding box of the object in the camera coordinate system.

2. The rapid monocular vision three-dimensional target detection method for coal mines underground according to claim 1, characterized in that, In step S1, the specific method for constructing the two-dimensional object detection dataset is as follows: S11. Use a monocular camera installed in a coal mine to collect object image data of different types of objects in different environments. Different environments include different lighting conditions and different poses of the objects; S12. Determine the number of object categories to be detected, use the labelImg tool to perform two-dimensional bounding box annotation on the collected images and export the labels. Each picture corresponds to a label, which contains the categories and two-dimensional bounding boxes of all the objects in the picture. Divide the annotated dataset into a training set and a test set, including the pictures and the corresponding label files.

3. A rapid monocular vision three-dimensional target detection method for coal mines underground according to claim 1, characterized in that, In step S2, the specific method for training the two-dimensional object detection model is as follows: S21. Use the object images in the training set as the input, and the object categories and two-dimensional bounding boxes as the output to train the neural network model, and test and adjust parameters on the test set to obtain a trained neural network model. This neural network model can use the object image as the input, and the output is the object categories in the object image and the two-dimensional bounding boxes in the image; S22. Use TensorRT to further improve the inference speed of the trained neural network model to obtain an accelerated neural network model.

4. A rapid monocular vision three-dimensional target detection method for underground coal mines according to claim 1, characterized in that, In step S3, the method for calibrating the camera internal and external parameters is as follows: S31. Use the checkerboard calibration method to calibrate the camera internal parameters to obtain the camera projection matrix P; S32. Use lidar to assist in calibrating the external parameter matrix. The external parameter refers to the transformation matrix from the world coordinate system to the camera coordinate system. The lidar and the camera have the same installation height. Use the MATLAB tool to calibrate the transformation matrix between the camera and the radar to obtain the transformation matrix T from the lidar to the camera L2C ; S33. Define the world coordinate system: The origin is the projection point of the radar origin on the ground. The xOy plane is the ground, the y-axis is along the roadway forward, the x-axis is to the right, and the z-axis is upward, satisfying the right-hand coordinate system. Measure the installation height h of the lidar. According to the definition of the world coordinate system, the translation vector t from the lidar coordinate system to the world coordinate system L2W =(0, 0, h); S34. Obtain the coordinates A of the two points on the roadway edge in the point cloud that are farthest apart L (x1, y1, z1), B L (x2, y2, z2), then and is parallel to the y-axis of the world coordinate system. The subscript L indicates that this coordinate is in the lidar coordinate system. From the dot product of vectors, to the direction vector of the y-axis of the radar coordinate system The cosine value of the angle between them is: Among them, normalize means converting the vector into a unit vector. to The rotation axis is obtained by cross multiplying the vectors: Record Obtained according to the Rodriguez rotation formula to the rotation matrix R1 of S35. Determine the ground normal vector, thereby determining the rotation between the other two axes of the two coordinate systems. In the acquired point cloud, use the plane fitting method to fit the ground part in the point cloud to obtain the normal vector of the ground in the radar coordinate system. Same as in S34 to For the calculation steps of the rotation matrix, calculate the vector to the normal vector of the xOy plane in the radar coordinate system The rotation matrix is R2, then the rotation matrix R from the radar coordinate system to the world coordinate system L2W is: R L2W = R1R2 Calculate the transformation matrix from the radar coordinate system to the world coordinate system as: The transformation matrix from the world coordinate system to the radar coordinate system is its inverse matrix T W2L = T' L2W ; S36. From T L2C and T W2L Calculate the transformation matrix T from the world coordinate system to the camera coordinate system W2C : T W2C = T W2L · T L2C 。 5. A rapid monocular vision three-dimensional target detection method for underground coal mines according to claim 1, characterized in that, The specific method for calculating the three-dimensional coordinates of the object in step S4 is as follows: S41. Use T W2C to transform the origin OW=(0, 0, 0) and the ground normal vector in the world coordinate system to the camera coordinate system: Denote O C =(x o , y o , z o ) and S42. Obtain the ground equation G(x, y, z) in the camera coordinate system from the point-normal form equation of the plane: G(x, y, z): n1(x - x o ) + n2(y - y o ) + n3(z - z o ) = 0 S43. Input the target image into the neural network model obtained in step S2 to obtain the categories and two-dimensional bounding boxes of all targets, and calculate the bottom center coordinates C 2d =(x 2d , y 2d ); S44. Perform back-projection using the pinhole camera projection model and calculate the three-dimensional coordinate representation of the two-dimensional coordinate C 2d : C 3d = (z 3d (x 2d - c x ) / f x , z 3d (y 2d - c y ) / f y , z 3d ) where f x and f y denote the focal lengths in the x-axis and y-axis directions of the camera, respectively, and are obtained from the camera projection matrix P, that is: f x = P[0][0], f y = P[1][1], z 3d ∈(0, ∞] represents the depth; S45. Substitute C 3d into the ground equation to calculate the unknown z 3d in C 3d : z 3d = (n1x o + n2y o + n3z o ) / (n1(x 2d - c x ) / f x + n2(y 2d - c y ) / f y + n3) S46. Use the obtained three-dimensional coordinates C 3d , calculate the center coordinates C of the target bottom surface, and calculate C 3d The coordinates in the world coordinate system are: C W = T C2W C 3d = (x W , y W , z W ) Among them, T C2W = T' W2C , where T' W2C is the inverse matrix of T W2C . The calculated C W has an offset from the target bottom center C' W . This offset changes with the change of the target course angle α. Taking the target course angle along the roadway direction, α is equal to the angle between the vector corresponding to the projection of the x-axis of the projection camera coordinate system on the ground and the y-axis of the world coordinate system. According to the geometric relationship, the magnitude of the offset offset is: offset = (|l·sinα| + |w·cosα|) / 2 Wherein, l and w are the length and width of the target respectively, and according to the definition of the world coordinate system, the bottom center coordinate C' of the target in the world coordinate system is calculated W : C′ W =(x W -|cosα·offset|, y W +|sinα·offset|, z W ) Calculate the coordinates of the center of the bottom of the object in the camera coordinate system as: C = T W2C C' W 。 6. The rapid monocular vision three-dimensional target detection method for underground coal mines according to claim 1, characterized in that The specific method for calculating the three-dimensional bounding box of the object in step S5 is as follows: S51. Calculate the coordinates of the 8 vertices of the three-dimensional bounding box in the target coordinate system. Define the target coordinate system: the xOz plane coincides with the target bottom surface, that is, coincides with the ground, the coordinate origin is located at the center point of the target bottom surface, the y-axis is perpendicular to the ground and points downward, the x-axis is along the target's forward direction, and the z-axis points to the left of the target's forward direction, conforming to the right-hand coordinate system. Let the origin of the target coordinate system be O V =(0, 0, 0), then the calculation of the coordinates of the 8 vertices of the three-dimensional bounding box is as follows: Where l, W, and h respectively represent the length, width, and height of the object; S52. Rotate the eight vertices into the camera coordinate system to obtain the three-dimensional bounding box in the camera coordinate system, and use the ground normal vector in the camera coordinate system obtained in step 4 and the target heading angle α to perform two rotations on the target, and then perform a translation to obtain the coordinates of the eight vertices in the camera coordinate system. Using the Rodriguez formula, similar to the calculation steps of the rotation matrix from to in step 3, calculate the rotation matrix from the direction vector (0, -1, 0) in the negative y-axis direction of the camera to n to be R Rotate the eight vertices to obtain: V′ = R α R n V Let the center coordinates of the target bottom surface in the camera coordinate system calculated in step 4 be C=(x c , y c , z c ), then: The obtained V″ is the coordinates of the eight vertices in the camera coordinate system; S53. Project the coordinates of the 8 vertices in the camera coordinate system onto the image coordinate system to obtain the two-dimensional coordinates of the three-dimensional bounding box on the image. First, use the camera projection matrix P for projection: V″′ = PV″ The x and y values of each point are divided by the pixel depth to obtain its coordinates in the image coordinate system: where x n , y n , z n respectively represent the coordinate values of the n-th vertex in V″′.

Citation Information

Patent Citations

  • Target detection method and electronic equipment

    CN113903028A

  • Three-dimensional environment target detection method based on multi-sensor fusion

    CN115049821A