A coal rock identification method based on multi-modal view frustum point cloud fusion

CN117636034BActive Publication Date: 2026-08-21CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311641762.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2026-08-21
Estimated Expiration
2043-12-01

AI Technical Summary

Technical Problem

[0005]针对上述现有技术存在的问题,本发明提供一种基于多模态视锥点云融合的煤岩识别方法,该方法能够快速准确地识别出煤岩位置,能够解决由于工作面煤尘大、光线暗、噪声多等恶劣环境导致的煤岩识别精度不高的问题,进而为采煤机的智能化开采提供技术支撑

Benefits of technology

[0047]本发明中,通过在采煤机摇臂上安装激光雷达和相机,可以便于在截割煤岩的过程中同步地采集到截割煤块过程中煤岩的激光点云和图像数据,进而能够为煤岩的快速精准识别提供可靠的数据支持;以Mask R-CNN网络为基础模型,选用MobileNetV3作为主干网络,并将Dilated-CBAM机制和Inception结构融入至MobileNetV3主干网络,这样,不仅可以减少模型体积和网络训练时间,还能有效增加网络对不同尺度特征的适应性,有助于更好的提取到目标的底层和边缘特征,进而大幅提升模型的检测性能;先利用相机所采集到的图像数据对改进后的Mask R-CNN网络模型进行训练,计算并输出图像中煤岩的目标边框位置,再根据激光雷达和相机之间的投影关系生成目标边框的视锥点云,并将生成的视锥点云和激光点云进行融合,这样,可以有效减少点云的计算规模和空间搜索范围,有利于点云网络更好的识别目标;先在Pointnet网络的基础上添加self-attention机制构成self-attention Pointnet网络,再利用self-attention Pointnet网络对视锥范围内的融合点云进行分割,并通过3D边框回归网络预测目标点云的边框参数,能够很好地获取点云空间局部特征信息,提升网络检测小目标物体和复杂场景的能力,有利于实现精确的三维检测。本发明基于激光雷达和相机融合技术,提出一种多模态视锥点云煤岩识别方法,可以充分利用激光点云的高度信息和图像的颜色信息,融合两者的优点,能够解决由于工作面煤尘大、光线暗、噪声多等恶劣环境导致的煤岩识别精度不高的问题,进而为采煤机的智能化开采提供技术支撑,对实现综采工作面的“无人化”或“少人化”开采具有重要的意义。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117636034B_ABST
    Figure CN117636034B_ABST
Patent Text Reader

Abstract

A coal rock identification method based on multi-modal view frustum point cloud fusion, performing cutting operation on coal rock test blocks, using laser radar and camera to collect laser point cloud and image data of coal rock in the process of cutting coal blocks in real time; using image data to train the improved Mask R-CNN network model, calculating and outputting the target bounding box position of coal rock in the image; generating the view frustum point cloud of the target bounding box according to the projection relationship, and fusing it with the laser point cloud; using the self-attention Pointnet network to segment the fused point cloud in the view frustum range, and predicting the bounding box parameters of the target point cloud through the 3D bounding box regression network; sequentially performing cutting operation on the remaining coal rock test blocks, and processing the collected laser point cloud and image data through steps two and three respectively to obtain the target view frustum fused point cloud, and then inputting it into the point cloud network for coal rock identification. The method can realize the rapid and accurate identification of coal rock.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of coal and rock identification technology, specifically relating to a coal and rock identification method based on multimodal view cone point cloud fusion. Background Technology

[0002] As one of the core pieces of equipment in fully mechanized mining faces, the coal mining machine is a crucial guarantee for achieving high-yield and high-efficiency coal mining. To improve the intelligence level of the coal mining machine, it is necessary to adaptively adjust the drum height and traction speed according to the distribution of coal and rock in the coal seam. However, fully mechanized mining faces are characterized by continuous movement and complex, ever-changing conditions. How to quickly and accurately identify the location of coal and rock strata in the cutting face remains a recognized technical challenge in the coal mining field and a key technological bottleneck restricting the intelligentization of coal mining.

[0003] Currently, scholars at home and abroad have conducted extensive research and proposed various methods for coal and rock identification, including: radioactive ray method, acoustic / vibration detection method, electromagnetic detection method, infrared detection method, image recognition method, point cloud recognition method, etc. The proposal of the above methods indicates that researchers have achieved certain results in the field of coal and rock identification, but the following problems still exist: (1) Radiation sources have certain radioactive hazards and weak penetration, which limits their application scope; (2) Due to the harsh environment of the coal mining face, such as high humidity, low light, and a lot of coal dust, the acquisition of images and signals is strongly interfered with, reducing the accuracy of acquisition; (3) Compared with other sensor signals, the color information and resolution of point clouds are low, resulting in unsatisfactory identification effects. Therefore, the above methods cannot achieve ideal identification effects in practical applications.

[0004] With the emergence of multiple perception modes, multimodal fusion technology has become an important component of computer vision. LiDAR point cloud and image fusion is a novel measurement technique proposed in recent years. It effectively combines LiDAR and cameras, integrating their advantages to enhance perception capabilities and achieve accurate environmental perception in all weather conditions. However, traditional fusion methods do not directly process point cloud data; instead, they extract features by projecting the point cloud onto a two-dimensional plane, resulting in the loss of spatial information from the point cloud. Therefore, there is an urgent need to provide a method that can fully utilize the complementarity of LiDAR point cloud location information and image color information for rapid and accurate identification of coal and rock. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a coal and rock identification method based on multimodal view cone point cloud fusion. This method can quickly and accurately identify the location of coal and rock, and can solve the problem of low coal and rock identification accuracy caused by harsh environments such as high coal dust, low light, and high noise in the working face, thereby providing technical support for intelligent mining of coal mining machines.

[0006] To achieve the above objectives, this invention provides a coal and rock identification method based on multimodal view cone point cloud fusion, specifically including the following steps:

[0007] Step 1: Before the coal mining machine starts cutting the coal and rock test block, install a lidar and a camera on the rocker arm of the coal mining machine, and establish communication connections between the lidar and the data processing equipment, and between the camera and the data processing equipment. According to the sequence of fully mechanized mining operations, perform the cutting operation on the first coal and rock test block, and at the same time use the lidar and camera to collect the laser point cloud and image data of the coal and rock in real time during the cutting process.

[0008] Step 2: Using the Mask R-CNN network as the base model, select MobileNetV3 as the backbone network, and integrate the Dilated-CBAM mechanism and Inception structure into the MobileNetV3 backbone network to obtain an improved Mask R-CNN network model; use the image data acquired by the camera to train the improved Mask R-CNN network model, calculate and output the target bounding box position of coal and rock in the image;

[0009] Step 3: Generate the view frustum point cloud of the target bounding box based on the projection relationship between the lidar and the camera, and fuse the generated view frustum point cloud with the lidar point cloud;

[0010] Step 4: Use the self-attention PointNet network to segment the fused point cloud within the view frustum, and use the 3D bounding box regression network to predict the bounding box parameters of the target point cloud;

[0011] Step 5: According to the sequence of fully mechanized mining operations, the remaining coal and rock test blocks are cut in sequence. Starting from the second coal and rock test block, the collected laser point cloud and image data are processed through Step 2 and Step 3 respectively to obtain the target's frustum fusion point cloud. Then, it is input into the self-attention Pointnet network and the 3D bounding box regression network to predict the shape and position information of coal and rock during the cutting process, so as to achieve rapid and accurate identification of coal and rock.

[0012] As a preferred option, the bounding box parameters of the target point cloud in step four include the target's center coordinates, length, width, height, and heading angle.

[0013] Furthermore, to improve the recognition performance, in step two, the method for outputting the target bounding box positions of coal and rock in the image using the improved Mask R-CNN network model is as follows:

[0014] S21: Input the image data acquired by the camera into the improved Mask R-CNN network model, and use the improved MobileNetV3 backbone network and feature pyramid network to extract multi-scale feature maps of the input image;

[0015] S22: Use the region proposal network to extract candidate region boxes that may contain the target on the feature map, and use the cropping and alignment operations of the ROIAlign layer to normalize the candidate region boxes;

[0016] S23: Perform regression and classification operations on the target to obtain the target's bounding box and confidence score.

[0017] Furthermore, to reduce model size and network training time, while improving the network's ability to distinguish between targets and background, and to broaden the network width and enhance its adaptability, the improvement process for the Mask R-CNN network model in step two is as follows:

[0018] MobileNetV3 is introduced as the backbone network of Mask R-CNN. At the same time, the traditional MobileNetV3 structure is modified by replacing the fifth block structure with max pooling and 3×3 convolution operations, and changing the number of channels in the last two blocks to 64 and 96 respectively, in order to extract more low-level and edge features.

[0019] The Dilated-CBAM attention mechanism is introduced and embedded into the depthwise separable convolutional structure of the MobileNetV3 model. After being concatenated to the second pointwise convolution of the inverse residual module with a stride of 1, it balances spatial and channel-dimensional features, enhancing the network's ability to distinguish between targets and background. The embedded module is named bneck_DC. In the Dilated-CBAM mechanism, target features are first input into the channel attention module, where average pooling and max pooling extract channel feature information. The obtained parameters are then stacked through a multilayer perceptron and activated by a sigmoid function to obtain the channel attention features M. c (F), Channel attention features M c (F) Calculate according to formula (1); then, input the features into the spatial attention module, aggregate spatial information through pooling operation, and obtain spatial attention features M using dilated convolutions with dilation rates of 4, 2 and 1 and sigmoid activation functions. s (F x Spatial attention feature M s (F x ) Calculate according to formula (2);

[0020]

[0021]

[0022] In the formula: σ(·) represents the sigmoid activation function, W0 and W1 represent the weights of the hidden layer and the output layer, respectively, and f4 3×3 f2 3×3 f1 3×3 These represent dilated convolutions with a kernel size of 3 and dilation rates of 4, 2, and 1, respectively.

[0023] The Inception structure is introduced and combined with the bneck structure with a stride of 2 to form a more efficient feature extraction module. The embedded module is named bneck_IN. The Inception structure has four branches. Features are input into the network structure of each branch for processing. Then, 3×3 and 5×5 DW convolutions are used for multi-scale information fusion. Spatial information pooling is achieved through Max Pooling. The concat operation is discarded. The features output from different branches are added together before output to broaden the network width and enhance the network adaptability.

[0024] Furthermore, in order to fully utilize the height information of the laser point cloud and the color information of the image, and to fuse the advantages of both to obtain a model with higher recognition accuracy, the fusion process of the view frustum point cloud and the laser point cloud in step three is as follows:

[0025] S31: Coordinate system calibration;

[0026] Before fusing the point cloud of the target frustum candidate region, the lidar coordinate system and the camera coordinate system are calibrated. The specific method is as follows:

[0027] (1) Transform the point (X) in the lidar coordinate system using rigid body transformation. w ,Y w Z w Transformed to camera coordinate system (X) c ,Y c Z c );

[0028] (2) The transformation between the camera coordinate system and the image coordinate system is achieved by using the central perspective projection method in the pinhole imaging principle;

[0029] (3) Using formula (3), the transformation from image coordinate system to pixel coordinate system is achieved through scaling and translation transformation, thereby completing the spatial alignment and registration of laser point cloud to camera image;

[0030]

[0031] In the formula: R and T represent rotation and translation matrices, f x and fy The image represents the camera's two-dimensional focal length, u0 and v0 represent the coordinates of the image origin, u and v represent the coordinates in the pixel coordinate system, and b represents the coordinates of the pixel coordinate system. x It is the offset between the two coordinate systems;

[0032] S32: View frustum generation;

[0033] The four corner points of the target bounding box are mapped to four rays in three-dimensional space through coordinate system transformation, thereby determining a tetrahedron with the camera optical center as a fixed point, and collecting all points in the view frustum to form a view frustum point cloud;

[0034] S33: Point cloud fusion:

[0035] Determine the positional relationship between the laser point cloud and the image target bounding box, and fuse the laser point cloud and the view frustum point cloud within the view frustum range.

[0036] Furthermore, to improve the segmentation effect of the fused point cloud, in the frustum generation process in step S32 of step three, the angle of the frustum is first adjusted to rotate it to the central viewpoint so that the central axis of the frustum is orthogonal to the image plane, realizing the transformation from the camera coordinate system to the frustum coordinate system; then, linear segmentation of the frustum is performed, and the segmentation height is set to u, 2u, 3u, ..., ku along the central axis of the frustum from near to far. For these generated point cloud cluster sequences, the self-attention Pointnet network is used to process them to extract the frustum features of each sequence, and the features of each sequence are combined to improve the segmentation effect of the fused point cloud.

[0037] Furthermore, in order to effectively enhance the model's attention to different features and increase the network's adaptability to features at different scales, the method for predicting the bounding box parameters of the target point cloud using the self-attention PointNet network and the 3D bounding box regression network in step four is as follows:

[0038] S41: Take the fused point cloud data extracted from the frustum region as input, and combine it with the one-hot classification vector of the two-dimensional object detection part to output the evaluation score of the category to which each point cloud belongs;

[0039] S42: Non-target point clouds are removed by masking operations, and instance segmentation is performed on the fused point cloud using a self-attention PointNet network;

[0040] S43: Utilize a 3D bounding box regression network and multilayer perception mechanism to achieve bounding box regression prediction of target point clouds.

[0041] Furthermore, to better capture local features of the point cloud space, in step four (S42), the method for instance segmentation of the fused point cloud using a self-attention PointNet network is as follows:

[0042] A self-attention mechanism is introduced on the basis of the PointNet network to construct a self-attention PointNet network. First, the input X is mapped to the query vector matrix Q, the key vector matrix K, and the value vector matrix V through three different linear transformations. Then, the original input is subjected to multiple non-shared parameter attention operations through formula (4) to calculate the output result head of the i-th attention head. i Then, the m heads are merged and spliced ​​using formula (5), and then the encoded result Y of the fused point cloud is obtained by formula (6) after linear mapping.

[0043] head i =Attention(QW Q ,KW K VW V (4);

[0044]

[0045] Y = MultiHead(X,X,X) (6);

[0046] In the formula, X is the fused point cloud set; W Q W K W V These are the weights for Q, K, and V, respectively; Attention(·) represents the attention module function for a single head. Where softmax(·) represents the normalization function, x t x s All are point clouds, d k Indicates the input dimension; For feature concatenation, MultiHead(·) represents the multi-head attention function.

[0047] In this invention, by installing a lidar and camera on the rocker arm of a coal mining machine, it is possible to simultaneously acquire lidar point cloud and image data of coal and rock during the coal cutting process, thereby providing reliable data support for rapid and accurate identification of coal and rock. Using the Mask R-CNN network as the base model and MobileNetV3 as the backbone network, the Dilated-CBAM mechanism and Inception structure are integrated into the MobileNetV3 backbone network. This not only reduces the model size and network training time but also effectively increases the network's adaptability to features at different scales, helping to better extract the underlying and edge features of the target, thus significantly improving the model's detection performance. First, the improved Mask R-CNN network model is trained using image data acquired by the camera to calculate and output the target bounding box position of the coal and rock in the image. Then, based on the projection relationship between the lidar and the camera, a view frustum point cloud of the target bounding box is generated, and the generated view frustum point cloud is fused with the lidar point cloud. This effectively reduces the computational scale and spatial search range of the point cloud, which is beneficial for the point cloud network to better identify the target. Finally, a self-attention mechanism is added to the PointNet network to construct a self-attention mechanism. The PointNet network, further enhanced by a self-attention PointNet network, segments the fused point cloud within the view frustum. A 3D bounding box regression network then predicts the bounding box parameters of the target point cloud, effectively acquiring local spatial feature information of the point cloud. This improves the network's ability to detect small objects and complex scenes, facilitating accurate 3D detection. This invention, based on LiDAR and camera fusion technology, proposes a multimodal view frustum point cloud coal and rock recognition method. It fully utilizes the height information of the LiDAR point cloud and the color information of the image, combining their advantages to solve the problem of low coal and rock recognition accuracy caused by harsh environments such as high coal dust, low light, and high noise levels at the working face. This provides technical support for intelligent mining of coal mining machines and is of great significance for achieving "unmanned" or "reduced-manpower" mining in fully mechanized mining faces. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the process of the present invention;

[0049] Figure 2 This is a diagram showing the framework structure of the improved Mask R-CNN network model in this invention;

[0050] Figure 3 This is a diagram illustrating the framework structure of the traditional MobileNetV3 backbone network in this invention.

[0051] Figure 4 This is a diagram illustrating the framework structure of the improved MobileNetV3 backbone network in this invention.

[0052] Figure 5 This is a schematic diagram of the viewing cone adjustment state in this invention;

[0053] Figure 6 This is a framework diagram of the self-attention Pointnet network in this invention. Detailed Implementation

[0054] The present invention will be further described below with reference to the embodiments.

[0055] like Figures 1 to 6 As shown, this invention provides a coal and rock identification method based on multimodal view cone point cloud fusion, specifically including the following steps:

[0056] Step 1: Before the coal mining machine starts cutting the coal and rock test block, install a lidar and a camera on the rocker arm of the coal mining machine, and establish communication connections between the lidar and the data processing equipment, and between the camera and the data processing equipment. According to the sequence of fully mechanized mining operations, perform the cutting operation on the first coal and rock test block, and at the same time use the lidar and camera to collect the laser point cloud and image data of the coal and rock in real time during the cutting process.

[0057] Step 2: Using the Mask R-CNN network as the base model, MobileNetV3 is selected as the backbone network. The Dilated-CBAM mechanism and Inception structure are integrated into the MobileNetV3 backbone network to obtain an improved Mask R-CNN network model, which highlights the image location information and expands the network width. The improved Mask R-CNN network model is trained using image data acquired by the camera to calculate and output the target bounding box position of coal and rock in the image.

[0058] Step 3: Generate the view frustum point cloud of the target bounding box based on the projection relationship between the lidar and the camera, and fuse the generated view frustum point cloud with the lidar point cloud to reduce the computational scale of the point cloud and the spatial search range.

[0059] Step 4: Use the self-attention PointNet network to segment the fused point cloud within the view frustum, and use a 3D bounding box regression network to predict the bounding box parameters of the target point cloud. As a preferred method, the bounding box parameters include information such as the center coordinates, length, width, height, and heading angle of the target.

[0060] Step 5: According to the sequence of fully mechanized mining operations, the remaining coal and rock test blocks are cut in sequence. Starting from the second coal and rock test block, the collected laser point cloud and image data are processed through Step 2 and Step 3 respectively to obtain the target's frustum fusion point cloud. Then, it is input into the self-attention Pointnet network and the 3D bounding box regression network to predict the shape and position of coal and rock during the cutting process, so as to achieve rapid and accurate identification of coal and rock.

[0061] To improve recognition accuracy, in step two, the method for outputting the target bounding box positions of coal and rock in the image using the improved Mask R-CNN network model is as follows:

[0062] S21: Input the image data acquired by the camera into the improved Mask R-CNN network model, and use the improved MobileNetV3 backbone network and Feature Pyramid Network (FPN) to extract multi-scale feature maps of the input image;

[0063] S22: Use the Region Proposal Network (RPN) to extract candidate region boxes that may contain the target on the feature map, and use the cropping and alignment operations of the ROI Align layer to normalize the candidate region boxes;

[0064] S23: Perform regression and classification operations on the target to obtain the target's bounding box and confidence score.

[0065] To reduce model size and network training time, and to improve the network's ability to distinguish between targets and background, as well as to broaden the network width and enhance its adaptability, the improvement process of the Mask R-CNN network model in step two is as follows:

[0066] MobileNetV3 is introduced as the backbone network of Mask R-CNN. As a lightweight and efficient model, MobileNetV3 eliminates the stacking of network layers and uses a combination of 1×1 PW and 3×3 DW to extract target features and fuse feature information, which greatly improves the efficiency of network parameter update and model training. At the same time, the structure of the traditional MobileNetV3 is modified. The fifth block structure is replaced with max pooling and 3×3 convolution operation, and the number of channels in the last two blocks is changed to 64 and 96 respectively to extract more low-level and edge features.

[0067] The Dilated-CBAM attention mechanism is introduced and embedded into the depthwise separable convolutional structure of the MobileNetV3 model. After being concatenated to the second pointwise convolution of the inverse residual module with a stride of 1, it can take into account features in both spatial and channel dimensions, improving the network's ability to distinguish between targets and background. The embedded module is named bneck_DC. In the Dilated-CBAM mechanism, the target features are first input into the channel attention module, where channel feature information is extracted through average pooling and max pooling. The obtained parameters are then stacked through a multilayer perceptron and activated by the sigmoid function to obtain the channel attention features M. c (F), Channel attention features M c (F) Calculate according to formula (1); then, input the features into the spatial attention module, aggregate spatial information through pooling operation, and obtain spatial attention features M using dilated convolutions with dilation rates of 4, 2 and 1 and sigmoid activation functions. s (F x Spatial attention feature M s (F x ) Calculate according to formula (2);

[0068]

[0069]

[0070] In the formula: σ(·) represents the sigmoid activation function, W0 and W1 represent the weights of the hidden layer and the output layer, respectively, and f4 3×3 f2 3×3 f1 3×3 These represent dilated convolutions with a kernel size of 3 and dilation rates of 4, 2, and 1, respectively.

[0071] The Inception structure is introduced and combined with the bneck structure with a stride of 2 to form a more efficient feature extraction module. The embedded module is named bneck_IN. The Inception structure has four branches. Features are input into the network structure of each branch for processing. Then, 3×3 and 5×5 DW convolutions are used for multi-scale information fusion. Spatial information pooling is achieved through Max Pooling. The concat operation is discarded. The features output from different branches are added together before output to broaden the network width and enhance the network adaptability.

[0072] To fully utilize the height information of the laser point cloud and the color information of the image, and to fuse their advantages to obtain a model with higher recognition accuracy, the fusion process of the view frustum point cloud and the laser point cloud in step three is as follows:

[0073] S31: Coordinate system calibration;

[0074] Before fusing the point cloud of the target frustum candidate region, the lidar coordinate system and the camera coordinate system are calibrated. The specific method is as follows:

[0075] (1) Transform the point (X) in the lidar coordinate system using rigid body transformation. w ,Y w Z w Transformed to camera coordinate system (X) c ,Y c Z c );

[0076] (2) The transformation between the camera coordinate system and the image coordinate system is achieved by using the central perspective projection method in the pinhole imaging principle;

[0077] (3) Using formula (3), the transformation from image coordinate system to pixel coordinate system is achieved through scaling and translation transformation, thereby completing the spatial alignment and registration of laser point cloud to camera image;

[0078]

[0079] In the formula: R and T represent rotation and translation matrices, f x and f y The image represents the camera's two-dimensional focal length, u0 and v0 represent the coordinates of the image origin, u and v represent the coordinates in the pixel coordinate system, and b represents the coordinates of the pixel coordinate system. x It is the offset between the two coordinate systems;

[0080] S32: View frustum generation;

[0081] The four corner points of the target bounding box are mapped to four rays in three-dimensional space through coordinate system transformation, thereby determining a tetrahedron with the camera optical center as a fixed point, and collecting all points in the view frustum to form a view frustum point cloud;

[0082] S33: Point cloud fusion:

[0083] Determine the positional relationship between the laser point cloud and the image target bounding box, and fuse the laser point cloud and the view frustum point cloud within the view frustum range.

[0084] To improve the segmentation effect of the fused point cloud, in the frustum generation process of step S32 in step three, the angle of the frustum is first adjusted. Since each extracted frustum point cloud has a different orientation in the camera coordinate system, in order to simplify the point cloud processing, the angle of the frustum needs to be adjusted to rotate the frustum to the center viewpoint so that the central axis of the frustum is orthogonal to the image plane, thereby realizing the transformation from the camera coordinate system to the frustum coordinate system. Then, linear segmentation of the frustum is performed. Considering that the point cloud data distribution within the frustum is uneven and equal-height segmentation has certain defects, it is necessary to perform linear segmentation of the frustum. Along the central axis of the frustum, the segmentation height is set to u, 2u, 3u, ..., ku from near to far. For these generated point cloud cluster sequences, the self-attention PointNet network is used to process them to extract the frustum features of each sequence, and the features of each sequence are combined to improve the segmentation effect of the fused point cloud.

[0085] To effectively enhance the model's attention to different features and increase the network's adaptability to features at different scales, the method for predicting the bounding box parameters of the target point cloud using a self-attention PointNet network and a 3D bounding box regression network in step four is as follows:

[0086] S41: Take the fused point cloud data extracted from the frustum region as input, and combine it with the one-hot classification vector of the two-dimensional object detection part to output the evaluation score of the category to which each point cloud belongs;

[0087] S42: Non-target point clouds such as background and cluttered point clouds are removed by masking operations, and the fused point cloud is segmented into instances using the self-attention Pointnet network;

[0088] S43: Utilizing a 3D bounding box regression network and a multilayer perceptron mechanism to predict the bounding box of the target point cloud. As a preferred method, the predicted parameters include center coordinates, length, width, height, and heading angle.

[0089] To better capture local features of the point cloud space, in step four (S42), the method for instance segmentation of the fused point cloud using the self-attention PointNet network is as follows:

[0090] Because the spatial local information in the PointNet network is not closely connected, its ability to detect small targets and complex scenes is limited, which is not conducive to accurate 3D detection. Therefore, a self-attention mechanism is introduced on the basis of the PointNet network to form a self-attention PointNet network, so as to achieve the purpose of better acquiring the spatial local features of the point cloud. In the self-attention mechanism, firstly, the input X is mapped to the query vector matrix Q, the key vector matrix K, and the value vector matrix V through three different linear transformations. Then, the original input is subjected to multiple attention operations without sharing parameters through formula (4) to calculate the output result head of the i-th attention head. i Then, the m heads are merged and spliced ​​using formula (5), and then the encoded result Y of the fused point cloud is obtained by formula (6) after linear mapping.

[0091] head i =Attention(QW Q ,KW K VW V (4);

[0092]

[0093] Y = MultiHead(X,X,X) (6);

[0094] In the formula, X is the fused point cloud set; W Q W K W V These are the weights for Q, K, and V, respectively; Attention(·) represents the attention module function for a single head. From this, we can observe that each time x is... t During encoding, other point cloud x s It also acts as a variable, affecting x. t The input result yields an encoded result that contains both local information and global features; where softmax(●) represents the normalization function, d k Indicates the input dimension; For feature concatenation, MultiHead(●) represents the multi-head attention function.

[0095] In this invention, by installing a lidar and camera on the rocker arm of a coal mining machine, it is possible to simultaneously acquire lidar point cloud and image data of coal and rock during the coal cutting process, thereby providing reliable data support for rapid and accurate identification of coal and rock. Using the Mask R-CNN network as the base model and MobileNetV3 as the backbone network, the Dilated-CBAM mechanism and Inception structure are integrated into the MobileNetV3 backbone network. This not only reduces the model size and network training time but also effectively increases the network's adaptability to features at different scales, helping to better extract the underlying and edge features of the target, thus significantly improving the model's detection performance. First, the improved Mask R-CNN network model is trained using image data acquired by the camera to calculate and output the target bounding box position of the coal and rock in the image. Then, based on the projection relationship between the lidar and the camera, a view frustum point cloud of the target bounding box is generated, and the generated view frustum point cloud is fused with the lidar point cloud. This effectively reduces the computational scale and spatial search range of the point cloud, which is beneficial for the point cloud network to better identify the target. Finally, a self-attention mechanism is added to the PointNet network to construct a self-attention mechanism. The PointNet network, further enhanced by a self-attention PointNet network, segments the fused point cloud within the view frustum. A 3D bounding box regression network then predicts the bounding box parameters of the target point cloud, effectively acquiring local spatial feature information of the point cloud. This improves the network's ability to detect small objects and complex scenes, facilitating accurate 3D detection. This invention, based on LiDAR and camera fusion technology, proposes a multimodal view frustum point cloud coal and rock recognition method. It fully utilizes the height information of the LiDAR point cloud and the color information of the image, combining their advantages to solve the problem of low coal and rock recognition accuracy caused by harsh environments such as high coal dust, low light, and high noise levels at the working face. This provides technical support for intelligent mining of coal mining machines and is of great significance for achieving "unmanned" or "reduced-manpower" mining in fully mechanized mining faces.

Claims

1. A coal and rock identification method based on multimodal view cone point cloud fusion, characterized in that, Specifically, the following steps are included: Step 1: Before the coal mining machine starts cutting the coal and rock test block, install a lidar and a camera on the rocker arm of the coal mining machine, and establish communication connections between the lidar and the data processing equipment, and between the camera and the data processing equipment. According to the sequence of fully mechanized mining operations, perform the cutting operation on the first coal and rock test block, and at the same time use the lidar and camera to collect the laser point cloud and image data of the coal and rock in real time during the cutting process. Step 2: Using the Mask R-CNN network as the base model, select MobileNetV3 as the backbone network, and integrate the Dilated-CBAM mechanism and Inception structure into the MobileNetV3 backbone network to obtain an improved Mask R-CNN network model; use the image data acquired by the camera to train the improved Mask R-CNN network model, calculate and output the target bounding box position of coal and rock in the image; The improvement process for the Mask R-CNN network model is as follows: MobileNetV3 is introduced as the backbone network of Mask R-CNN. At the same time, the traditional MobileNetV3 structure is modified by replacing the fifth block structure with max pooling and 3×3 convolution operations, and changing the number of channels in the last two blocks to 64 and 96 respectively, in order to extract more low-level and edge features. The Dilated-CBAM attention mechanism is introduced and embedded into the depthwise separable convolutional structure of the MobileNetV3 model. After being concatenated to the second pointwise convolution of the inverse residual module with a stride of 1, it takes into account both spatial and channel dimensions of features, improving the network's ability to distinguish between targets and background. The embedded module is named bneck_DC. In the Dilated-CBAM mechanism, target features are first input into the channel attention module, where channel feature information is extracted through average pooling and max pooling. The obtained parameters are then stacked through a multilayer perceptron and activated by the sigmoid function to obtain channel attention features. Channel attention features The calculation is performed according to formula (1); then, the features are input into the spatial attention module, spatial information is aggregated through pooling operations, and spatial attention features are obtained by dilated convolutions with dilation rates of 4, 2 and 1 and sigmoid activation functions. Spatial attention features Calculate according to formula (2); (1); (2); In the formula: This represents the sigmoid activation function, where W0 and W1 represent the weights of the hidden and output layers, respectively. These represent dilated convolutions with a kernel size of 3 and dilation rates of 4, 2, and 1, respectively. An Inception structure is introduced and combined with a bneck structure with a stride of 2 to form a more efficient feature extraction module. The embedded module is named bneck_IN. The Inception structure has four branches; features are input into each branch network structure for processing, and then... , The DW convolution is used to perform multi-scale information fusion. Spatial information pooling is achieved through Max Pooling. The concat operation is abandoned. The features output by different branches are added together before output to broaden the network width and enhance the network adaptability. Step 3: Generate the view frustum point cloud of the target bounding box based on the projection relationship between the lidar and the camera, and fuse the generated view frustum point cloud with the lidar point cloud; Step 4: Use the self-attention PointNet network to segment the fused point cloud within the view frustum, and use the 3D bounding box regression network to predict the bounding box parameters of the target point cloud; Step 5: According to the sequence of fully mechanized mining operations, the remaining coal and rock test blocks are cut in sequence. Starting from the second coal and rock test block, the collected laser point cloud and image data are processed through Step 2 and Step 3 respectively to obtain the target's frustum fusion point cloud. Then, it is input into the self-attention Pointnet network and the 3D bounding box regression network to predict the shape and position information of coal and rock during the cutting process, so as to achieve rapid and accurate identification of coal and rock.

2. The coal and rock identification method based on multimodal view cone point cloud fusion according to claim 1, characterized in that, In step four, the bounding box parameters of the target point cloud include the target's center coordinates, length, width, height, and heading angle.

3. A coal and rock identification method based on multimodal view cone point cloud fusion according to claim 1 or 2, characterized in that, In step two, the method for outputting the target bounding box positions of coal and rock in the image using the improved Mask R-CNN network model is as follows: S21: Input the image data acquired by the camera into the improved Mask R-CNN network model, and use the improved MobileNetV3 backbone network and feature pyramid network to extract multi-scale feature maps of the input image; S22: Use the Region Proposal Network to extract candidate region boxes that may contain the target on the feature map, and use the cropping and alignment operations of the ROI Align layer to normalize the candidate region boxes; S23: Perform regression and classification operations on the target to obtain the target's bounding box and confidence score.

4. The coal and rock identification method based on multimodal view cone point cloud fusion according to claim 3, characterized in that, In step three, the fusion process of the view frustum point cloud and the laser point cloud is as follows: S31: Coordinate system calibration; Before fusing the point cloud of the target frustum candidate region, the lidar coordinate system and the camera coordinate system are calibrated. The specific method is as follows: (1) By rigid body transformation, the points in the lidar coordinate system ( Transformed to camera coordinate system ); (2) The transformation between the camera coordinate system and the image coordinate system is achieved by using the central perspective projection method in the pinhole imaging principle; (3) Using formula (3), the transformation from image coordinate system to pixel coordinate system is achieved through scaling and translation transformation, thereby completing the spatial alignment and registration of laser point cloud to camera image; (3); In the formula: R and T represent rotation and translation matrices, f x and f y The image represents the camera's two-dimensional focal length, u0 and v0 represent the coordinates of the image origin, u and v represent the coordinates in the pixel coordinate system, and b represents the coordinates of the pixel coordinate system. x It is the offset between the two coordinate systems; S32: View frustum generation; The four corner points of the target bounding box are mapped to four rays in three-dimensional space through coordinate system transformation, thereby determining a tetrahedron with the camera optical center as a fixed point, and collecting all points in the view frustum to form a view frustum point cloud; S33: Point cloud fusion: Determine the positional relationship between the laser point cloud and the image target bounding box, and fuse the laser point cloud and the view frustum point cloud within the view frustum range.

5. The coal and rock identification method based on multimodal view cone point cloud fusion according to claim 4, characterized in that, In step S32, during the generation of the view frustum, the angle of the view frustum is first adjusted to rotate it to the center viewpoint, making the central axis of the view frustum orthogonal to the image plane, thus realizing the transformation from the camera coordinate system to the view frustum coordinate system. Then, linear segmentation of the view frustum is performed. Along the central axis of the view frustum, the segmentation height is set to u, 2u, 3u, ..., ku from near to far. For these generated point cloud cluster sequences, the self-attention PointNet network is used to process them to extract the view frustum features of each sequence. The features of each sequence are then combined to improve the segmentation effect of the fused point cloud.

6. The coal and rock identification method based on multimodal view cone point cloud fusion according to claim 5, characterized in that, In step four, the method for predicting the bounding box parameters of the target point cloud using the self-attention PointNet network and the 3D bounding box regression network is as follows: S41: Take the fused point cloud data extracted from the frustum region as input, and combine it with the one-hot classification vector of the two-dimensional object detection part to output the evaluation score of the category to which each point cloud belongs; S42: Non-target point clouds are removed by masking operations, and instance segmentation is performed on the fused point cloud using a self-attention PointNet network; S43: Utilize a 3D bounding box regression network and multilayer perception mechanism to achieve bounding box regression prediction of target point clouds.

7. The coal and rock identification method based on multimodal view cone point cloud fusion according to claim 6, characterized in that, In step four, S42, the method for instance segmentation of the fused point cloud using the self-attention PointNet network is as follows: A self-attention mechanism is introduced on the basis of the PointNet network to construct a self-attention PointNet network. First, the input X is mapped to the query vector matrix Q, the key vector matrix K, and the value vector matrix V through three different linear transformations. Then, the original input is subjected to multiple non-shared parameter attention operations through formula (4) to calculate the output result head of the i-th attention head. i Then, the m heads are merged and spliced ​​using formula (5), and then the encoding result Y of the fused point cloud is obtained by formula (6) after linear mapping. (4); (5); (6); In the formula, X is the fused point cloud set; W Q W K W V The weights for Q, K, and V are respectively; Attention ( ● ) represents the attention module function for a single head. , where softmax ( ● () represents the normalization function, x t x s All are point clouds, d k Indicates the input dimension; For feature splicing, MultiHead ( ● ) represents the multi-head attention function.