Coal mine conveyor belt foreign matter detection method based on deep learning

Through multi-view data enhancement, contextual background feature fusion module, conveyor belt area loss and variable focus loss, the robustness of viewing angle changes and external background interference in foreign matter detection in coal mine conveyor belts is solved, and high-precision and high-safe detection effects are achieved.

CN120182210APending Publication Date: 2025-06-20SHANXI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510257299.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of robustness of viewing angle changes and external background interference in foreign matter detection on coal mine conveyor belts, resulting in the impact of detection accuracy and safety.

Method used

Multi-view data augmentation (MVDA) technology is used to simulate viewing angle changes, and the design of the contextual background feature fusion module (CFP) integrates surrounding coal characteristics to reduce false detection of foreign objects outside the conveyor belt, and introduces conveyor belt area loss (CBAL) and variable focus loss (VFL) to improve detection accuracy.

Benefits of technology

It significantly improves detection capability and multi-lens adaptability, reduces false detection outside the conveyor belt, improves detection accuracy in the conveyor belt area, and enhances attention to difficult samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182210A_ABST
    Figure CN120182210A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of coal mine intelligent monitoring, and particularly relates to a coal mine conveyor belt foreign matter detection method based on deep learning. In order to overcome background interference outside a conveyor belt and highlight feature information of foreign matters, multi-view data enhancement (MVDA) is applied to simulate view angle change and learn more view angle robustness features, a context background feature fusion module (CFP) is further designed, features of surrounding coal are integrated into potential foreign matters, and wrong detection of the foreign matters outside the conveyor belt is reduced. Moreover, a conveyor belt area loss (CBAL) is designed and focuses on the conveyor belt area, and interference from complex backgrounds outside the conveyor belt is reduced. And finally, introducing variable focus loss (VFL) to enhance the attention to the difficult sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent monitoring in coal mines, and particularly relates to a method for detecting foreign objects on a coal mine conveyor belt based on deep learning. Background Art

[0002] With the acceleration of the automation and intelligentization process in the coal mining industry, safe production and efficient operation have become one of the important goals of modern coal mines. As a core device in the coal production and transfer process, the stable operation of the coal conveyor belt is directly related to the production efficiency of the coal mine. In actual production, foreign objects such as anchor bolts, gangue, and large coal lumps often appear on the conveyor belt. The existence of these foreign objects not only increases the difficulty of coal cleaning and processing, but also often causes damage and breakage of the conveyor belt, and even leads to accidents such as personnel safety. Therefore, the removal of foreign objects on the conveyor belt, as an important link to ensure the safe operation of the equipment, has received extensive attention.

[0003] In order to reduce the impact of foreign objects on the conveyor belt, coal mines have adopted various technical means to clean and screen the sundries on the conveyor belt. First, manual inspection relies on manual monitoring and cleaning, but it is limited by high labor costs, low recognition efficiency, and is greatly affected by environmental conditions, and is prone to safety accidents. In addition, coal mines also use compound coal-gangue separation equipment to screen and filter the coal gangue and foreign objects mixed in the conveyor belt, but it often causes congestion at the connection of the conveyor belt due to foreign objects or large coal lumps, and even causes problems such as tearing and damage of the conveyor belt at the connection, making it difficult to completely eliminate the harm of foreign objects on the conveyor belt.

[0004] In recent years, object detection methods based on deep neural networks have made remarkable progress in various fields. Deep learning models can automatically learn the features in data and are particularly suitable for object detection tasks in complex environments. Existing object detection networks such as YOLO (You Only Look Once), Faster R-CNN, etc. have performed well in object detection tasks in many fields. However, when directly applied to the detection of foreign objects on coal mine conveyor belts, due to the background differences in camera images under different working environments, as well as the influence of noises such as light, dust, or water vapor, the model trained on one camera's data is difficult to be applied to other cameras. In addition, it is also easily affected by the external background of the conveyor belt and produces incorrect predictions. Therefore, for the detection of foreign objects on coal mine conveyor belts, how to overcome the interference of the background outside the conveyor belt and highlight the feature information of foreign objects is also an urgent problem for us to solve. Summary of the Invention

[0005] The present invention provides a foreign object detection method for coal mine conveyor belts based on deep learning. This method applies multi-view data augmentation (MVDA) to simulate perspective changes, enabling the learning of more perspective-robust features. It also designs a context background feature fusion module (CFP) to integrate the features of the surrounding coal into potential foreign objects, reducing false detections of foreign objects outside the conveyor belt. Additionally, a conveyor belt area loss (CBAL) is designed to focus on the conveyor belt area and reduce interference from complex backgrounds outside the conveyor belt. Finally, a variable focal loss (VFL) is introduced to enhance the attention to difficult samples.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A foreign object detection method for coal mine conveyor belts based on deep learning, comprising the following steps:

[0008] Step 1, collect foreign object images of the coal mine conveyor belt and perform multi-view data augmentation on them to generate enhanced images;

[0009] Step 2, input the enhanced images into a pre-trained foreign object detection model for coal mine conveyor belts, and output the positions and categories of the predicted targets; the foreign object detection model for coal mine conveyor belts includes a backbone network, a query network, and a classification and regression network; the backbone network uses the ResNet50 architecture to process the enhanced images, extract multi-scale features, and generate multi-scale feature maps; the query network includes an encoder, a decoder, and a target-background feature fusion module. The encoder processes the multi-scale feature maps from the backbone network through a self-attention mechanism to capture the global context information in the image and obtain enhanced feature maps. The decoder is responsible for generating and using learnable query vectors to extract target features from the enhanced feature maps. The context background feature fusion module extracts the background features around the target and combines them with the extracted target features; the classification and regression network consists of a classification module and a regression module, which are respectively responsible for the category recognition of the target and the prediction of the bounding box position.

[0010] Further, the multi-view data augmentation in step 1 is specifically as follows:

[0011] Step 1.1, define four transformation matrices, namely the perspective transformation matrix P, the scaling matrix S, the rotation matrix R, and the translation matrix T, and the expressions are as follows:

[0012]

[0013] where P x and P y are parameters randomly generated in the range of [-0.001, 0.001]; S x and S y respectively represent the shear strengths along the corresponding axes, Sx , S y is in the range of [-10°, 10°]; θ represents the image rotation angle, and θ is in the range of [-15°, 15°]; T x , T y represents the translation distances in the x and y directions, and T x is in the range of [-0.1w, 0.1w], and T y is in the range of [-0.1h, 0.1h], and is randomly generated within 10% of the image width w and height h;

[0014] Step 1.2, combine these transformation matrices to form the total transformation matrix A aug , as follows:

[0015] A aug = P × S × R × T

[0016] X' = A aug × X

[0017] B' = A aug × B

[0018] where X' and B' are the transformed image and the corresponding bounding box, and A aug × X applies an affine transformation to each pixel coordinate of the image, and A aug × B transforms the coordinate box of the true target corresponding to the image.

[0019] Furthermore, the backbone network in step 2 is specifically:

[0020] The enhanced image extracts 4 feature maps of different scales through ResNet50, gradually screening out information at different levels in the image. The first three layers come from layer3, layer4, and layer5 of the ResNet network, with downsampling rates of 8, 16, and 32 respectively, and then a 1*1 convolution is used to unify the feature dimensions;

[0021] By performing a 3*3 convolution on the features of layer3, a feature with a downsampling rate of 64 and 256 dimensions is obtained. Among them, the feature map of the l-th layer represents a feature map with c l channels, a height of H l , and a width of W l .

[0022] Furthermore, the encoder in step 2 is specifically:

[0023] Input the multi-scale feature maps from the backbone network, map all the feature maps to the same number of channels, and introduce positional encoding. Generate learnable spatial coordinate embeddings for each position in the multi-scale feature maps through sine-cosine positional encoding, and add them to the features at each position in the feature maps;

[0024] For the features at each position in the multi-scale feature maps, generate query vectors, key vectors, and value vectors through linear transformation. At the same time, generate a set of dynamic offsets based on the reference points, which are initialized as grid points with a uniform distribution. Predict the offsets and attention weights of each reference point from the concatenation of the query vectors and the positional encoding through the deformable attention mechanism, and perform weighted summation of the attention weights and the offset feature values. The calculation formula is as follows:

[0025]

[0026] where, x ∈ R C×H×W is the given input feature map, z q is the feature of each query vector in the encoder, P q is the reference point of each query vector in the encoder, M represents the number of attention heads, Δp mqk is the learnable offset, A mqk is the attention weight, W m is the projection matrix of the output of the m-th head of the multi-head attention, W‘ m x(·) is to bilinearly interpolate the features at the offset coordinates and map them to the feature space of the m-th attention head through the projection matrix W m ;

[0027] After completing the attention calculation, pass through a residual connection and a normalization layer, and a feed-forward network. Use the output of the feed-forward network and the output of the previous residual connection and normalization layer as the input of the next residual connection and normalization layer;

[0028] Repeat the above calculation L times, and use the output each time as the input for the next time. Continuously optimize the feature representation in this iterative manner, and finally obtain the multi-scale feature map z final .

[0029] Furthermore, the decoder in step 2 is specifically:

[0030] Generate learnable query vectors q init and the enhanced features z final of the encoder through random initialization as the input. Each query vector embeds spatial coordinate information through positional encoding to generate an initial reference point as the target candidate position for the decoder to focus on;

[0031] Perform self-attention operation on the randomly initialized query vectors q initSuppress redundant candidate bounding boxes by learning the dependencies between interactive learning objectives, and use residual connections and layer normalization;

[0032] Perform deformable attention operation, associate the query vector after residual connection and layer normalization with the multi-scale feature map z output by the encoder final to dynamically correct the offset of the reference point, and upsample on the multi-scale feature map z final to extract features of different resolutions. The feature values at each sampling point are multiplied by their corresponding weights and then summed. The calculation formula is:

[0033]

[0034] where z q represents the feature of each query vector in the decoder, which is generated by the multi-scale feature z final output by the encoder; represents the reference point of each query vector in the decoder, which is used to dynamically generate the spatial reference coordinates of the target and is initialized as a randomly generated target candidate position; is a set of multi-scale feature maps; represents the normalized coordinate space that maps the reference point to the l-th layer feature map, and Δp mlqk represents the learnable offset of the m-th attention head, the l-th layer feature map, and the k-th sampling point. A mlqk represents the attention weight, that is, the importance of the k-th sampling point in the m-th head and the l-th layer. W‘ m x l (·) means extracting features by bilinear interpolation of the offset coordinates on the l-th layer feature map and mapping them to the feature space of the m-th attention head through the projection matrix W m ;

[0035] After completing the attention calculation, stabilize the gradient propagation through two residual connections and layer normalization, and further extract high-order features through a feed-forward network;

[0036] After L-layer iterative calculation, finally obtain the feature F containing the predicted target position and category information target , which is used for subsequent object classification and regression.

[0037] Furthermore, the context background feature fusion module in step 2 is specifically:

[0038] Use the decoder to generate the feature F containing the predicted target position and category information target to generate the target bounding box B target =(x center , Y center , w, h), and expand it to the enlarged bounding box B expand , and the calculation formula is as follows:

[0039] w expanded = w · scale

[0040] h expanded = h · scale

[0041] B expand = (x center , y center , w expanded , h expanded )

[0042] where w expand and h expand are the width and height;

[0043] Based on B target and B expand , two masks M target and M expand matching their sizes are generated. Through a subtraction operation, the surrounding mask M context is obtained. The calculation formula is as follows:

[0044] M context = M expand - M target ;

[0045] Multiply the mask M context element-wise with the feature map Z final and adjust it to the same dimension as F target through a max pooling operation. The calculation formula is as follows:

[0046] F context = Pooling(Z final ⊙ M context );

[0047] Combine the feature F target with the feature F context . The calculation formula is as follows:

[0048] F integrate = F target + F context

[0049] The potential target feature F integrate containing context information is input to the classification module for classification tasks.

[0050] Furthermore, the classification module in step 2 is specifically:

[0051] Use the potential target feature F integrate output by the context background feature fusion module and containing context information, a high-dimensional feature is mapped to a category space through a fully connected layer, and the output values of each category are normalized using the softmax function to generate the category probability corresponding to each target.

[0052] Further, the regression module in step 2 is specifically as follows:

[0053] Use the feature F containing the predicted target position and category information output by the decoder target , map the feature to a 4D space through a multi-layer perceptron, corresponding to the center point and width and height of the bounding box respectively, and finally output a tensor of dimension (N,4) representing the bounding box coordinates of each candidate target.

[0054] Furthermore, it also includes the training of the foreign object detection model for the coal mine conveyor belt, and uses the conveyor belt area loss, variable focus loss, position regression loss, and category classification loss to optimize the model parameters, specifically as follows:

[0055] The conveyor belt area loss:

[0056] Manually annotate the conveyor belt area Poly from different camera perspectives, expressed as:

[0057] Poly = [(x1,y1),(x2,y2),…,(x n ,y n )]

[0058] where, (x1,y1),(x2,y2),…,(x n ,y n ) represent the vertices of the conveyor belt polygon;

[0059] Evaluate whether the center point of the predicted bounding box is within the conveyor belt area, that is, calculate the number of intersection points Num of the ray emitted horizontally from the center Pred of each predicted bounding box j =(x c ,y c ) with the polygon boundary, and the calculation formula is as follows: points where, sign represents the function to judge the positive and negative of the value. When Num

[0060]

[0061] is odd, the prediction box is within the conveyor belt area, and when Num points is even, the prediction box is outside the conveyor belt area; points Calculate the loss of the prediction box located outside the conveyor belt area, and the calculation formula is as follows:

[0062] Outside(Pred

[0063] j j, Poly) = 1 - (Num points mod 2)

[0064]

[0065] where m is the total number of prediction boxes, and mod 2 is the modulo operation;

[0066] The variable focus loss:

[0067] By dynamically adjusting the weights of positive and negative samples and the focus factor, false alarms are reduced, and the calculation formula is as follows:

[0068]

[0069] where p represents the predicted IoU score, q represents the target score, α is the weight balance factor for positive and negative samples, and γ is the focus factor;

[0070] The overall loss function:

[0071] Loss = L cls + L reg + L cbal + L vfl

[0072] where L cls represents the class classification loss, which is calculated by the cross - entropy loss function, and L reg represents the location regression loss, including L1 loss and generalized IoU loss.

[0073] Compared with the prior art, the present invention has the following advantages:

[0074] By introducing multi - view data augmentation (MVDA), context background feature fusion module (CFP), conveyor belt area loss (CBAL), and variable focus loss (VFL), the present invention significantly improves the detection ability and multi - camera adaptation performance. Specifically, through multi - view data augmentation, foreign object features of more views of the conveyor belt can be learned; the context background feature fusion module enhances context association by integrating features around foreign objects, thereby reducing false detections outside the conveyor belt; the conveyor belt area loss function further focuses on the conveyor belt area through loss drive, further reducing false detections outside the conveyor belt area and improving the detection accuracy of the conveyor belt area; introducing variable focus loss enhances the model's attention to difficult samples. Finally, through a large number of experiments, the effectiveness of the present invention is verified, showing excellent performance in foreign object detection and multi - camera adaptability, having significant industrial application potential, especially in improving coal mine production safety and efficiency, demonstrating great application value. Brief Description of the Drawings

[0075] Figure 1The multi-view data augmentation of the present invention simulates the perspective changes of different cameras;

[0076] Figure 2 It is a schematic diagram of the foreign object detection model for the coal mine conveyor belt of the present invention;

[0077] Figure 3 It is a schematic diagram of the context background feature fusion module of the present invention;

[0078] Figure 4 It is the error type of the present invention;

[0079] Figure 5 It is the visualization of the feature map of the present invention;

[0080] Figure 6 It is the change curve of the CBAL value of the present invention. Specific implementation manners

[0081] To further elaborate on the technical solution of the present invention, the present invention will be further described below through embodiments.

[0082] A method for detecting foreign objects on a coal mine conveyor belt based on deep learning in this embodiment includes the following steps:

[0083] Step 1, collect images of foreign objects on the coal mine conveyor belt and perform multi-view data augmentation on them to generate enhanced images;

[0084] Due to the difficulty of collecting foreign object data on the coal mine conveyor belt, in order to enable the model to learn the foreign object features from multiple perspectives, we apply various geometric transformations to the images to simulate the camera perspectives that may appear in reality, so as to enhance the generalization ability of the model at different camera angles, as Figure 1 shown.

[0085] First, define four transformation matrices, namely the perspective transformation matrix P, the scaling matrix S, the rotation matrix R, and the translation matrix T, and the expressions are as follows:

[0086]

[0087] Among them, P x and P y are parameters randomly generated within the range of [-0.001, 0.001]; S x , S y respectively represent the shear strengths along the corresponding axes, and S x , S y are within the range of [-10°, 10°]; θ represents the image rotation angle, and θ is within the range of [-15°, 15°]; T x , T y represent the translation distances in the x and y directions, and T xWithin the range of [-0.1w, 0.1w], T y Within the range of [-0.1h, 0.1h], it is randomly generated within 10% of the image width w and height h;

[0088] Secondly, these transformation matrices are combined to form the total transformation matrix A aug , as follows:

[0089] A aug = P × S × R × T

[0090] X' = A aug × X

[0091] B' = A aug × B

[0092] where X' and B' are the transformed image and the corresponding bounding box, A aug × X applies an affine transformation to each pixel coordinate of the image, rather than performing a direct matrix multiplication operation on the entire image, and A aug × B transforms the coordinate box of the true target corresponding to the image;

[0093] Finally, by applying the above transformation to each image, the enhanced image is obtained;

[0094] Step 2, input the enhanced image into a pre-trained foreign object detection model for coal mine conveyor belts, and output the position and category of the predicted target; the foreign object detection model for coal mine conveyor belts (as Figure 2 shown) includes a backbone network, a query network, and a classification regression network; the backbone network uses the ResNet50 architecture to process the enhanced image, extracts multi-scale features, and generates multi-scale feature maps; the query network includes an encoder, a decoder, and a target-background feature fusion module. The encoder processes the multi-scale feature maps from the backbone network through a self-attention mechanism to capture the global context information in the image and obtain enhanced feature maps. The decoder is responsible for generating and using learnable query vectors to extract target features from the enhanced feature maps. The context-background feature fusion module combines the background features around the target by extracting them and combines them with the extracted target features; the classification regression network consists of a classification module and a regression module, which are responsible for the category recognition of the target and the prediction of the bounding box position respectively.

[0095] 1) Backbone network

[0096] Extract useful information from the input image, and extract features from the data of foreign objects on the conveyor belt under multiple cameras.

[0097] The enhanced image extracts four feature maps of different scales through ResNet50, gradually screening out information at different levels in the image. The first three come from layer3, layer4, and layer5 of the ResNet network, with downsampling rates of 8, 16, and 32 respectively. Then, a 1*1 convolution is used to unify the feature dimensions for each of them; by performing a 3*3 convolution on the features of layer3, a feature with a downsampling rate of 64 and 256 dimensions is obtained. Among them, the feature map of the l-th layer denotes having c l channels, a height of H l , and a width of W l feature map.

[0098] 2) Query network

[0099] The input is the multi-scale feature maps extracted from the backbone network. By generating and using learnable query vectors, and through self-attention and cross-attention calculations, target information is extracted. This network includes an encoder, a decoder, and a context background feature fusion module;

[0100] 2.1) Encoder

[0101] Process the multi-scale feature maps from the backbone network through the self-attention mechanism to capture the global context information in the image.

[0102] The input is the multi-scale feature maps from the backbone network. Since feature maps of different scales may have different numbers of channels (i.e., feature dimensions, 256 dimensions), for unified processing, all feature maps are mapped to the same number of channels. To retain the spatial position information in the feature maps, these feature maps are flattened, and position encoding (Pos embedding ) is introduced. Through sine-cosine position encoding, learnable spatial coordinate embeddings are generated for each position in the multi-scale feature maps and added to the features at each position in the feature maps;

[0103] For the features at each position in the multi-scale feature maps, query vectors (Q), key vectors (K), and value vectors (V) are generated through linear transformation. At the same time, a set of dynamic offsets are generated based on the reference point (R points ). The reference point is initialized as grid points with a uniform distribution. Through the deformable attention mechanism, the offset (Δp) and attention weight (α) of each reference point are predicted from the concatenation of the query vector and the position encoding. The attention weight is weighted and summed with the offset feature values, and the calculation formula is as follows:

[0104]

[0105] where x ∈ R C×H×W is the given input feature map, zq is the feature of each query vector in the encoder, P q is the reference point of each query vector in the encoder, M represents the number of attention heads, each head is used to independently learn the features of different patterns, and each head m weights and sums K << (H × W) sampling points instead of calculating all positions, which significantly reduces the computational amount, Δp mqk is a learnable offset, A mqk is the attention weight, W m is the projection matrix of the output of the m-th multi-head attention, W‘ m x(·) extracts features by bilinear interpolation of the offset coordinates and maps them to the feature space of the m-th attention head through the projection matrix W m ;

[0106] After completing the attention calculation, it passes through a residual connection and normalization layer (Add&Norm), a feed-forward network (FFN), and uses the output of the feed-forward network and the output of the previous residual connection and normalization layer as the input of the next residual connection and normalization layer;

[0107] Repeat the above calculation L (L = 6) times, and each time use the output as the input of the next time. Through this iterative method, continuously optimize the feature representation, and finally obtain the multi-scale feature map z output by the encoder final .

[0108] 2.2) Decoder

[0109] Responsible for generating and using learnable query vectors to extract target information from the enhanced features output by the encoder. mainly through the cross-attention mechanism, enabling the query vectors to find and focus on the target area in the feature map.

[0110] In order to further mine the target position and category information based on the features extracted by the encoder, the decoder randomly initializes and generates a learnable query vector q init and the enhanced features z of the encoder final as inputs. Each query vector embeds spatial coordinate information through positional encoding to generate an initial reference point, which serves as a candidate position for the decoder to focus on the target;

[0111] Perform self-attention operation. The randomly initialized generated query vectors q init interact with each other to learn the dependencies between targets, suppress redundant candidate boxes, and stabilize the gradient propagation through residual connection and layer normalization (Add&Norm);

[0112] Perform deformable attention operation. The query vectors after residual connection and layer normalization are combined with the multi-scale feature map z output by the encoder finalAssociate, dynamically correct the offset of the reference point, and perform upsampling on the multi-scale feature map z final Extract features of different resolutions. After multiplying the feature values of each sampling point by their corresponding weights and summing them, the calculation formula is as follows:

[0113]

[0114] Among them, z q represents the feature of each query vector in the decoder, which is generated by the multi-scale feature z final output by the encoder, represents the reference point of each query vector in the decoder, which is used to dynamically generate the spatial reference coordinates of the target and is initialized as a randomly generated target candidate position, is a set of multi-scale feature maps, represents mapping the reference point to the normalized coordinate space of the l-th layer feature map, Δp mlqk represents the learnable offset of the m-th attention head, the l-th layer feature map, and the k-th sampling point, A mlqk represents the attention weight, that is, it represents the importance of the k-th sampling point in the m-th head and the l-th layer, W‘ m x l (·) means that on the l-th layer feature map, bilinearly interpolate the offset coordinates to extract features and map them to the feature space of the m-th attention head through the projection matrix W m ;

[0115] After completing the attention calculation, stabilize the gradient propagation through two residual connections and layer normalization (Add&Norm), and further extract high-order features through the feed-forward network (FFN);

[0116] After L (L = 6) layers of iterative calculations, finally obtain the feature F target containing the predicted target position and class information, which is used for subsequent target classification and regression.

[0117] 2.3) Context background feature fusion module

[0118] By extracting the background features around the target and fusing them with the target features captured by the query vector output by the decoder, the model can learn the association relationship between the target features and their surrounding features.

[0119] Use the feature F target generated by the decoder containing the predicted target position and class information to generate the target bounding box B target =(x center ,Y center ,w,h), and expand it to the enlarged bounding box B expand , and the calculation formula is as follows:

[0120] w expanded = w·scale

[0121] h expanded = h·scale

[0122] B expand = (x center , y center , w expanded , h expanded )

[0123] where w expand and h expand are the width and height;

[0124] Based on B target and B expand , two masks M target and M expand matching their sizes are generated. Through a subtraction operation, the surrounding mask M context is obtained. The calculation formula is as follows:

[0125] M context = M expand - M target ;

[0126] Multiply the mask M context element-wise with the feature map Z final and adjust it to the same dimension as F target through a max pooling operation. The calculation formula is as follows:

[0127] F context = Pooling(Z final ⊙ M context );

[0128] Combine the feature F target with the feature F context . The calculation formula is as follows:

[0129] F integrate = F target + F context

[0130] The potential target feature F integrate containing context information is input to the classification module for the classification task.

[0131] 3) Classification regression network

[0132] Perform class prediction and bounding box regression for each query vector. It mainly consists of a classification module and a regression module, which are responsible for the class recognition of the target and the prediction of the bounding box position respectively.

[0133] 3.1) Classification module

[0134] Use the potential target feature F containing context information output by the context background feature fusion module integrate , map the high-dimensional feature to the category space through a fully connected layer, and use the softmax function to normalize the output values of each category to generate the category probability corresponding to each target.

[0135] 3.2) Regression module

[0136] Use the feature F containing predicted target position and category information output by the decoder target , map the feature to a 4D space through a multi-layer perceptron (MLP) composed of several fully connected layers and non-linear activation functions, corresponding to the center point and width and height of the bounding box respectively, and finally output a tensor of dimension (N,4) representing the bounding box coordinates of each candidate target.

[0137] In addition, this embodiment also includes the training of the foreign object detection model for the coal mine conveyor belt, and uses the conveyor belt area loss, variable focus loss, position regression loss, and category classification loss to optimize the model parameters, as follows:

[0138] Conveyor belt area loss (CBAL):

[0139] To prevent foreign objects outside the conveyor belt area from being misdetected, a conveyor belt area loss (CBAL) mechanism is introduced. This mechanism adopts a penalty mechanism to prompt the model to focus on the targets within the conveyor belt area, thereby improving its generalization performance.

[0140] Manually annotate the conveyor belt area Poly from different camera perspectives, expressed as:[[]]

[0141] Poly = [(x1,y1),(x2,y2),…,(x n ,y n )]

[0142] where, (x1,y1),(x2,y2),…,(x n ,y n ) represent the vertices of the conveyor belt polygon;

[0143] Evaluate whether the center point of the predicted bounding box is within the conveyor belt area, that is, calculate the number of intersection points Num of the ray emitted horizontally from the center Pred of each predicted bounding box j =(x c ,y c ) with the polygon boundary, and the calculation formula is as follows:[[]] points The calculation formula is as follows:[[]]

[0144]

[0145] Among them, sign represents a function for judging the positive and negative of a numerical value. When Num points is odd, the prediction box is located within the conveyor belt area. When Num points is even, the prediction box is located outside the conveyor belt area;

[0146] Calculate the loss of the prediction boxes located outside the conveyor belt area. The calculation formula is as follows:

[0147] Outside(Pred j , Poly) = 1 - (Num points mod 2)

[0148]

[0149] Among them, m is the total number of prediction boxes, and mod 2 is the modulo operation. By taking the average of the number of predictions outside the conveyor belt during the training process, the conveyor belt area loss value is obtained, which is used to drive the model to focus on the conveyor belt area during the learning process. When all prediction boxes are correctly located within the conveyor belt area, the loss value gradually becomes zero, indicating the best model performance. In the analysis of error types and loss curves in the experimental part, it is proved that this model effectively reduces the misdetection of the background outside the conveyor belt area and improves the detection accuracy and generalization ability of the model.

[0150] Variable Focus Loss (VFL):

[0151] By dynamically adjusting the weights of positive and negative samples and the focus factor, false alarms are reduced, and the calculation formula is as follows:

[0152]

[0153] Among them, p represents the predicted IoU score, q represents the target score, α is the weight balance factor for positive and negative samples, and γ is the focus factor. For foreground points, q is defined as the IoU between the predicted bounding box and the ground truth box, while for background points, q is set to 0. α is the weight balance factor for negative samples, and γ is the focus factor, which is used to adjust the emphasis on samples with different confidences. VFL reduces the impact of negative samples with q = 0 through the scaling factor p γ and α is used to balance the contributions of positive and negative samples and fine-tune their relative importance. For positive samples with q > 0, their contributions remain intact and are weighted according to the target q. Higher true IoU values will increase their influence, ensuring that the training focuses on high-quality positive samples, thereby improving the average precision (AP).

[0154] Overall loss function:

[0155] By combining CBAL, VFL, the location regression loss and classification loss of object detection, the overall loss function of the CAFOD model is obtained, as shown in the following formula:

[0156] Loss = L cls + L reg + L cbal + L vfl

[0157] Among them, L cls represents the category classification loss, which is calculated by the cross-entropy loss function, and L reg represents the location regression loss, including L1 loss and generalized IoU loss.

[0158] Example 2

[0159] This example evaluates the performance of the proposed model, including overall performance comparison and detection visualization, model complexity analysis, multi-camera performance analysis, ablation study, and feature visualization.

[0160] 1. Data description

[0161] To train the model, explosion-proof cameras and lighting lamps for coal mines were deployed on the conveyor belts of three different working faces underground in the coal mine, and training data was obtained through the acquisition and processing of the monitoring video images. Specifically, three conventional coal-conveying belts were selected as foreign object detection conveyor belts at three different working faces underground in the coal mine. Mining cameras and lighting lamps were used and installed 2 meters directly above the conveyor belts. Ten segments of conveyor belt coal-conveying videos at different times were collected respectively. Training picture data was obtained by intercepting 5 adjacent frames of the video frames of non-coal foreign objects on the conveyor belt, and then the target boxes were first labeled using the Labelme software in the coco dataset format. Finally, a total of 4449 conveyor belt foreign object samples were obtained under three cameras, which were divided into a training set (70%) and a test set (30%). This dataset consists of four types of foreign objects: large coal blocks, anchor bolts, iron nets, and wooden blocks. And for the positioning of the conveyor belt positions in each camera, it was completed by using the interactive graphical user interface of the OpenCV library. First, the image data stored in RGB was displayed in the OpenCV window, and the four vertices of the conveyor belt corners were selected on the image using the mouse event callback function. The coordinate points clicked each time were recorded and marked with red dots on the image, and the coordinate data was stored in a txt file. In this way, the vertex coordinates of the conveyor belt areas under different cameras were finally obtained.

[0162] 2. Experimental settings

[0163] The proposed method and other benchmark models were trained and evaluated in an experimental environment of NVIDIA Tesla P100, Cuda 12.0, and Pytorch 2.0.0. The experiment was carried out with a learning rate of 0.00002 and a batch size of 4.

[0164] 3. Numerical Analysis of Comparative Experiment Results

[0165] To verify the detection performance of the method we proposed for non - coal foreign objects on the coal mine conveyor belt, we used the above - mentioned data for experiments and selected O2F, Deformable - DETR, CEASE, Sparse - rcnn, AFPN, CentripetalNet, Autoassign, and VarifocalNet as comparison methods.

[0166] To analyze the overall performance difference between the proposed method and existing methods, the AP50 (average precision under the condition that the IOU (Intersection over Union) is equal to 0.5) is used as the overall detection performance evaluation index. We used existing SOTA methods to conduct evaluation and comparison experiments on our dataset. By training these models on the same experimental environment and dataset, we compared their comprehensive performance for non - coal foreign object detection and the performance for different categories of non - coal foreign objects. The performance comparison by category is shown in Table 1.

[0167] Table 1 Results of Performance Comparison by Category

[0168]

[0169] The results show that the CAFOD model proposed in the present invention is superior to existing state - of - the - art methods, achieving the highest detection accuracy of 80.39%. The CAFOD model of the present invention uses MVDA to capture more diverse features, thus solving the limitations in the case of multiple cameras or multiple perspectives. The CFP module fuses the local background (surrounding coal) features with potential targets, reducing false detections outside the conveyor belt.

[0170] 4. Model Complexity Analysis

[0171] In coal mine applications, due to the high - speed operation of the conveyor belt, the inference speed of foreign object detection is crucial. Therefore, the complexities of each model were compared. All models were trained on the entire dataset to ensure fairness. The experiments were conducted under the same hardware configuration: Intel(R)Xeon(R)CPU E5 - 2666v3@2.90GHz, 20GB RAM, and 16GB NVIDIA Tesla P100 GPU. Table 2 shows the parameter scale, GPU memory usage, RAM consumption, and calculation time during training and inference of each model.

[0172] Table 3 Results of Model Complexity Analysis

[0173]

[0174]

[0175] The results show that during training, each model generally requires a large amount of time (measured in hours) to optimize parameters, while in the inference stage, all models can process one image within 1 second. This indicates that the CAFOD model meets the real-time detection requirements of conveyor belt applications and also adapts to the requirements of edge devices.

[0176] 5. Multi-camera Adaptability Evaluation

[0177] To verify the applicability of the model of the present invention in the complex environment of coal mines, especially the stability and robustness of the model when facing changes in coal mine application scenarios, a cross-camera detection experiment was conducted. The test aimed to simulate various changes in the working face environment that might be encountered in actual applications, as well as the scene differences under different camera perspectives. Specifically, the model was trained using the data of a single camera (Camera 1) in the above dataset, and inference was performed on the data of multiple cameras (Camera 2 and Camera 3). With MAP50 as the evaluation metric, the generalization ability and stability of this evaluation model were evaluated. The experimental results are shown in Table 3.

[0178] Table 3 Results of Multi-camera Adaptability Evaluation

[0179]

[0180]

[0181] The results show that all the comparison models can achieve good accuracy on the subset of Camera 1. However, due to the significant differences between the backgrounds of other cameras and the training data, the detection performance of the model drops sharply on these images. This indicates that the generalization ability of the model is weak and it is difficult to adapt to changes in different camera scenarios. However, by improving the baseline method Deformable DETR, CBAL can enable the model to effectively perceive the conveyor belt area. CFP associates coal with foreign objects through the integration of context features, effectively reducing the impact of the drastic changes in the backgrounds of Camera 2 and Camera 3. Therefore, the model of the present invention can still maintain relatively stable performance in different scenarios without fine-tuning.

[0182] 6. Ablation Experiment Analysis

[0183] To analyze the effectiveness of the proposed module, ablation experiments were conducted using the above dataset, and CAFOD was compared with four methods. The experimental results are shown in Table 4.

[0184] Table 4 Results of Ablation Experiments

[0185]

[0186] The experimental results show that MVDA, CFP, CBAL, and VFL have significantly improved the detection effects on different foreign objects. Specifically, by introducing MVDA, Method B has significantly enhanced the detection performance for anchor bolts and wire meshes. Method C utilizes CFP to associate coal with foreign objects through context feature integration, improving the model's ability to identify potential foreign objects and thus enhancing the detection performance for various targets. Method D adopts a penalty-driven strategy to penalize predictions outside the conveyor belt area, increasing the model's attention to the conveyor belt area, especially in reducing the influence of targets outside the conveyor belt, resulting in a significant improvement in the detection performance of wire meshes. Finally, CAFOD introduces VFL, combines the category and location scores of the prediction targets, and dynamically adjusts the loss weights of samples, further enhancing the overall performance of the model.

[0187] 7. Error Type Analysis

[0188] The TIDE evaluator is introduced to analyze different types of errors to evaluate the contribution of each module to the performance of CAFOD. The results are shown in Table 5.

[0189] Figure 4 The error types in include:

[0190] (1) Classification error (Cls): IoUmax > 0.5, but the classification is incorrect;

[0191] (2) Localization error (Loc): The classification is correct, but 0.1 < IoUmax < 0.5;

[0192] (3) Classification and localization error (Cls+Loc): 0.1 < IoUmax < 0.5 and the classification is incorrect;

[0193] (4) Missed detection error (Missed): All ground truths that are not detected (except Cls errors and Loc errors);

[0194] (5) Background error (Bkgd): All ground truths with IoUmax < 0.1.

[0195] Table 5 Results of Error Type Analysis

[0196]

[0197] The experimental results show that Method B enables the model to learn more diverse features and make more accurate predictions of the target location, reducing classification errors and localization errors. Method C integrates the foreign object with its contextual features and correlates the foreign object with the background features, reducing the background false detection of the model and lowering the missed detection and background errors. Method D is designed specifically for foreign object detection on the conveyor belt. Through a penalty-driven mechanism, it further focuses on the conveyor belt area, enhancing the activation intensity of the features in this area, thereby improving the performance of the model in terms of location and category and reducing the false detection of targets outside the conveyor belt. Finally, VFL enhances the weights of difficult samples during the training process by dynamically adjusting the loss weights, improving the classification and localization capabilities of the model and further reducing classification and localization errors.

[0198] 8. Feature Map Visualization

[0199] To demonstrate the contribution of CFP, we visualized the feature maps in Figure 5 . By comparing the heatmaps with and without CFP added, it can be found that the boundary of Method A without the CFP module is not clear enough and the activation intensity of the background is relatively high, while the heatmap of CAFOD can better focus on the foreign objects on the conveyor belt in a complex background and noise environment. This is because CFP captures the features around the potential foreign objects and learns the association between the foreign objects and their surrounding features. Figure 6 The change of the CBAL value in

[0200] The above shows and describes the main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention.

[0201] In addition, it should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A coal mine conveyor belt foreign body detection method based on deep learning, characterized in that: The following steps are involved: Step 1, collect the coal mine conveyor belt foreign body image, and perform multi-view data enhancement on it to generate an enhanced image; Step 2, input the enhanced image into a pre-trained coal mine conveyor belt foreign body detection model, and output the predicted target location and category; the coal mine conveyor belt foreign body detection model includes a backbone network, a query network and a classification regression network; the backbone network uses the ResNet50 architecture to process the enhanced image, which is used to extract multi-scale features and generate a multi-scale feature map; the query network includes an encoder, a decoder and a target background feature fusion module, the encoder processes the multi-scale feature map from the backbone network through a self-attention mechanism to capture the global context information in the image and obtain an enhanced feature map, the decoder is responsible for generating and using a learnable query vector to extract target features from the enhanced feature map, and the context background feature fusion module extracts background features around the target and combines them with the extracted target features; The classification and regression network consists of a classification module and a regression module, which are responsible for target category recognition and bounding box position prediction respectively.

2. According to the method of detecting foreign matter in coal mine conveyor belt based on deep learning in claim 1, it is characterized in that: The multi-view data enhancement in step 1 is specifically as follows: Step 1.1, define four transformation matrices, namely the view transformation matrix P, the scaling matrix S, the rotation matrix R and the translation matrix T. The expressions are as follows: Among them, P x and P y is a parameter randomly generated in the range of [-0.001, 0.001]; S x , S y They represent the shear strength along the corresponding axis, S x , S y In the range of [-10°, 10°]; θ represents the image rotation angle, θ is in the range of [-15°, 15°]; T x , T y represents the translation distance in the x and y directions, T x In the range of [-0.1w, 0.1w], T y In the range [-0.1h, 0.1h], it is randomly generated within 10% of the image width w and height h; Step 1.2, combine these transformation matrices to form the total transformation matrix A aug , as shown below: A aug =P×S×R×T X'=A aug ×X B'=A aug ×B Among them, X' and B' are the transformed image and the corresponding bounding box, A aug ×X is the affine transformation applied to each pixel coordinate of the image, A aug ×B is the transformation of the coordinate frame of the real target corresponding to the image.

3. According to the method of detecting foreign matter in coal mine conveyor belt based on deep learning in claim 2, it is characterized in that: The backbone network in step 2 is specifically: The enhanced image is extracted through ResNet50 to extract 4 layers of feature maps of different scales, and the information of different levels in the image is gradually filtered out. The first three layers come from layer3, layer4, and layer5 of the ResNet network, with downsampling rates of 8, 16, and 32 respectively, and then a 1*1 convolution is used to unify the feature dimensions; By performing a 3*3 convolution on the features of layer3, we get features with a downsampling rate of 64,256 dimensions, where the feature map of the lth layer is Indicates that it has C l channels, with a height of H l , width W l feature map.

4. The method for detecting foreign matter in a coal mine conveyor belt based on deep learning according to claim 3 is characterized in that: The encoder in step 2 is specifically: Input the multi-scale feature map from the backbone network, map all feature maps to the same number of channels, introduce position encoding, generate a learnable spatial coordinate embedding for each position in the multi-scale feature map through sine-cosine position encoding, and add it to the features of each position in the feature map; For the features at each position in the multi-scale feature map, a query vector, a key vector, and a value vector are generated through linear transformation. At the same time, a set of dynamic offsets are generated based on the reference points. The reference points are initialized as uniformly distributed grid points. The offset and attention weight of each reference point are predicted from the concatenation of the query vector and the position encoding through the deformable attention mechanism. The attention weight is weighted and summed with the offset feature value. The calculation formula is as follows: Where x∈R C×H×W is a given input feature map, z q is the feature of each query vector in the encoder, P q is the reference point for each query vector in the encoder, M represents the number of attention heads, and Δp mqk is the learnable offset, A mqk is the attention weight, W m is the projection matrix of the m-th multi-head attention output, W' m x(·) is the feature extracted by bilinear interpolation of the offset coordinates, and is obtained by the projection matrix W m Mapped to the feature space of the mth attention head; After completing the attention calculation, it passes through a residual connection and normalization layer, a feedforward network, and uses the output of the feedforward network and the output of the previous residual connection and normalization layer as the input of the next residual connection and normalization layer; Repeat the above calculation L times, and use the output as the input for the next time. Through this iterative method, the feature representation is continuously optimized, and finally the multi-scale feature map z output by the encoder is obtained. final .

5. A coal mine conveyor belt foreign body detection method based on deep learning according to claim 4, characterized in that: The decoder in step 2 is specifically: Generate a learnable query vector q with random initialization init and encoder enhanced features z final As input, each query vector is embedded with spatial coordinate information through position encoding to generate an initial reference point as the target candidate position that the decoder focuses on; Perform self-attention operation and randomly initialize the generated query vector q init The dependencies between targets are learned through interaction, redundant candidate boxes are suppressed, and residual connections and layer normalization are performed; Perform a deformable attention operation to connect the residual connection and the layer normalized query vector with the multi-scale feature map z output by the encoder final Associate, dynamically correct the offset of the reference point, and in the multi-scale feature map z final Sampling is performed on the , and features of different resolutions are extracted. The feature value of each sampling point is multiplied by its corresponding weight and then summed. The calculation formula is: Among them, z q Represents the features of each query vector in the decoder, and the multi-scale features z output by the encoder final generate, Represents the reference point of each query vector in the decoder, which is used to dynamically generate the spatial reference coordinates of the target, initialized to generate random target candidate positions, is a collection of multi-scale feature maps, Represents the normalized coordinate space mapping the reference point to the feature map of the lth layer, Δp mlqk A represents the learnable offset of the mth attention head, the lth layer feature map, and the kth sampling point. mlqk represents the attention weight, that is, the importance of the k-th sampling point in the m-th head and the l-th layer, W' m x l (·) indicates that on the feature map of the first layer, bilinear interpolation is performed on the offset coordinates to extract features, and the projection matrix W is used to extract features. m Mapped to the feature space of the mth attention head; After completing the attention calculation, the gradient propagation is stabilized through two residual connections and layer normalization, and high-order features are further extracted through the feed-forward network; After L layers of iterative calculations, we finally get the feature F containing the predicted target location and category information. target , used for subsequent target classification and regression.

6. A coal mine conveyor belt foreign body detection method based on deep learning according to claim 5, characterized in that: The context background feature fusion module in step 2 is specifically: The decoder is used to generate features F containing the predicted target location and category information target , used to generate the target bounding box B target =(x center ,Y center ,w,h), and expand it into an enlarged bounding box B expand , the calculation formula is as follows: w expanded =w·scale h expanded =h·scale B expand =(x center ,y center ,w expanded ,h expanded ) Among them, w expand and h expand for width and height; According to B target and B expand , generate two masks M that match their size target and M expand , through the subtraction operation, we get the surrounding mask M context , the calculation formula is as follows: M context =M expand -M target ; The mask M context With the feature map Z final Multiply element by element, and adjust the maximum pooling operation to match F target For the same dimensions, the calculation formula is as follows: F context =Pooling(Z final ⊙M context ); The feature F target With feature F context Combined, the calculation formula is as follows: F integrate =F target +F context Latent target feature F containing contextual information integrate It is input into the classification module for classification task.

7. A coal mine conveyor belt foreign body detection method based on deep learning according to claim 6, characterized in that: The classification module in step 2 is specifically: The latent target feature F containing context information output by the context background feature fusion module integrate , high-dimensional features are mapped to the category space through a fully connected layer, and the output values ​​of each category are normalized using the softmax function to generate the category probability corresponding to each target.

8. A coal mine conveyor belt foreign body detection method based on deep learning according to claim 7, characterized in that: The regression module in step 2 is specifically: Use the decoder output feature F containing the predicted target location and category information target , the features are mapped to 4-dimensional space through a multi-layer perceptron, corresponding to the center point and width and height of the bounding box respectively, and finally a tensor of dimension (N, 4) is output, representing the bounding box coordinates of each candidate target.

9. A coal mine conveyor belt foreign body detection method based on deep learning according to any one of claims 1 to 8, characterized in that: It also includes the training of the coal mine conveyor belt foreign body detection model, using the conveyor belt area loss, variable focus loss, position regression loss, and category classification loss to optimize the model parameters, as follows: The conveyor belt area loss: Manually mark the conveyor belt area Poly from different camera perspectives, expressed as: Poly=[(x1,y1),(x2,y2),…,(x n ,y n )] Among them, (x1,y1),(x2,y2),…,(x n ,y n ) represents the vertices of the conveyor belt polygon; Evaluate whether the center point of the predicted bounding box is within the conveyor belt area, that is, calculate the center point Pred from each predicted bounding box j =(x c ,y c )Num number of intersections between the horizontally emitted ray and the polygon boundary points , the calculation formula is as follows: Among them, sign represents a function to determine the positive or negative value. points When Num is an odd number, the prediction box is located in the conveyor belt area. points When it is an even number, the prediction box is outside the conveyor belt area; Calculate the loss of the prediction box outside the conveyor belt area, the calculation formula is as follows: Outside(Pred j ,Poly)=1-(Num points mod2) Where m is the total number of prediction boxes and mod2 is the modulo operation; The variable focus loss: By dynamically adjusting the weights and focus factors of positive and negative samples, false positives can be mitigated and reduced. The calculation formula is as follows: Among them, p represents the predicted IoU score, q represents the target score, α is the weight balance factor of positive and negative samples, and γ is the focus factor; Overall loss function: Loss=L cls +L reg +L cbal +L vfl Among them, L cls represents the category classification loss, calculated by the cross entropy loss function, L reg Represents the position regression loss, including L1 loss and generalized IoU loss.

Citation Information

Cited By

  • Foreign matter detection method based on radar point cloud and visual image fusion

    CN122416128A