Belt Damage Detection Method and System, Equipment, and Storage Medium for Roadheader Mucking
By using the combination method of MaskRCNN network and packet convolution module in the damage detection of slag belts of the tunneling machine, the problems of poor detection accuracy and high error recognition rate in the prior art are solved, and high accuracy damage detection is achieved.
Patent Information
- Application Number
- CN202310237416.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-03-13
AI Technical Summary
The prior art has problems with poor detection accuracy and high misidentification rate in the detection of slag belts from the tunneling machine, mainly due to the insufficient deep feature extraction capability of the Alexnet model.
The MaskRCNN damage recognition model based on the ResNet50 skeleton network was used for belt damage detection, and the 3x3 convolution layer in the ResNet50 skeleton network was replaced by a transversely densely connected packet convolution module during the model training process.
The accuracy of belt damage detection is improved, the false recognition rate and false detection rate are reduced, and the high-accuracy damage detection is achieved.
Smart Images

Figure CN116309426B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of conveyor belt damage detection, and in particular to a method and system for detecting damage to a slag belt of a roadheader, electronic equipment, and a computer-readable storage medium. Background Art
[0002] Continuous belt conveyors are important material transportation equipment in the fields of steel smelting, mining, power plants, building materials, etc. They have many advantages such as long transportation distance, high transportation efficiency and low operating cost. Therefore, in the construction process of tunnel boring machines, most construction sites use belt conveyors to discharge slag over long distances. In the process of conveying slag, the continuous belt conveyor will cause different degrees and types of damage to the belt surface due to long continuous working time, abnormal wear of the rollers on the belt, and the impact of crushed stone and slag on the belt under the action of external forces. If these damages are not handled in time, the transportation efficiency will be reduced at the least, and the frame structure will be damaged at the worst, seriously threatening the personal safety of construction workers. Therefore, it is of great significance to detect belt damage in time and make corresponding decisions and handle it.
[0003] At present, patent CN113682762A discloses a method and system for belt tear detection based on machine vision and deep learning, which uses a CCD industrial camera, an image transmission module based on the GigE Vision interface, a deep learning model training module, a tear recognition module and a motion control module to form a belt tear detection system. The patent enhances the quality of the grayscale image collected by the camera through Fourier transform, and then transmits the enhanced image to the trained Alexnet belt tear recognition model to perform belt tear detection. If the belt tear is not detected, the next round of detection is performed. If the belt tear is detected, an alarm and braking command are sent to the controller, and then the tear image is manually confirmed and saved in the local deep learning data set, and the tear recognition detection model is updated. However, due to the harsh working environment of the roadheader slag conveyor belt, the damage shapes are different, and the damage accounts for a small proportion of the image pixels, and the deep feature extraction capability of the Alexnet model is insufficient, resulting in poor belt damage detection accuracy and high misrecognition rate. Summary of the invention
[0004] The present invention provides a method and system for detecting damage to a slag belt of a roadheader, electronic equipment, and a computer-readable storage medium, so as to solve the technical problems of poor damage detection accuracy and high misidentification rate in the existing damage identification using the Alexnet model.
[0005] According to one aspect of the present invention, a method for detecting damage to a slag belt of a roadheader is provided, comprising the following contents:
[0006] Collect the surface images of the conveyor belt during operation, and label the contours of the damaged areas in the images to create a damage detection sample dataset;
[0007] Construct a Mask R-CNN damage recognition model based on the ResNet50 backbone network;
[0008] Use the damage detection sample dataset to train the Mask R-CNN damage recognition model until the model converges. During the model training process, replace the 3x3 convolutional layer that composes the residual module in the ResNet50 backbone network with a grouped convolutional module with lateral dense connections;
[0009] Obtain the image of the surface of the conveyor belt to be detected, input it into the trained Mask R-CNN damage recognition model, and output the damage detection result.
[0010] Further, the grouped convolutional module includes four convolutional paths. When performing grouped convolution, first divide the input feature map into four groups X1, X2, X3, and X4 in the channel order, and the number of channels in each path is 1 / 4 of the input feature map. In the first convolutional path, the input feature Figure X 1 undergoes an identity mapping to obtain the output feature map Y1. In the second convolutional path, first add the input feature Figure X 2 to the output feature map Y1, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y2. In the third convolutional path, first add the input feature Figure X 3 to the output feature map Y1 and the output feature map Y2, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y3. In the fourth convolutional path, first add the input feature Figure X 4 to the output feature map Y1, the output feature map Y2, and the output feature map Y3, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y4. Finally, sequentially splice the output feature maps Y1, Y2, Y3, and Y4 to obtain the final output feature map.
[0011] Further, the dense convolution reparameterization module is composed of a transformation structure and a 3x3 repConv unit connected in sequence. Among them, the transformation structure is composed of three cascaded 1x1 Conv-BN layers, and skip connections with BN layers are added between any two 1x1 Conv-BN layers. The 3x3 repConv unit is composed of four parallel convolution paths. The first convolution path is composed of a 1x1 Conv layer, a BN layer, a 3x3 Avgpool layer, and a BN layer connected in sequence. The second convolution path is composed of a 1x3 Conv layer and a BN layer connected in sequence. The third convolution path is composed of a 3x1 Conv layer and a BN layer connected in sequence. The fourth convolution path is composed of a 1x1 Conv layer, a BN layer, a 3x3 Conv layer, and a BN layer connected in sequence. The output feature maps of the four convolution paths are added to obtain the final output feature map.
[0012] Further, after the model training is completed, the following also includes:
[0013] Reparameterize the dense convolution reparameterization module into a 3x3 Conv layer to convert the damage identification training model with a multi-branch structure into a single-link damage identification inference model.
[0014] Further, the process of reparameterizing the dense convolution reparameterization module into a 3x3 Conv layer is specifically as follows:
[0015] Reparameterize the transformation structure into a single 3x3 Conv layer;
[0016] Reparameterize the 3x3 repConv unit into a single 3x3 Conv layer;
[0017] Merge the two reparameterized 3x3 Conv layers to obtain the final single 3x3 Conv layer.
[0018] Further, the process of reparameterizing the transformation structure into a single 3x3 Conv layer is specifically as follows:
[0019] Merge the 1x1 Conv-BN layers into a 3x3 Conv layer;
[0020] Merge the convolutional layer with skip connections and the corresponding skip connection branches into a single 3x3 Conv layer;
[0021] Merge the two sequentially connected 3x3 Conv layers into a single 3x3 Conv layer, and so on, finally obtaining a single 3x3 Conv layer equivalent to the transformation structure.
[0022] Further, the process of reparameterizing the 3x3 repConv unit into a single 3x3 Conv layer is specifically as follows:
[0023] In the first convolutional path, first, the 3x3 Avgpool layer is equivalent to a special weighted convolutional layer with all weights being 1 / 9. Then, the 1x1 Conv layer, BN layer, 3x3 Avgpool layer, and BN layer are combined to obtain a single 3x3 Conv layer;
[0024] In the second convolutional path, the 1x3 Conv layer is padded with 0s to a 3x3 Conv layer, and then combined with the BN layer to form a 3x3 Conv layer;
[0025] In the third convolutional path, the 3x1 Conv layer is padded with 0s to a 3x3 Conv layer, and then combined with the BN layer to form a 3x3 Conv layer;
[0026] In the fourth convolutional path, the 1x1 Conv layer, BN layer, 3x3 Conv layer, and BN layer are combined to form a 3x3 Conv layer;
[0027] The four 3x3 Conv layers are combined into a single 3x3 Conv layer.
[0028] In addition, the present invention also provides a detection system for the damage of the slag discharge belt of a roadheader, including:
[0029] A sample collection module, which is used to collect the surface image of the belt during operation, and label the contour of the damaged area in the image to make a damage detection sample data set;
[0030] A model construction module, which is used to construct a Mask R-CNN damage recognition model based on the ResNet50 backbone network;
[0031] A model training module, which is used to train the Mask R-CNN damage recognition model with the damage detection sample data set until the model converges. Among them, during the model training process, a grouped convolutional module with horizontal dense connections is used to replace the 3x3 convolutional layer that composes the residual module in the ResNet50 backbone network;
[0032] A damage detection module, which is used to obtain the image of the surface of the belt to be detected, input it into the trained Mask R-CNN damage recognition model, and output the damage detection result.
[0033] In addition, the present invention also provides an electronic device, including a processor and a memory. A computer program is stored in the memory, and the processor is used to execute the steps of the method described above by calling the computer program stored in the memory.
[0034] In addition, the present invention also provides a computer-readable storage medium, which is used to store a computer program for detecting the damage of the slag discharge belt of a roadheader. When the computer program runs on a computer, it executes the steps of the method described above.
[0035] The present invention has the following effects:
[0036] For the method for detecting damage to the slag discharging belt of a roadheader in the present invention, a damage recognition model is constructed using a Mask R-CNN network for belt damage detection. Through multi-task joint training, a single deep learning model can simultaneously complete the tasks of belt defect detection and defect segmentation. The damage detection results include the number of damages, classification confidence, position bounding box, and pixel region corresponding to the segmented damage. Moreover, during the model training process, the 3x3 convolutional layer that constitutes the residual module in the ResNet50 backbone network is replaced with a grouped convolutional module with lateral dense connections. Grouped convolution can improve the multi-scale feature extraction ability of the model at a finer granularity level. At the same time, lateral dense connections are used in the grouped convolution to strengthen feature reuse, avoid gradient disappearance, and improve the recognition accuracy of the model. Therefore, the method for detecting damage to the slag discharging belt of a roadheader in the present invention has the advantages of high damage detection accuracy, low false recognition rate, and low false detection rate.
[0037] In addition, the system for detecting damage to the slag discharging belt of a roadheader in the present invention also has the above advantages.
[0038] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0040] Figure 1 is a schematic flow chart of the method for detecting damage to the slag discharging belt of a roadheader in a preferred embodiment of the present invention.
[0041] Figure 2 is a schematic structural layout diagram of a detection device using the damage detection method of the present invention in a preferred embodiment of the present invention.
[0042] Figure 3 is a schematic network structure diagram of the feature extraction backbone network in a preferred embodiment of the present invention.
[0043] Figure 4 is a schematic network structure diagram of the GCIM module in a preferred embodiment of the present invention.
[0044] Figure 5 is a schematic network structure diagram of replacing the 3x3 Conv layer of the residual module with a grouped convolutional module in a preferred embodiment of the present invention.
[0045] Figure 6 It is a schematic diagram of the network structure of the dense convolution reparameterization module in the preferred embodiment of the present invention.
[0046] Figure 7 It is another schematic diagram of the process of the method for detecting damage to the slag discharge belt of a roadheader in the preferred embodiment of the present invention.
[0047] Figure 8 It is Figure 7 a schematic diagram of the sub-process of step S3a in
[0048] Figure 9 It is Figure 8 a schematic diagram of the sub-process of step S31a in
[0049] Figure 10 It is Figure 8 a schematic diagram of the sub-process of step S32a in
[0050] Figure 11 It is a schematic diagram of the module structure of the system for detecting damage to the slag discharge belt of a roadheader in another embodiment of the present invention. Specific embodiments
[0051] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the following.
[0052] It can be understood that, as Figure 1 shown, the preferred embodiment of the present invention provides a method for detecting damage to the slag discharge belt of a roadheader, including the following:
[0053] Step S1: Collect the surface images during the operation of the belt, and mark the contours of the damaged areas in the images to make a damage detection sample dataset;
[0054] Step S2: Construct a Mask R-CNN damage recognition model based on the ResNet50 backbone network;
[0055] Step S3: Use the damage detection sample dataset to train the Mask R-CNN damage recognition model until the model converges. Among them, during the model training process, the 3x3 convolutional layer that composes the residual module in the ResNet50 backbone network is replaced by a grouped convolutional module with horizontal dense connections;
[0056] Step S4: Obtain the image of the surface of the belt to be detected, input it into the trained Mask R-CNN damage recognition model, and output the damage detection result.
[0057] It can be understood that the slag discharging belt damage detection method of this embodiment uses the MaskRCNN network to construct a damage recognition model for belt damage detection. Through the method of multi-task joint training, a single deep learning model can simultaneously complete the belt defect detection and defect segmentation tasks. The damage detection results include the number of damages, classification confidence, position bounding box, and the pixel region corresponding to the segmented damage. Moreover, during the model training process, the 3x3 convolutional layer that composes the residual module in the ResNet50 backbone network is replaced by a grouped convolutional module with lateral dense connections. Grouped convolution can improve the multi-scale feature extraction ability of the model at a finer granularity level. At the same time, lateral dense connections are used in the grouped convolution, strengthening the reuse of features, avoiding gradient disappearance, and improving the recognition accuracy of the model. Therefore, the slag discharging belt damage detection method of the present invention has the advantages of high damage detection accuracy, low false recognition rate, and low false detection rate.
[0058] It can be understood that the schematic structural layout diagram of the detection device using the above-mentioned slag discharging belt damage detection method is as Figure 2 shown. The detection device includes a belt cleaning device, a belt air drying device, a camera, a light source, an industrial network cable, a network switch, an industrial control computer, and a video display. The belt cleaning device is located directly below the muck belt, and the water outlet is perpendicular to the horizontal plane of the belt, and is used to wash away the residue and dust on the belt surface, making the surface of the belt to be photographed as clean as possible without muck. The belt air drying device is located directly below the muck belt behind the cleaning device, and the air outlet is perpendicular to the horizontal plane of the belt, and is used to reduce the water marks on the belt surface. The camera is vertically installed at the middle position 1.5 meters directly below the belt behind the cleaning device through a bracket to photograph the side of the belt carrying muck. The light source is installed obliquely above the camera through a bracket, with one light source on each side of the camera, providing a stable illumination environment with uniform brightness distribution for the camera shooting area. The camera, network switch, and industrial control computer are connected in sequence through network cables. The camera collects images of the belt during operation in real time and transmits them to the industrial control computer. The belt damage detection model deployed on the industrial control computer is used to detect and segment the damage images on the belt surface in real time and display them on the video display.
[0059] It can be understood that in step S1, when the slag discharging belt conveyor is started, the belt cleaning device and the belt blowing device are turned on to clean the belt surface. Then, the camera is used to collect the RGB images of the belt during operation, and the Labelme tool is used to contour label the damage areas in the images to obtain the corresponding json label file and Mask image of the images, so as to make the damage detection sample set E. The present invention uses the belt cleaning device in cooperation with the belt air drying device to preprocess the camera shooting area, reducing the interference of the residual muck and dust on the belt on the imaging of the damage area, and can improve the accuracy of damage detection. It can be understood that in other embodiments of the present invention, when the surface of the belt conveyor is clean, the cleaning and blowing steps can also be omitted.
[0060] In addition, in the step S1, the damage detection sample set E is also randomly divided into a training sample set E1, a validation sample set E2, and a test sample set E3. Among them, the training sample set E1, the validation sample set E2, and the test sample set E3 are all obtained from the damage detection sample set E by non-repetitive random sampling, and their respective proportions of the damage detection sample set E are 70%, 15%, and 15%.
[0061] It can be understood that in the step S2, a Mask R-CNN network model with ResNet50 as the Backbone is built using the PyTorch deep learning framework, and the parameters of the model anchor boxes anchors are initialized. Among them, the scale is set to (16, 32, 64, 128, 256), and the ratio is set to [0.5, 1.2]. Among them, the existing Mask R-CNN network architecture mainly includes a Backbone (i.e., the feature extraction backbone network) and an RPN (Region Proposal Network). Among them, the Backbone part is constructed using the ResNet50 network. One of the improvements of the present invention to the existing Mask R-CNN network lies in: a pyramid feature fusion module FPN is introduced in the Neck part of the model, and a bottom-up pyramid structure including a context fusion module GCIM is added at the output end of the FPN.
[0062] Specifically, the improved feature extraction backbone network structure of the present invention is as Figure 3As shown in the figure, in the ResNet50 network structure, four sets of feature maps {C2, C3, C4, C5} are sequentially generated from the second prediction layer Block2 to the fifth prediction layer Block5, and their resolutions are {32x32, 64x64, 128x128, 256x256} respectively. First, a top-down pyramid structure is constructed by sequentially connecting the M5 feature map to the M2 feature map. The M5 feature map is generated from the C5 feature map through a 1x1Conv-BN-Relu layer, and has the same resolution and number of channels as C5. The M5 feature map generates a feature map M5' through a 2-fold linear upsampling operation. The C4 feature map generates a feature map C4' through a 1x1Conv-BN-Relu layer with a stride of 1. The M5' and C4' with the same resolution are added element by element for feature fusion processing to obtain a feature map M4 containing context information. The generation principles of the M3 and M2 feature maps are the same as that of M4, and will not be elaborated here. Then, the P5, P4, P3, and P2 feature maps are generated from the corresponding M5, M4, M3, and M2 feature maps through a 3x3Conv-BN-Relu layer with a stride of 1, and the P6 feature map is generated from the P5 feature map through a 2x2Maxpool operation with a stride of 2. Then, the feature map P2 undergoes a 3x3Conv-BN-Relu operation with a stride of 1 to obtain a feature map F2, and F2 has the same resolution and number of channels as P2. Then, the feature map P3 undergoes a BN-Relu operation and is input into the GCIM module together with F2, and finally a feature map F3 is obtained. The generation principles of the feature maps F4, F5, and F6 are the same as that of the feature map F3, so they will not be elaborated here. Then, F2, F3, F4, F5, and F6 are the feature maps finally generated by the Neck part of the model.
[0063] Among them, the schematic diagram of the network structure of the GCIM module is as Figure 4 shown. The GCIM module fuses features in two stages, namely Stage1 and Stage2. In Stage1, the input feature map Low_F is the low-resolution feature map generated by FPN, such as the feature map P2, and the input feature map High_F is the high-resolution feature map output by the previous convolutional layer in the buttom-up pyramid structure, such as the feature maps F2, F3, F4, F5. First, the input feature map High_F passes through a 3x3Conv-BN-Relu layer with a stride of 1 and a global pooling layer Global pool with a stride of 2 to obtain a feature map f s1 , then the input feature map Low_F undergoes a 3x3Conv-BN-Relu operation with a stride of 1 to obtain a feature map f s2 , and finally f s1 and f s2Perform an element-wise multiplication operation to obtain the intermediate output feature map f s3 . In Stage2, first, perform a deconvolution operation on f s1 to obtain the feature map f s4 , and then perform a pixel-wise addition operation on f s3 and f s4 to obtain the final output feature map Output_F.
[0064] It can be understood that based on the existing MaskRCNN network structure, the present invention introduces an FPN module in the Neck part of the model, and adds a bottom-up pyramid structure including a context fusion module GCIM at the output end of the FPN module, so as to achieve deep fusion of the underlying structural features and high-level semantic features, further enriching the expression ability of the feature map.
[0065] Optionally, another improvement of the present invention to the existing MaskRCNN network is that: in the RPN network part, the DIoU-NMS algorithm is used to replace the existing NMS algorithm to eliminate duplicate candidate boxes, and the regression loss function of the model is modified according to DIoU. Among them, DIoU represents the distance-based intersection over union of bounding boxes, which represents the overlap rate between the predicted bounding box and the ground truth bounding box. Among them, the specific steps of the DIoU-NMS algorithm are as follows:
[0066] Step 1: Set the candidate box threshold T, create sets Ω1 and Ω2, and initialize them as empty sets;
[0067] Step 2: For a number of candidate boxes generated based on anchors in the RPN, first eliminate the candidate boxes that exceed the image range, and then sort the remaining candidate boxes in descending order according to the foreground / background confidence, and select multiple candidate boxes with higher rankings and put them into the set Ω1; specifically, for 260488 candidate boxes [b1, b2,..., b 260488 , first eliminate the candidate boxes that exceed the image range, and then sort the remaining candidate boxes in descending order according to the foreground / background confidence, and select the first 2000 candidate boxes and put them into the set Ω1;
[0068] Step 3: Take the candidate box with the highest confidence in the set Ω1 and put it into the set Ω2, and then calculate the DIoU value of each candidate box in the set Ω1 and the candidate box in the set Ω2. The calculation formula is as follows:
[0069] Among them, b i represents the current candidate box in the set Ω1, and b jDenote the current candidate box in set Ω2, ρ represents the Euclidean distance between the centers of two candidate boxes, and c represents the diagonal distance of the smallest closed region that can contain both candidate boxes. If the calculated DIoU value is greater than the threshold T, then the candidate box b is removed. i If the DIoU value is less than or equal to the threshold T, then the candidate box b is retained. i In set Ω1;
[0070] Step 4: Repeat the above Step 3 until the number of candidate boxes in set Ω1 is 0 and then end.
[0071] It can be understood that the traditional NMS algorithm uses an IoU (Intersection over Union)-based metric to evaluate the position accuracy of the bounding box bbox. Although it can reflect the distance when the predicted bounding box and the ground truth bounding box intersect, when the two do not intersect or touch, it cannot reflect the distance between them, and at this time the loss function L re is zero, resulting in no gradient backpropagation in the object detection branch during backpropagation training. In addition, there will be a problem of divergence in the regression of the target bounding box during the model training process, which will affect the recognition accuracy of the model. Therefore, the present invention introduces the distance, overlap rate, and size between the predicted bounding box and the ground truth bounding box into the loss function of the object detection branch of the network, and uses the distance-based Intersection over Union of Bounding Boxes DIoU to modify the regression loss function of the model. Among them, the expression of the regression loss function is as follows:
[0072] where n represents the number of predicted bounding boxes, b represents the predicted bounding box, and b gt represents the labeled bounding box corresponding to the predicted bounding box, ρ represents the Euclidean distance between the centers of the predicted bounding box and the labeled bounding box, and c represents the diagonal distance of the smallest closed region that can contain both the predicted bounding box and the labeled bounding box.
[0073] It can be understood that the present invention adopts the DIoU-NMS algorithm to remove the duplicate candidate boxes generated in the RPN network, and uses the regression loss function based on DIoU, which can effectively avoid the problem of divergence in the regression of the target bounding box during the model training process and improve the damage detection accuracy of the model.
[0074] It can be understood that in the step S3, first use the image enhancement technology to preprocess the training image input to the model, and then use the Stochastic Gradient Descent SGD combined with the transfer learning technology to iteratively train the model offline until the model converges to obtain a belt damage detection model containing the optimal weight ω. The specific training process is as follows: First, set the hyperparameters of the model, and set the initial learning rate lr to 1x10 -3, the learning rate decay method is set to exponential decay, the decay coefficient ξ is set to 0.005, the batch size batch_size is set to 32, and the number of epochs is set to 500. Then, the image is randomly rotated, translated, horizontally and vertically flipped, and randomly scaled and cropped according to the probability θ, and the contrast, brightness, and saturation of the image are randomly adjusted to achieve image enhancement and prevent overfitting in model training. Then, the weight obtained by pre-training on the ImageNet public dataset is used to initialize the model parameter ω0. Then, the image size is scaled to 1024x1024x3 and fed into the model for offline iterative training. When epochs reach 500 or the mAP accuracy of the model on the validation sample set E2 reaches the specified value, the training can be ended to obtain a belt damage detection model containing the optimal weight ω. Then, the model accuracy is tested on the test sample set E3.
[0075] Among them, in order to improve the deep feature extraction ability of the model, during the model training process, the present invention adopts a grouped convolution module with horizontal dense connection (g = 4) and a convolutional structure reparameterization method to reconstruct the residual module in ResNet50. Specifically, as Figure 5 shown, the grouped convolution module includes four groups of convolutional paths. When performing grouped convolution, the input feature map is first divided into four groups X1, X2, X3, and X4 in the channel order, and the number of channels in each path is 1 / 4 of the input feature map. In the first group of convolutional paths, the input feature Figure X 1 passes through an identity mapping to obtain the output feature map Y1. In the second group of convolutional paths, first add the input feature Figure X 2 to the output feature map Y1, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y2. In the third group of convolutional paths, first add the input feature Figure X 3 to the output feature map Y1 and the output feature map Y2, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y3. In the fourth group of convolutional paths, first add the input feature Figure X 4 to the output feature map Y1, the output feature map Y2, and the output feature map Y3, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y4. Finally, the output feature maps Y1, Y2, Y3, and Y4 are sequentially concatenated to obtain the final output feature map.
[0076] It can be understood that the present invention uses a grouped convolution module to replace the 3x3 convolutional layer in the residual module, which can improve the multi-scale feature extraction ability of the model at a finer granularity level, and the horizontal dense connection method is adopted in the grouped convolution, which strengthens the reuse of features and avoids gradient disappearance.
[0077] Among them, as Figure 6As shown in the figure, the dense convolutional reparameterization module 3x3repDenseConv consists of a transformation structure and a diversified branch block 3x3repConv connected in sequence. Among them, the transformation structure consists of three cascaded 1x1Conv-BN layers, which can achieve network over-parameterization, is beneficial to increasing the model capacity, and a skip connection with a BN layer is added between any two 1x1Conv-BN layers, which can reduce the redundancy of features. The 3x3repConv unit consists of four parallel convolutional paths. The first convolutional path consists of a 1x1Conv layer, a BN layer, a 3x3Avgpool layer, and a BN layer connected in sequence. The second convolutional path consists of a 1x3Conv layer and a BN layer connected in sequence. The third convolutional path consists of a 3x1Conv layer and a BN layer connected in sequence. The fourth convolutional path consists of a 1x1Conv layer, a BN layer, a 3x3Conv layer, and a BN layer connected in sequence. The output feature maps of the four convolutional paths are added to obtain the final output feature map. By adopting a diversified branch block to replace a single convolutional layer, the feature extraction and expression ability of the model are greatly improved, thereby improving the recognition accuracy of the model. And 1x3Conv layers and 3x1Conv layers are introduced in the diversified branch block, which is beneficial to improving the robustness of the model to flipping and translation.
[0078] It can be understood that, as Figure 7 shown, after the model training is completed, the tunneling machine slag conveyor belt damage detection method further includes the following content:
[0079] Step S3a: Reparameterize the dense convolutional reparameterization module into a 3x3Conv layer to convert the damage recognition training model with a multi-branch structure into a damage recognition inference model with a single-link structure.
[0080] Specifically, as Figure 8 shown, in the step S3a, the process of reparameterizing the dense convolutional reparameterization module into a 3x3Conv layer is specifically as follows:
[0081] Step S31a: Reparameterize the transformation structure into a single 3x3Conv layer;
[0082] Step S32a: Reparameterize the 3x3repConv unit into a single 3x3Conv layer;
[0083] Step S33a: Combine the two reparameterized 3x3Conv layers to obtain the final single 3x3Conv layer.
[0084] It can be understood that, as Figure 9 shown, in the step S31a, the process of reparameterizing the transformation structure into a single 3x3Conv layer is specifically as follows:
[0085] Step S311a: Merge the 1x1 Conv-BN layer into a 3x3 Conv layer;
[0086] Step S312a: Merge the convolutional layer with skip connection and the corresponding skip connection branch into a single 3x3 Conv layer;
[0087] Step S313a: Merge two sequentially connected 3x3 Conv layers into a single 3x3 Conv layer, and so on, finally obtaining a single 3x3 Conv layer equivalent to the transformation structure.
[0088] Specifically, in Step S311a, first pad the 1x1 Conv layer with 0s to form a 3x3 Conv layer. The weights of the padded 3x3 Conv layer are denoted as W conv , the biases are denoted as b conv , the mean of the BN layer is denoted as μ, the standard deviation is denoted as σ, and the scaling factor is denoted as γ. Then the weights of the merged 3x3 Conv layer are The biases are
[0089] It can be understood that in Step S312a, the convolutional layer with skip connection and the corresponding skip connection branch are fused into a single 3x3 Conv layer. There are three types of skip connections: passing through one convolutional layer, passing through two convolutional layers, and passing through three convolutional layers. The schematic diagram is as Figure 6 shown. The skip connection can be equivalent to a 1x1 Conv layer with special weights where the convolutional kernel parameter for the current channel is 1. Pad the 1x1 Conv layer with 0s to form a 3x3 Conv layer, and then fuse it with the BN layer. Denote the weights of the convolutional layer with skip connection as W, the biases as b, the weights of the fused layer as W skip , the biases as b skip , the mean of the BN layer is denoted as μ, the standard deviation is denoted as σ, and the scaling factor is denoted as γ. Then the weights of the 3x3 Conv layer obtained after merging the skip connection and the convolutional layer are The biases are The specific merging order is as follows: First, merge the skip connections passing through only a single convolutional layer and the corresponding convolutional layers in sequence from top to bottom to obtain three 3x3 Conv layers. The weights of the merged layers are denoted as W1, W2, and W3 in sequence, and the biases are denoted as b1, b2, and b3 in sequence. Then, merge the skip connections passing through two convolutional layers and the corresponding convolutional layers in sequence from top to bottom. Denote the weights of the two skip connections as W skip1 and W skip2 , the biases as b skip1 and b skip2 , then the weights of the two merged convolutional layers are and The biases are and and To represent the transposition of the first and second dimensions of W2 and W3, * represents a two-dimensional convolution operation. Finally, the skip connection passing through three convolutional layers is merged with the corresponding convolutional layer, and the weight of the skip connection is denoted as W skip3 , and the bias is b skip3 , then the weight of the merged convolutional layer bias represents the transposition of for the first and second dimensions.
[0090] It can be understood that in the step S313a, two adjacent 3x3Conv layers connected in sequence are merged into a single 3x3Conv layer in the cascading order. Denote the weight of the first 3x3Conv layer as W1 and the bias as b1, and the weight of the second 3x3Conv layer as W2 and the bias as b2. Then the weight of the merged single 3x3Conv layer is W = W2 * W1 T , and the bias is b = b2 + (b1 × W2), where * represents a two-dimensional convolution operation, and W1 T represents the transposition of the first and second dimensions of W1. The above merging process is repeated for multiple subsequent cascaded 3x3Conv layers, and finally a single 3x3Conv layer equivalent to the entire transformation structure is merged.
[0091] It can be understood that as Figure 10 shown, in the step S32a, the process of reparameterizing the 3x3repConv unit into a single 3x3Conv layer is specifically as follows:
[0092] Step S321a: In the first convolutional path, first, the 3x3Avgpool layer is equivalent to a special weight convolutional layer with all weight values being 1 / 9, then the 1x1Conv layer is merged with the BN layer, and the 3x3Avgpool layer is merged with the BN layer to obtain two corresponding 3x3Conv layers. Finally, the two 3x3Conv layers are merged to obtain a single 3x3Conv layer; among them, the merging process refers to step S31a and will not be elaborated here;
[0093] Step S322a: In the second convolutional path, the 1x3Conv layer is padded with 0s to form a 3x3Conv layer, and then merged with the BN layer to form a 3x3Conv layer;
[0094] Step S323a: In the third convolutional path, the 3x1Conv layer is padded with 0s to form a 3x3Conv layer, and then merged with the BN layer to form a 3x3Conv layer;
[0095] Step S324a: In the fourth convolutional path, the 1x1Conv layer, BN layer, 3x3Conv layer, and BN layer are merged into a 3x3Conv layer; the merging process refers to Step S31a and will not be elaborated here;
[0096] Step S325a: Merge the four 3x3Conv layers into a single 3x3Conv layer. For example, denote the weight of the 3x3Conv layer in the first convolutional path as W1 and the bias as b1, and the weight of the 3x3Conv layer in the second convolutional path as W2 and the bias as b2. Then the weight after fusion of the 3x3Conv layers in the two convolutional paths is W add = W1 + W2, and the bias is b add = b1 + b2; and so on. The weight of the finally fused single 3x3Conv layer is calculated as W1 + W2 + W3 + W4, and the bias is b1 + b2 + b3 + b4.
[0097] It can be understood that in Step S33a, the 3x3Conv layer obtained by reparameterizing the transformation structure is merged with the 3x3Conv layer obtained by reparameterizing the 3x3repConv unit to obtain the final single 3x3Conv layer. The specific merging process refers to Step S33a and will not be elaborated here. By reparameterizing all 3x3repConv units in the residual modules of ResNet50 into a single 3x3Conv layer, the final belt damage detection inference model can be obtained.
[0098] It can be understood that in the model training stage of the present invention, the 3x3Conv layer of the residual module is reconstructed by using the conversion structure and the diversified branch block, which can improve the feature extraction ability of the model, and is reparameterized into a single-link structure in the inference stage, thereby accelerating the inference speed of the model and improving the damage detection efficiency.
[0099] In addition, as Figure 11 shown, another embodiment of the present invention further provides a tunneling machine slag discharge belt damage detection system, preferably adopting the above-mentioned damage detection method. The system includes:
[0100] A sample collection module, which is used to collect the surface image of the belt during operation, and mark the contour of the damage area in the image to make a damage detection sample data set;
[0101] A model construction module, which is used to construct a MaskRCNN damage recognition model based on the ResNet50 backbone network;
[0102] A model training module, which is used to train the MaskRCNN damage recognition model using a damage detection sample dataset until the model converges. During the model training process, a grouped convolution module with lateral dense connections is used to replace the 3x3 convolution layer that constitutes the residual module in the ResNet50 backbone network.
[0103] A damage detection module, which is used to obtain an image of the surface of the belt to be detected, input it into the trained MaskRCNN damage recognition model, and output a damage detection result.
[0104] It can be understood that the tunneling machine slag discharge belt damage detection system in this embodiment uses the MaskRCNN network to construct a damage recognition model for belt damage detection. Through the method of multi-task joint training, a single deep learning model can simultaneously complete the tasks of belt defect detection and defect segmentation. The damage detection result includes the number of damages, classification confidence, position bounding box, and the pixel region corresponding to the segmented damage. Moreover, during the model training process, a grouped convolution module with lateral dense connections is used to replace the 3x3 convolution layer that constitutes the residual module in the ResNet50 backbone network. Grouped convolution can improve the multi-scale feature extraction ability of the model at a finer granularity level. At the same time, dense connections are used in the grouped convolution, which strengthens the reuse of features, avoids gradient disappearance, and improves the recognition accuracy of the model. Therefore, the tunneling machine slag discharge belt damage detection system of the present invention has the advantages of high damage detection accuracy, low false recognition rate, and low false detection rate.
[0105] Optionally, the tunneling machine slag discharge belt damage detection system further includes:
[0106] A model conversion module, which is used to reparameterize the dense convolution reparameterization module into a 3x3 Conv layer to convert the damage recognition training model with a multi-branch structure into a damage recognition inference model with a single-link structure.
[0107] In addition, another embodiment of the present invention further provides an electronic device, including a processor and a memory. A computer program is stored in the memory. The processor is used to execute the steps of the method as described above by calling the computer program stored in the memory.
[0108] In addition, another embodiment of the present invention further provides a computer-readable storage medium, which is used to store a computer program for tunneling machine slag discharge belt damage detection. The computer program executes the steps of the method as described above when running on a computer.
[0109] The forms of common computer-readable storage media include: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media with a pattern of holes, random access memories (RAMs), programmable read-only memories (PROMs), erasable programmable read-only memories (EPROMs), flash erasable programmable read-only memories (FLASH-EPROMs), any other memory chips or cartridges, or any other media readable by a computer. Instructions can further be transmitted or received by a transmission medium. The term transmission medium can include any tangible or intangible medium that can be used to store, encode, or carry instructions for execution by a machine, and includes digital or analog communication signals or other intangible media that facilitate the communication of the above instructions. Transmission media include coaxial cables, copper wires, and optical fibers, which include the wires of a bus used to transmit a computer data signal.
[0110] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0111] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript, etc.
[0112] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementation in the process Figure 1One or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks
[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions in the process Figure 1 One or more processes and / or blocks Figure 1 specified in one or more blocks
[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the process Figure 1 One or more processes and / or blocks Figure 1 specified in one or more blocks
[0115] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application
[0116] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations
Claims
1. A method for detecting damage to the slag discharge belt of a roadheader, characterized in that, It includes the following contents: Collect the surface images of the belt during operation, and label the contours of the damaged areas in the images to produce a damage detection sample dataset; Construct a MaskRCNN damage recognition model based on the ResNet50 backbone network; Use the damage detection sample dataset to train the MaskRCNN damage recognition model until the model converges. During the model training process, replace the 3x3 convolutional layer that composes the residual module in the ResNet50 backbone network with a grouped convolutional module with lateral dense connections; Obtain the image of the surface of the belt to be detected, input it into the trained MaskRCNN damage recognition model, and output the damage detection result; The grouped convolutional module includes four groups of convolutional paths. When performing grouped convolution, first divide the input feature map into four groups X1, X2, X3, and X4 in the channel order. The number of channels in each path is 1 / 4 of the input feature map. In the first group of convolutional paths, the input feature map X1 passes through an identity mapping to obtain the output feature map Y1. In the second group of convolutional paths, first add the input feature map X2 and the output feature map Y1, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y2. In the third group of convolutional paths, first add the input feature map X3 and the output feature maps Y1 and Y2, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y3. In the fourth group of convolutional paths, first add the input feature map X4 and the output feature maps Y1, Y2, and Y3, and then input the added result into the dense convolutional reparameterization module to obtain the output feature map Y4. Finally, sequentially splice the output feature maps Y1, Y2, Y3, and Y4 to obtain the final output feature map; The dense convolutional reparameterization module is composed of a transformation structure and a 3x3repConv unit connected in sequence. Among them, the transformation structure is composed of 3 cascaded 1x1Conv-BN layers, and a skip connection with a BN layer is added between any two 1x1Conv-BN layers. The 3x3repConv unit is composed of four parallel convolutional paths. The first convolutional path is composed of a 1x1Conv layer, a BN layer, a 3x3Avgpool layer, and a BN layer connected in sequence. The second convolutional path is composed of a 1x3Conv layer and a BN layer connected in sequence. The third convolutional path is composed of a 3x1Conv layer and a BN layer connected in sequence. The fourth convolutional path is composed of a 1x1Conv layer, a BN layer, a 3x3Conv layer, and a BN layer connected in sequence. Add the output feature maps of the four convolutional paths to obtain the final output feature map.
2. The method for detecting damage to the slag discharge belt of a roadheader according to claim 1, characterized in that, After the model training is completed, it also includes the following contents: Reparameterize the dense convolutional reparameterization module into a 3x3Conv layer to convert the damage recognition training model with a multi-branch structure into a single-link damage recognition inference model.
3. The method for detecting damage to the slag discharge belt of a roadheader according to claim 2, characterized in that, The process of reparameterizing the dense convolutional reparameterization module into a 3x3Conv layer is specifically as follows: Reparameterize the transformation structure into a single 3x3Conv layer; Reparameterize the 3x3 repConv unit into a single 3x3 Conv layer; Merge the two reparameterized 3x3 Conv layers to obtain the final single 3x3 Conv layer.
4. The method for detecting damage to the slag discharge belt of a roadheader according to claim 3, characterized in that, The process of reparameterizing the transformation structure into a single 3x3 Conv layer is specifically as follows: Merge the 1x1 Conv-BN layer into the 3x3 Conv layer; Merge the convolutional layer with skip connection and the corresponding skip connection branch into a single 3x3 Conv layer; Merge two sequentially connected 3x3 Conv layers into a single 3x3 Conv layer, and so on, finally obtaining a single 3x3 Conv layer equivalent to the transformation structure.
5. The method for detecting damage to the slag discharge belt of a roadheader according to claim 3, characterized in that, The process of reparameterizing the 3x3 repConv unit into a single 3x3 Conv layer is specifically as follows: In the first convolutional path, first equivalent the 3x3 Avgpool layer to a special weight convolutional layer with all weight values being 1 / 9, and then merge the 1x1 Conv layer, BN layer, 3x3 Avgpool layer and BN layer to obtain a single 3x3 Conv layer; In the second convolutional path, pad the 1x3 Conv layer with 0s to a 3x3 Conv layer, and then merge it with the BN layer into a 3x3 Conv layer; In the third convolutional path, pad the 3x1 Conv layer with 0s to a 3x3 Conv layer, and then merge it with the BN layer into a 3x3 Conv layer; In the fourth convolutional path, merge the 1x1 Conv layer, BN layer, 3x3 Conv layer and BN layer into a 3x3 Conv layer; Merge the four 3x3 Conv layers into a single 3x3 Conv layer.
6. A system for detecting damage to the slag discharge belt of a roadheader, adopting the method for detecting damage to the slag discharge belt of a roadheader according to any one of claims 1 to 5, characterized in that, It includes: A sample acquisition module, which is used to acquire the surface image during the operation of the belt, and label the contour of the damaged area in the image to make a damage detection sample data set; A model construction module, which is used to construct a MaskRCNN damage recognition model based on the ResNet50 backbone network; A model training module, which is used to train the MaskRCNN damage recognition model with the damage detection sample data set until the model converges. Among them, during the model training process, the 3x3 convolutional layer that composes the residual module in the ResNet50 backbone network is replaced by a grouped convolutional module with lateral dense connections; A damage detection module, which is used to obtain the image of the surface of the belt to be detected, input it into the trained MaskRCNN damage recognition model, and output the damage detection result.
7. An electronic device, characterized in that, It includes a processor and a memory. A computer program is stored in the memory. The processor is used to execute the steps of the method according to any one of claims 1 to 5 by calling the computer program stored in the memory.
8. A computer-readable storage medium for storing a computer program for detecting damage to the slag conveyor belt of a roadheader, characterized in that, When the computer program runs on a computer, it executes the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Belt tearing detection method and system based on machine vision and deep learning
CN113682762A
Automobile surface scratch detection method based on improved Mask RCNN
CN112802005A
Feature extraction method based on target detection grouping residual structure
CN115294326A