Mine stone length detection method based on U-net

Through the U-net network model combined with the encoder and decoder structure, the rapid and accurate detection of stone length is achieved, and the low efficiency and safety hazards of stone size detection in mining operations are solved, and the accuracy and safety of detection are improved.

CN120339363APending Publication Date: 2025-07-18SICHUAN DUOWEI INTELLIGENT CLOUD VALLEY CO LTD +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510409333.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In mining operations, stone size detection relies on manual measurement efficiency and poses safety risks, especially when the stone size exceeds the plate feeder processing limit, it may damage the equipment. It is difficult for existing machine vision and deep learning technologies to quickly and accurately detect stone lengths in complex environments.

Method used

The mine stone length detection method based on U-net is adopted, and the global feature extraction module, convolution module, feature fusion module and self-attention mechanism module are combined with the encoder and decoder structure, and the stone feature information is extracted using multi-scale fusion features to construct a mine stone size detection network model, and the stone length is judged through semantic segmentation.

Benefits of technology

It improves the accuracy and safety of stone length detection, reduces false detection and missed detection, enhances the generalization ability and stability of the model, protects the plate feeder from damage, and improves the efficiency of mining operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339363A_ABST
    Figure CN120339363A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of mining industry, in particular to a mine stone length detection method based on U-net. Comprising the following steps: S1, acquiring an image data set of a feeding port of the plate feeding machine and marking the image data set; s2, constructing a mine stone size detection network model based on a U-net network; s3, based on the image data set obtained in the S1 and the mine stone size detection network model constructed in the S2, carrying out model training; and S4, inputting a test set of mine stone images into the network model trained in the S3 for prediction, and detecting the stone size according to a semantic segmentation result. According to the method, the accuracy of mine stone length detection is effectively improved, redundant information in feature extraction is reduced, local detail feature information of the image is effectively extracted by using multi-scale fusion features, and the generalization ability and stability of the model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mining, and specifically relates to a method for detecting the length of mine stones based on U-net. Background Art

[0002] In traditional mine operations, the size detection of stones mainly relies on manual measurement. This method is not only inefficient but also has safety hazards. Especially when the size of the stone exceeds the processing limit of the apron feeder, it may cause damage to the equipment, affecting production safety and efficiency.

[0003] With the development of industrial automation and mining technology, more and more research has begun to focus on using machine vision and deep learning technologies for automatic detection of mine stones. For example, existing research has proposed a method for detecting the size of ore blocks based on deep learning. This method uses a residual neural network structure to form a CSPDarkNet21 backbone feature extraction network under the Darknet framework, speeds up the model training and prediction by simplifying the feature layer, and uses CIOU to calculate the loss value to make the training more stable. In addition, there is also research that has proposed a method for segmenting conveyor belt ore images based on U-Net and Res_UNet models. These methods effectively overcome the problems of excessive parameters and gradient dispersion caused by the deepening of the network layers by improving the network structure, and improve the training speed and accuracy of the model.

[0004] Although existing research has made certain progress in ore image processing, in the actual mine operation environment, how to quickly and accurately detect the length of stones and how to apply these technologies to protect the apron feeder from damage still face challenges. Especially in a complex mine environment, the diversity of the size, shape and position of stones, as well as the change of lighting conditions, all bring difficulties to the automatic detection of stone length. Summary of the Invention

[0005] The present invention provides a method for detecting the length of mine stones based on U-net, aiming to be able to effectively segment the stones and accurately judge their lengths to protect the apron feeder from damage and improve the safety and efficiency of mine operations.

[0006] The present invention provides a method for detecting the length of mine stones based on U-net, including the following steps:

[0007] S1. Obtain the image dataset of the apron feeder inlet and perform annotation;

[0008] S2. Construct a mine stone size detection network model based on the U-net network; the mine stone size detection network model includes a global feature extraction module, a convolution module, a feature fusion module, a self-attention mechanism module, a downsampling layer, and an upsampling layer;

[0009] The global feature extraction module is used to capture local feature information in the image;

[0010] The convolution module is used to extract detailed feature information of the image shadow;

[0011] The feature fusion module is used to effectively fuse the global feature information and the detailed feature information of the image;

[0012] The self-attention mechanism module is used to enhance the model's ability to recognize stone features;

[0013] The mine stone size detection network model uses the convolution module to extract features, performs size transformation of the feature map through the downsampling layer and the upsampling layer, and finally outputs the segmentation mask of the stone;

[0014] S3. Based on the image dataset obtained in S1, and based on the mine stone size detection network model constructed in the input S2, the model is trained;

[0015] S4. Input the test set of the mine stone image into the network model trained in S3 for prediction, and detect the stone size according to the semantic segmentation result.

[0016] As a further improvement of the present invention, the S1 includes:

[0017] S11. Obtain a mine stone image dataset, which includes stone images of different sizes, shapes and positions, as well as corresponding annotation information;

[0018] S12. Scale each image in the dataset into an image patch of size m×n to obtain an image set;

[0019] S13. Perform model training on the obtained image set through k-fold cross-validation.

[0020] As a further improvement of the present invention, the S2 includes:

[0021] S21. The global feature extraction module is composed of consecutive convolution modules and downsampling layers, and is used to gradually extract the global feature information of the image; the convolution modules correspond to feature extraction at different levels, expand the input channel number to capture features from the shallow layer to the deep layer; the downsampling layer reduces the spatial dimension of the feature map through convolution operations, and at the same time increases the channel number to transfer more abstract global feature information to the decoder stage;

[0022] S22. The convolutional module consists of two groups of n×n convolutional layers. Each convolutional module first extracts features through an n×n convolutional kernel, followed by a BatchNorm2d layer and a Dropout layer, and the LeakyReLU activation function is introduced to introduce non-linearity; another n×n convolutional kernel extracts features again, followed by the same normalization and activation layers, and the outputs of the convolutional layers are fused to form a convolutional layer that can capture different features;

[0023] S23. After the global feature extraction module extracts rich detailed information of the image in the encoder stage, it outputs detailed feature information of different sizes to the feature fusion module. In the decoder stage, the detailed feature information is fused with the corresponding global feature information of the same convolutional layer again through skip connections to generate effective context guidance information;

[0024] S24. After the decoder stage, the size of the upsampled image feature map is obtained; a 1×1 convolutional layer is applied to the feature information of the last layer to change the number of channels of the output image, and then thresholding is performed through the Sigmoid function to obtain the final prediction map of the shadow distribution;

[0025] S25. The self-attention mechanism module includes a channel attention module, and the channel attention module further enhances the feature map through channel attention and spatial attention.

[0026] As a further improvement of the present invention, in S22, the m-th convolutional layer of each convolutional module in the network can be expressed as:

[0027]

[0028] where V is the input feature map, is the index of the m-th convolutional layer, w m is the convolutional kernel weight at the m-th step, b m is the corresponding bias term, and ReLU is the rectified linear activation function applied to the feature map after convolution; each convolutional module contains two consecutive convolutional operations, and the output of the entire convolutional module can be expressed as:

[0029]

[0030] and represent the weights of the first and second convolutional layers respectively, and represent the biases of the first and second convolutional layers respectively.

[0031] As a further improvement of the present invention, in S23, the formula of the feature fusion module is:

[0032] F fused=U(x, f) = Cat(F up (x), F skip )

[0033] F fused represents the features after passing through this convolutional module, Cat represents feature fusion, and F up (x) is the feature map obtained from the upsampling layer. The size of the feature map is doubled through nearest neighbor interpolation and the number of channels is adjusted through a 1x1 convolution. F skip is the feature map at the corresponding stage of the encoder and is directly passed to the decoder stage through a skip connection; x refers to the feature map of the decoder, f refers to the feature map of the encoder, and U(x, f) is the feature map in which x and f are fused.

[0034] As a further improvement of the present invention, in S25, the formula of the channel attention module is:

[0035]

[0036] where Q is the Query matrix, K is the Key matrix, V is the Value matrix, Z is the finally calculated attention score, and d k represents the dimension of each vector in the Key matrix.

[0037] As a further improvement of the present invention, S3 includes:

[0038] S31. Training stage: Use the deep learning framework Pytorch 1.7 to build the network architecture for remote sensing image shadow detection. The model uses the binary cross-entropy loss function and the Adam optimizer for hyperparameter setting, and saves the best model in each round of training;

[0039] S32. Stone detection performance evaluation metrics: Four semantic segmentation evaluation metrics of semantic segmentation, Accuracy, F1-score, Precision, and Recall, are used to evaluate the performance of the model for stone detection;

[0040] S33. The data obtained from the dataset is input into the network model for mine stone detection constructed in S2, and then the model is trained using the backpropagation algorithm.

[0041] As a further improvement of the present invention, in S31, the calculation formula of the binary cross-entropy loss function is:

[0042]

[0043] where N represents the total number of pixels in the image; y i represents the label value of the i-th pixel in the remote sensing image, with the positive class being 1 and the negative class being 0; yi ∈ {0, 1} represents the i-th pixel in the remote sensing image, and p(y i ) represents the probability that the i-th pixel is predicted as the positive class. m is the weight of the positive class samples. The binary weighted cross-entropy loss function is used to balance the positive and negative samples, calculate the loss between the prediction result and the actual result, and continuously optimize the parameters of the model through the backpropagation algorithm.

[0044] As a further improvement of the present invention, in S32, the calculation process of the four semantic segmentation evaluation indicators is as follows:

[0045] The formula for accuracy is as follows:

[0046]

[0047] The formula for precision is as follows:

[0048]

[0049] The formula for recall is as follows:

[0050]

[0051] The formula for F1-score is as follows:

[0052]

[0053] Among them, TP: True Positive, that is, the total number of positive class pixels correctly identified; FN: False Negative, that is, the positive class pixels misidentified as negative class pixels; FP: False Positive, that is, the negative class pixels misidentified as positive class pixels; TN: True Negative, that is, the number of negative class pixels correctly identified.

[0054] As a further improvement of the present invention, S4 includes:

[0055] For the stones obtained from the semantic segmentation result, determine their maximum length through the diagonal of the circumscribed rectangle, and simplify the calculation by building a physical model to obtain the actual length of the stone. The calculation formula is as follows:

[0056]

[0057] L actual represents the actual length of the stone, L image represents the length calculated from the language segmentation result, w represents the width of the actual feeding port, h1 represents the length of the actual feeding port, l1 represents the width of the upper base of the trapezoid of the feeding port in the image, l2 represents the width of the lower base of the trapezoid of the feeding port in the image, and H1 represents the length of the hypotenuse of the trapezoid of the feeding port in the image.

[0058] The beneficial effects of the present invention are as follows: This method can more effectively detect stones, suppress non-stone areas, reduce false detections and missed detections, especially in complex areas, the detection performance of stone areas is better. It effectively improves the accuracy of detecting the length of stones in mines, reduces redundant information in feature extraction, and effectively extracts local detailed feature information of images using multi-scale fusion features, enhancing the generalization ability and stability of the model, which is of great significance for improving the safety and efficiency of mine operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flowchart of the automatic detection method for the size of mine stones based on the U-net network of the present invention;

[0060] Figure 2 is an overall network framework diagram of the automatic detection method for the size of mine stones based on the U-net network of the present invention;

[0061] Figure 3 is a structural schematic diagram of the feature fusion model of the present invention;

[0062] Figure 4 is the prediction result of the model in the experiment of the present invention;

[0063] Figure 5 is a schematic diagram of calculating the true length of a stone by building a physical model in the present invention;

[0064] Figure 6 is the original image of testing stone detection in the present invention;

[0065] Figure 7 is the prediction result diagram of each group of the model in the present invention;

[0066] Figure 8 is a picture at the inlet of the apron feeder in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further describes the present invention in detail with reference to the accompanying drawings and embodiments.

[0068] The present invention relates to the real-time monitoring of stones from mines using machine vision technology. Through the U-net network, semantic segmentation of stones is performed to determine whether the length of the stones exceeds the safe processing range of the apron feeder, thereby protecting the apron feeder from damage and improving the safety and efficiency of mine operations.

[0069] Specifically, as Figures 1 to 4 shown, a method for detecting the length of mine stones based on U-net of the present invention includes the following steps:

[0070] S1. Obtain an image dataset of the inlet of the apron feeder and perform annotation.

[0071] Step S1 includes:

[0072] S11. Obtain a mine stone image dataset. The mine stone image dataset contains stone images of different sizes, shapes, and positions, as well as corresponding annotation information, which is used to train and verify the accuracy of the model.

[0073] S12. Perform data augmentation on the original dataset. In the existing mine stone image dataset, the image size resolution is 3840×2160. Each image in the dataset is scaled into an image patch of size m×n to obtain an image set; in this embodiment, it is scaled into small image patches of 256×256. In order to explore the true effectiveness of the model of the present invention, no redundant data augmentation is performed except for cropping the data.

[0074] S13. Obtain 1288 images in the image set, and perform model training through k-fold cross-validation for model evaluation.

[0075] S2. Construct a mine stone size detection network model based on the U-net network; the mine stone size detection network model includes a global feature extraction module, a convolution module, a feature fusion module, a self-attention mechanism module, a downsampling layer, an upsampling layer, and a final output layer;

[0076] The global feature extraction module is used to capture local feature information in the image;

[0077] The convolution module is used to extract detailed feature information of the image shadow;

[0078] The feature fusion module is used to effectively fuse the global feature information and detailed feature information of the image;

[0079] The self-attention mechanism module is used to enhance the model's ability to recognize stone features and improve the accuracy of segmentation;

[0080] The mine stone size detection network model uses the convolution module to extract features, performs size transformation of the feature map through the downsampling layer and the upsampling layer, and finally outputs the segmentation mask of the stone.

[0081] The mine stone size detection network model consists of two parts: an encoder (downsampling path) and a decoder (upsampling path). The encoder is responsible for gradually extracting the features of the image, while the decoder is responsible for restoring these feature maps to the original image size and outputting the segmentation mask of the stone. This structure allows the model to capture the global features of the image in the encoder and restore the detailed information of the image in the decoder, thus achieving accurate stone size detection. The model uses the encoder-decoder structure to extract global feature information through consecutive convolutional modules and downsampling layers, and realizes the size transformation and fusion of feature maps through upsampling and skip connections. Finally, the segmentation mask of the stone is output through the Sigmoid function to determine the maximum length of the stone.

[0082] Specifically, step S2 includes:

[0083] S21. The global feature extraction module consists of consecutive convolutional modules (Conv_Block) and downsampling layers (DownSample), and is used to gradually extract the global feature information of the image. These convolutional modules correspond to feature extraction at different levels, gradually expanding from the input channel number 3 to 1024 to capture features from the shallow layer to the deep layer. The downsampling layer reduces the spatial dimension of the feature map through convolutional operations while increasing the number of channels to transfer more abstract global feature information to the decoder stage. Such a design enables the encoder to capture the global feature information in the image and provides a basis for subsequent stone size detection. After each convolutional module (Conv_Block), a channel attention module (CBAM) is integrated. This module further enhances the feature map through channel attention and spatial attention, enabling the model to pay more attention to the key features of the stone and improving the accuracy and robustness of the model for stone length detection.

[0084] S22. The convolutional module consists of two groups of 3×3 convolutional layers with a stride of 1. Each Conv_Block first extracts the detailed feature information of the image shadow through a 3×3 convolutional kernel, followed by a batch normalization layer BatchNorm2d and a Dropout layer, and then a LeakyReLU activation function to introduce non-linearity. After that, another 3 × 3 convolutional kernel extracts features again, followed by the same normalization and activation layers. The convolutional kernel is a component of the convolutional layer. The outputs of these convolutional layers are fused to form a convolutional layer that can capture different features; the m-th convolutional layer of each Conv_Block in the network can be expressed as:

[0085]

[0086] where V is the input feature map, is the index of the m-th convolutional layer, w m is the convolutional kernel weight at the m-th step, b mis the corresponding bias term. The ReLU is a rectified linear activation function applied to the feature map after convolution, ensuring non-linear transformation and enhancing the feature extraction ability of the model. Since each Conv_Block contains two consecutive convolution operations, the output of the entire Conv_Block can be expressed as:

[0087]

[0088] and represent the weights of the first and second convolutional layers respectively, and represent the biases of the first and second convolutional layers respectively. The feature fusion module is used to effectively fuse the global feature information and detailed feature information of the image. In the U-Net model, such Conv_Blocks are reused. By outputting detailed feature information of different sizes to the feature fusion module, in the decoder stage, the detailed feature information is fused with the corresponding global feature information of the same convolutional layer again through skip connections, generating effective context guidance information.

[0089] Through convolution with a convolution kernel size of 3×3, a stride of 2, and a padding layer of 1, the role of the pooling layer is replaced to perform image downsampling and reduce the size of the input data. After four downsamplings, the pixel sizes of the image feature maps are [128×64×64], [256×32×32], [512×16×16], and [1024×8×8] respectively.

[0090] S23. After the global feature extraction module extracts rich detailed information of the image in the encoder stage, it outputs detailed feature information of different sizes to the feature fusion module. In the decoder stage, the detailed feature information is fused with the corresponding global feature information of the same convolutional layer again through skip connections, generating effective context guidance information. The decoder part fuses the feature map extracted by the encoder with the feature map of the decoder through upsampling and skip connections. The upsampling layer enlarges the size of the feature map, while the skip connection directly passes the feature map in the encoder to the corresponding decoder layer to achieve refined restoration of the features. This structure enables the decoder to not only restore the detailed information of the image but also utilize the global feature information in the encoder, improving the accuracy of stone size detection. F fused represents the feature after this convolution module. After feature fusion, it is expressed as:

[0091] F fused = U(x,f) = Cat(F up (x),F skip )

[0092] where Cat represents feature fusion, F up(x) is the feature map upsampled by the UpSample layer. The size of the feature map is doubled through nearest neighbor interpolation and the number of channels is adjusted by a 1x1 convolution. F skip is the feature map of the corresponding stage of the encoder, which is directly passed to the decoder stage through a skip connection. x refers to the feature map of the decoder, f refers to the feature map of the encoder, and U(x,f) is the feature map in which x and f are fused.

[0093] The fused feature is further normalized and passed through an activation function. The activation layer is implemented using the LeakyReLU function. The normalization layer and the activation layer are used to increase numerical stability and activate the output in a non-linear manner. The resulting image features will have four times the original number of channels. Therefore, preliminary fusion is required to reduce parameters and avoid overfitting to capture important information. In this embodiment, the multi-scale asymmetric inner convolution module consists of inner convolutions with inner convolution kernel sizes of 1×1, 1×3, 3×1, and 5×5 respectively. The stride is set to 1 for each. They are respectively passed through a normalization layer and an activation layer, and then the different convolution kernels are fused to form a multi-scale asymmetric inner convolution layer.

[0094] S24. After passing through the decoder stage, the sizes of the image feature maps after four upsamplings are [64×128×128], [128×64×64], [256×32×32], and [512×16×16] in sequence. A 1×1 convolution layer is applied to the feature information of the last layer to change the number of channels of the output image, and then thresholding is performed through the Sigmoid function to obtain the prediction map of the final shadow distribution.

[0095] Through skip connections, effective fusion of the global feature information and local information of the image is achieved to enhance the effective object context at an appropriate scale for shadow detection. In this way, the obtained image features can adapt to the detection of shadow regions at different scales. The pixel sizes of the image feature maps after four upsamplings are [512×16×16], [256×32×32], [128×64×64], and [64×128×128]. After the last upsampling, a 1×1 convolution layer is used and thresholding is performed through the Sigmoid function to obtain the prediction map of the final shadow distribution.

[0096] For the design of the final output layer, in the last stage of the decoder, a 1×1 convolution layer is used to change the number of channels of the output image to match the number of channels of the stone segmentation mask. Then, binary classification is performed on each pixel through the Sigmoid function, that is, to judge whether the length of the stone exceeds the safe handling range of the apron feeder. The output value range of the Sigmoid function is between 0 and 1, indicating the probability of each pixel belonging to the stone. Based on this probability, a threshold can be set to determine the pixels with a probability higher than the threshold as stones, thereby obtaining the segmentation mask of the stones.

[0097] S25. The self-attention mechanism module includes a channel attention (CBAM) module, which further enhances the feature map through channel attention and spatial attention, enabling the model to pay more attention to the key features of the stone blocks, and improving the accuracy and robustness of the model for detecting the length of the stone blocks.

[0098]

[0099] Among them, Q is the Query matrix, K is the Key matrix, V is the Value matrix, Z is the finally calculated attention score, and d k represents the dimension of each vector in the Key matrix. The Query matrix, Key matrix, and Value matrix respectively represent the query, key, and value matrices, which are extracted from the input features through linear transformation.

[0100] S3. Based on the image dataset obtained in S1, and based on the mine stone size detection network model constructed in the input S2, the model is trained.

[0101] Step S3 includes:

[0102] S31. Use the deep learning framework Pytorch 1.7 to build the network architecture for remote sensing image shadow detection. In the experiment, all models use the binary cross-entropy loss function and the Adam optimizer, and the hyperparameters are set (the initial learning rate is 0.001, the total number of iterations is 30, the batch size is 4, and the L2 weight regularization is 0.0005). The best model is saved in each round of training; among them, the calculation formula of the binary cross-entropy loss function is:

[0103]

[0104] Among them: N represents the total number of pixels in the image; y i represents the label value of the i-th pixel in the remote sensing image, the positive class is 1, and the negative class is 0; y i ∈{0,1} represents the i-th pixel in the remote sensing image, and p(y i ) represents the probability that the i-th pixel is predicted as the positive class. m is the weight coefficient used to balance positive and negative samples. m is the weight of the positive class samples, which is used to adjust the degree of attention of the model to the loss of positive class pixels. The value of m obtained from the experiment of the present invention is 0.65; the binary weighted cross-entropy loss function is adopted to balance positive and negative samples, calculate the loss between the prediction result and the actual result, and continuously optimize the parameters of the model through the backpropagation algorithm.

[0105] S32. Stone detection performance evaluation metrics: To intuitively and effectively analyze the shadow detection and segmentation accuracy of the proposed model, four semantic segmentation evaluation metrics commonly used in semantic segmentation, namely Accuracy, F1-score, Precision, and Recall, are adopted to evaluate the performance of the model in stone detection.

[0106] The formula for Accuracy is as follows:

[0107]

[0108] The formula for Precision is as follows:

[0109]

[0110] The formula for Recall is as follows:

[0111]

[0112] The formula for F1-score is as follows:

[0113]

[0114] Among them, TP: True Positive, that is, the total number of positive class pixels correctly identified. FN: False Negative, that is, the positive class pixels misidentified as negative class pixels. FP: False Positive, that is, the negative class pixels misidentified as positive class pixels; TN: True Negative, that is, the number of negative class pixels correctly identified. For the four semantic segmentation evaluation metrics, the higher the value, the better the detection result. However, for the actual application scenario, the focus is on the recall rate.

[0115] S33. The data obtained from the dataset is input into the network model for mine stone detection constructed in Step 2, and then the model is trained using the backpropagation algorithm;

[0116] S4. The test set of the mine stone images is input into the network model trained in S3 for prediction, and the stone size is detected based on the semantic segmentation result.

[0117] Specifically, for the stones obtained from the semantic segmentation result, the maximum length is determined by the diagonal of the circumscribed rectangle. Due to the camera shooting angle and the inclination angle of the actual plate feeder inlet, there is a deviation between the stone length calculated directly from the semantic segmentation result and the actual stone length. The true stone length can be obtained by simplifying the calculation through building a physical model, as shown in Appendix Figure 5 ; Its calculation formula is as follows:

[0118]

[0119] L actualRepresents the actual length of the stone block, L image Represents the length calculated from the language segmentation result. w represents the width of the actual feeding inlet, h1 represents the length of the actual feeding inlet, l1 represents the width of the upper base of the trapezoidal feeding inlet in the image, l2 represents the width of the lower base of the trapezoidal feeding inlet in the image, and H1 represents the length of the hypotenuse of the trapezoidal feeding inlet in the image.

[0120] The actual width of the apron feeder is w = 2.8 meters. Set OMU = One Meter Upper, OMB = One Meter Bottom, OMF = One Meter Reference.

[0121] S41. First, obtain the lengths of the one-meter lines at the inlet of the apron feeder and the outlet at the bottom of the picture, as shown in the appendix Figure 8 .

[0122] L1 / L = OMU / 1 meter;

[0123] OMU = L1×1 meter / L = L1×1 meter / 2.8 meters;

[0124] L2 / L = OMB / 1 meter;

[0125] OMB = L2×1 meter / L = L2×1 meter / 2.8 meters;

[0126] S42. Find the position of the large stone block. Appendix Figure 8 The reference length of the one-meter line in it: OMF = OMU+(OMB - OMU)×(h1 / h).

[0127] S43. Compare whether the diagonal length x of the circumscribed rectangle of the large stone block is greater than OMF. If it is greater, an alarm is issued.

[0128] Input the data-augmented dataset into the trained network model for image shadow detection, use the test pictures for prediction, perform shadow detection on the images, and obtain the mask map.

[0129] The effects of the present invention can be further illustrated by the following simulation experiments.

[0130] 1) Simulation experiment conditions:

[0131] The hardware platform for the simulation experiment of the present invention is: the processor is AMD Ryzen9 3900X, the main memory is 32GHz, the external memory is 2T, the graphics card is NVIDIA GeForce RTX3090, and the video memory is 24GB;

[0132] The software platform for the simulation experiment of the present invention is: Ubuntu 18.04 LTS operating system and python 3.7, CUDA11.5, CUDNN8.3.1;

[0133] 2) Simulation experiment content and result analysis:

[0134] The simulation experiment of the present invention is carried out according to the following steps respectively by using the method of the present invention and the ablation experiment method of the prior art;

[0135] The comparative experiment method is to train the mine stone detection network model based on U-net with Adam as the optimizer and binary cross-entropy as the loss function;

[0136] Under the same experimental environment, 100 iterations of training are completed. After 100 iterations, the model reaches the convergence state. Comparing two common semantic segmentation networks, FCN and U-net++, the evaluation index results are shown in Table 1. From the recall rate and other evaluation index data, it can be seen that the present model is more excellent in stone detection. The recall rates of the two common semantic segmentation networks are about 0.80 and 0.75 respectively, while the highest recall rate of the disclosed method of the present invention is 0.96 and it is stable at 0.90, with the recall rate increased by about 10%. Obviously, the method disclosed by the present invention performs best in the stone detection task;

[0137] Table 1 Evaluation indexes of the segmentation performance of each model

[0138]

[0139] Appendix Figure 6 is the original image for testing stone detection, Figure 7 (fcn_1_1), Figure 7 (fcn_1_2), Figure 7 (fcn_1_3), Figure 7 (fcn_1_4) and Figure 7 (fcn_1_5) are the prediction results of each group of the FCN model. Figure 7 (u2net_1_1), Figure 7 (u2net_1_2), Figure 7 (u2net_1_3), Figure 7 (u2net_1_4) and Figure 7 (u2net_1_5) are the prediction results of each group of the U-net++ model. Figure 7 (unet_1_1), Figure 7 (unet_1_2), Figure 7 (unet_1_3), Figure 7 (unet_1_4) and Figure 7 (unet_1_5) are the prediction results of each group of the U-net model. Figure 7 (ours_1_1), Figure 7 (ours_1_2), Figure 7(ours_1_3), Figure 7 (ours_1_4) and Figure 7 (ours_1_5) are the prediction results of each group of the model.

[0140] The method of the present invention detects more stones, suppresses non-stone areas, and reduces more false detections and missed detections. For some stone areas that are easily flooded by complex areas, the method of the present invention has better detection performance for these areas.

[0141] The beneficial effects of the present invention are as follows:

[0142] First, by using the U-net network model, the present invention effectively combines the encoder and decoder structures, realizes feature extraction from shallow to deep layers, and feature restoration from deep to shallow layers, improving the accuracy of detecting the length of mine stones.

[0143] Second, the present invention constructs a convolutional module. The cross-shaped receptive field can reduce redundant information in capturing representative features, and effectively extract local detail feature information of the image by using multi-scale fusion features to form effective context information.

[0144] Third, the U-net model of the present invention uses standard convolutional operations to extract image features. These operations have spatial specificity, can capture local features, and effectively transfer long-distance pixel dependencies through skip connections. To reduce redundant information between channels, the model introduces batch normalization and Dropout layers in the Conv_Block, which not only enhances the generalization ability of the model but also helps improve the stability of the model. In the encoder stage, the model uses a 3x3 convolutional kernel and a downsampling layer (Down Sample) with a stride of 2 to replace the traditional pooling layer. This design reduces the size of the feature map, thereby improving the computational efficiency of the network, reducing the number of parameters, and accelerating the convergence speed of the model. In addition, to prevent the loss of low-resolution feature information during the image size reduction process, the model performs downsampling through a convolutional operation with a stride of 2, effectively maintaining the integrity of the stone detail features. The present invention further adds a self-attention mechanism. Through the Channel Attention Module and the Spatial Attention Module, the recognition ability of the model for stone features is enhanced, and the accuracy of segmentation is improved.

[0145] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for detecting the length of mine stones based on U-net, characterized in that, It includes the following steps: S1. Obtain the image dataset at the feeding port of the apron feeder and perform annotation; S2. Construct a mine stone size detection network model based on the U-net network; the mine stone size detection network model includes a global feature extraction module, a convolution module, a feature fusion module, a self-attention mechanism module, a downsampling layer, and an upsampling layer; The global feature extraction module is used to capture local feature information in the image; The convolution module is used to extract detailed feature information of the image shadow; The feature fusion module is used to effectively fuse the global feature information and the detailed feature information of the image; The self-attention mechanism module is used to enhance the model's ability to recognize stone features; The mine stone size detection network model extracts features using the convolution module, performs size transformation of the feature map through the downsampling layer and the upsampling layer, and finally outputs the segmentation mask of the stone; S3. Based on the image dataset obtained in S1, train the model based on the mine stone size detection network model constructed in S2; S4. Input the test set of the mine stone image into the network model trained in S3 for prediction, and detect the stone size according to the semantic segmentation result.

2. The method for detecting the length of mine stones based on U-net according to claim 1, characterized in that, The S1 includes: S11. Obtain the mine stone image dataset, which contains stone images of different sizes, shapes, and positions, as well as corresponding annotation information; S12. Scale each image in the dataset into an image patch of size m×n to obtain an image set; S13. Perform model training on the obtained image set through k-fold cross-validation.

3. The method for detecting the length of mine stones based on U-net according to claim 1, wherein The S2 includes: S21. The global feature extraction module consists of consecutive convolution modules and a downsampling layer, and is used to gradually extract the global feature information of the image; the convolution modules correspond to feature extraction at different levels, expand the input channel number to capture features from the shallow layer to the deep layer; the downsampling layer reduces the spatial dimension of the feature map through convolution operations, and at the same time increases the channel number to transfer more abstract global feature information to the decoder stage; S22. The convolution module consists of two groups of n×n convolutional layers. Each convolution module first extracts features through an n×n convolution kernel, a batch normalization layer BatchNorm2d and a Dropout layer, and the LeakyReLU activation function introduces non-linearity; another n×n convolution kernel extracts features again, followed by the same normalization and activation layers, and fuses the output of the convolutional layers to form a convolutional layer that can capture different features; S23. After the global feature extraction module extracts rich detailed information of the image in the encoder stage, it outputs detailed feature information of different sizes to the feature fusion module, and in the decoder stage, the detailed feature information is fused with the corresponding global feature information of the same convolutional layer again through skip connections to generate effective context guidance information; S24. After passing through the decoder stage, obtain the size of the image feature map after multiple upsamplings; apply a 1×1 convolutional layer to the feature information of the last layer to change the channel number of the output image, and then perform thresholding through the Sigmoid function to obtain the prediction map of the final shadow distribution; S25. The self-attention mechanism module includes a channel attention module, which further enhances the feature map through channel attention and spatial attention.

4. The method for detecting the length of mine stones based on U-net according to claim 3, characterized in that, In S22, the m-th convolutional layer of each convolutional module in the network can be expressed as: where V is the input feature map, m is the index of the m-th convolutional layer, w m is the convolutional kernel weight at the m-th step, b m is the corresponding bias term, and ReLU is the rectified linear activation function applied to the feature map after convolution; each convolutional module contains two consecutive convolutional operations, and the output of the entire convolutional module can be expressed as: and represent the weights of the first and second convolutional layers respectively, and represent the biases of the first and second convolutional layers respectively.

5. The method for detecting the length of mine stones based on U-net according to claim 3, characterized in that In S23, the formula of the feature fusion module is: F fused = U(x, f) = Cat(F up (x), F skip ) F fused represents the features after passing through this convolutional module. Cat represents feature fusion, and F up (x) is the feature map obtained from the upsampling layer. The size of the feature map is doubled through nearest neighbor interpolation and the number of channels is adjusted through a 1x1 convolution. F skip is the feature map at the corresponding stage of the encoder and is directly passed to the decoder stage through a skip connection; x refers to the feature map of the decoder, f refers to the feature map of the encoder, and U(x,f) is the feature map in which x and f are fused.

6. The method for detecting the length of mine stones based on U-net according to claim 3, characterized in that, In S25, the formula of the channel attention module is: Among them, Q is the Query matrix, K is the Key matrix, V is the Value matrix, Z is the finally calculated attention score, and d k represents the dimension of each vector in the Key matrix.

7. The method for detecting the length of mine stones based on U-net according to claim 1, wherein S3 includes: S31. Training phase: Build the network architecture for remote sensing image shadow detection using the deep learning framework Pytorch 1.

7. The model uses the binary cross-entropy loss function and the Adam optimizer for hyperparameter setting, and saves the best model in each round of training. S32. Stone detection performance evaluation metrics: Four semantic segmentation evaluation metrics of semantic segmentation, Accuracy, F1-score, Precision, and Recall, are used to evaluate the performance of the model for stone detection. S33. The data obtained from the dataset is input into the network model for mine stone detection constructed in S2, and then the model is trained using the backpropagation algorithm.

8. The method for detecting the length of mine stones based on U-net according to claim 7, characterized in that, In S31, the calculation formula of the binary cross-entropy loss function is: where N represents the total number of pixels in the image; y i represents the label value of the i-th pixel in the remote sensing image, with the positive class being 1 and the negative class being 0; y i ∈{0,1} represents the i-th pixel in the remote sensing image, p(y i ) represents the probability that the i-th pixel is predicted as the positive class, and m is the weight of the positive class samples; The binary weighted cross-entropy loss function is used to balance the positive and negative samples, calculate the loss between the prediction result and the actual result, and continuously optimize the parameters of the model through the backpropagation algorithm.

9. The method for detecting the length of mine stones based on U-net according to claim 7, characterized in that In S32, the calculation process of the four semantic segmentation evaluation metrics is: The formula of Accuracy is as follows: The formula of Precision is as follows: The formula of Recall is as follows: The formula of F1-score is as follows: Among them, TP: True Positive, that is, the total number of positive class pixels correctly identified; FN: False Negative, that is, the positive class pixels misidentified as negative class pixels; FP: False Positive, that is, the negative class pixels misidentified as positive class pixels; TN: True Negative, that is, the number of negative class pixels correctly identified.

10. The method for detecting the length of mine stones based on U-net according to claim 1, wherein, S4 includes: For the stones obtained from the semantic segmentation results, determine their maximum length through the diagonal of the circumscribed rectangle, and simplify the calculation by building a physical model to obtain the true stone length. The calculation formula is as follows: L actual represents the actual length of the stone block, L image represents the length calculated from the language segmentation result, w represents the width of the actual feeding port, h1 represents the length of the actual feeding port, l1 represents the width of the upper base of the trapezoid of the feeding port in the image, l2 represents the width of the lower base of the trapezoid of the feeding port in the image, and H1 represents the length of the hypotenuse of the trapezoid of the feeding port in the image.

Citation Information

Cited By

  • JPEG compression artifact learning module and image tampering detection system

    CN121121426A

  • A JPEG compression artifact learning module and image tamper detection system

    CN121121426B