Broken rock granularity identification method based on machine vision

By introducing a deformable convolution module, hollow convolution and attention residual shrinking unit that integrates attention into the backbone network, the problems of low detection accuracy and large calculation amount in gravel particle size recognition are solved, and efficient and accurate gravel particle size recognition is achieved.

CN120088642AActive Publication Date: 2025-06-03SOUTHWEST PETROLEUM UNIV

Patent Information

Application Number
CN202510066176.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-03
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art has low detection accuracy and excessive calculation amount in the recognition of gravel particle size, making it difficult to effectively deal with noise and complex images.

Method used

A method for identifying broken rocks based on machine vision is proposed. By introducing a deformable convolution module with fused attention into the backbone network, hollow convolution is used to expand the receptive field, and combined with an efficient attention residual shrinking unit, the accuracy and efficiency of recognition are improved.

Benefits of technology

It significantly improves the accuracy and efficiency of gravel particle size recognition, reduces calculation complexity and resource consumption, enhances the adaptability to different lighting and environmental conditions, and achieves faster and more accurate particle size analysis of gravel piles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088642A_ABST
    Figure CN120088642A_ABST
Patent Text Reader

Abstract

The invention discloses a broken rock granularity identification method based on machine vision. The method comprises the following steps: S1, obtaining a broken rock pile image in a quarry; s2, based on the lightweight edge segmentation model, combining a deformable convolution module fusing attention and an attention residual shrinkage unit, and replacing a maximum pooling layer with a cavity convolution layer to obtain an improved lightweight edge segmentation model; s3, segmenting the macadam pile image, and distinguishing macadam edges and backgrounds in the macadam pile image to obtain a macadam edge graph; s4, carrying out edge refinement on the broken stone edge graph by adopting a table look-up method, and carrying out area filtering by combining an area filter so as to refine edges and eliminate interference edges; and S5, carrying out macadam edge contour extraction on the macadam edge graph, calculating the area and perimeter of the macadam, and finally obtaining the granularity distribution condition of the macadam. According to the detection method, the calculation amount is remarkably reduced while high accuracy is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and in particular to a method for identifying the particle size of crushed rocks based on machine vision. Background Art

[0002] At present, a large number of research results have been achieved in the application research of visual technology for crushed stone particle size segmentation. These studies have not only improved the efficiency of crushed stone processing, but also provided strong support for the intelligent development of the crushed stone particle size detection industry. Scholars have mainly carried out research and exploration in three aspects. The first is an image algorithm for accurately detecting the outline of crushed stones; the second is a reasonable evaluation model that converts the two-dimensional information of crushed stones into three-dimensional information and then evaluates the particle size distribution; the third is new technologies including neural networks, deep learning, and genetic algorithms. The above methods help to understand the crushed stone particle size distribution in different scenarios, but most of them rely too much on the low-order visual information of image pixels and manual debugging to select the best threshold parameters. When the image has problems such as high noise, high complexity, and poor contrast, it is still relatively troublesome to process. Coupled with the inevitable human influence during the segmentation process, it will lead to too large segmentation errors and poor generalization performance of the segmentation model.

[0003] With the increasingly mature application of artificial intelligence technology in the field of crushed stones, the focus of detecting the crushed stone particle size distribution in crushed stone pile images has gradually shifted towards using deep learning methods to achieve. For example, on the basis of DUNet, a residual structure is introduced to propose a new crushed stone image segmentation model (RDU-Net), which has a significant improvement in segmentation accuracy and better generalization ability; in order to extract the features of crushed stones of different sizes and shapes, residual deformable convolutions are introduced into the HED model, so as to learn the image features of different sizes and shapes, and achieve accurate segmentation of crushed stones, further verifying the feasibility of using deep learning image processing for crushed stone pile particle size. However, the above deep learning methods still have problems of low accuracy and excessive computational complexity. Summary of the Invention

[0004] In order to solve the problems of low detection accuracy and excessive computational complexity existing in the prior art when identifying the particle size of crushed stones, the present invention proposes a method for identifying the particle size of crushed rocks based on machine vision. A deformable convolution module with fused attention is introduced on the backbone network, and dilated convolution is used to expand the receptive field, and an efficient attention residual shrinkage unit is combined to improve the noise processing ability, thereby improving the accuracy and efficiency of crushed stone particle size identification to solve the above problems.

[0005] The present application discloses a method for identifying the particle size of crushed rocks based on machine vision, including the following steps: S1. Obtain an image of a crushed stone pile in a quarry; S2. Based on the lightweight edge segmentation model, combine the deformable convolution module with fused attention and the attention residual shrinkage unit, and use the dilated convolution layer to replace the max pooling layer to obtain an improved lightweight edge segmentation model; S3. Use the improved lightweight edge segmentation model obtained in S2 to segment the gravel pile image, distinguish the gravel edge and the background in the gravel pile image, and obtain the gravel edge map; S4. Adopt the table lookup method to refine the edges of the gravel edge map, and combine the area filter for area filtering to refine the edges and eliminate the interfering edges; S5. Extract the gravel edge contour of the gravel edge map, calculate the area and perimeter of the gravel, and finally obtain the particle size distribution of the gravel.

[0006] Preferably, the gravel pile images are from different scenarios including construction sites and mines, and are taken under different lighting conditions, angles and backgrounds.

[0007] Preferably, the deformable convolution module with fused attention is composed of a deformable convolution module in parallel with a hybrid attention module.

[0008] Preferably, the attention residual shrinkage unit uses one-dimensional convolution to replace the soft threshold of the residual shrinkage unit, and introduces an attention mechanism with spatial weights at the identity mapping of the residual shrinkage unit.

[0009] Preferably, the structure of the improved lightweight edge segmentation model is as follows: Add the deformable convolution module with fused attention before the second module and the third module of the lightweight edge segmentation model respectively, add the attention residual shrinkage unit before the context-aware fusion block of the lightweight edge segmentation model, and use the dilated convolution layer to replace the max pooling layer of the lightweight edge segmentation model.

[0010] Preferably, S3 includes the following steps: S31. Perform data annotation, data preprocessing and data augmentation on the obtained gravel pile image dataset; S32. Divide the dataset obtained in S31 into a training set and a test set, and use the data in the training set to train the improved lightweight edge segmentation model, and select the Dice coefficient loss as the loss function; S33. Use the trained model to segment the images in the test set, and evaluate the model performance by calculating the difference between the segmentation result and the ground truth mask. The evaluation metrics include pixel accuracy, IoU and Dice coefficient; S34. Apply the improved lightweight edge segmentation model to the actual gravel pile monitoring task.

[0011] Preferably, the edge refinement of the gravel edge map using the table lookup method includes the following steps: Perform binarization on the gravel edge map to obtain a binary image, and count each edge pixel point and its neighborhood in the binary image; Traverse the entire image to divide the edge pixel points into skeleton points and boundary points, retain the skeleton points and delete the boundary points to obtain an edge skeleton map.

[0012] Preferably, the area filtering in combination with the area filter includes the following steps: Invert the edge skeleton map obtained according to the table lookup method, and count the areas of all connected regions in the inverted image and the number of connected regions ; Set a threshold , if the area of the connected region is less than this threshold, that is , then this connected region is inverted, and finally all edge branches in the edge skeleton are removed.

[0013] Preferably, S5 includes the following steps: Based on the gravel edge map, use the findContours function in OpenCV to accurately find the gravel contours in the image, so that each gravel forms an independent connected region; Use the minAreaRect function to obtain the parameters of the minimum circumscribed rectangle of each gravel; Calculate the area of each connected region through the contourArea function, calculate the perimeter of each connected region through the arcLength function, and calculate the diameter of the connected region through the area and perimeter of the connected region; According to the obtained pixel sizes of the area, perimeter and diameter, convert them into the actual sizes of the gravel according to the proportionality coefficient, and count the number of gravels in each particle size level to obtain the particle size distribution of the gravel pile.

[0014] Advantages of the present invention: (1) Lightweight neural network design: The lightweight neural network adopted by the present invention significantly reduces the computational complexity and resource consumption of the model, making the gravel particle size detection more efficient. Compared with traditional heavy networks, the detection method of the present invention significantly reduces the amount of calculation while maintaining high accuracy.

[0015] (2) Optimization of image segmentation algorithm: The present invention improves the accuracy of gravel pile image segmentation by optimizing the image segmentation algorithm. Under the same conditions, the segmentation accuracy is significantly improved compared with the prior art.

[0016] (3) Real-time monitoring ability: The present invention can complete the particle size analysis of the gravel pile image in a short time, while the prior art takes too long.

[0017] (4)Environmental adaptability: The network model of the present invention has better adaptability to different lighting and weather conditions. In the experiments simulating different environmental conditions, the system of the present invention has maintained stable performance under various conditions, while the performance of the prior art has decreased significantly under certain conditions.

[0018] (5)Easy to deploy and maintain: Due to the lightweight of the model, the system of the present invention is easier to deploy in resource-constrained environments and has lower maintenance costs. Description of the Drawings

[0019] Figure 1 Flow chart of the method for identifying the particle size of broken rocks based on machine vision according to the embodiment of the present invention; Figure 2 Schematic structural diagram of the deformable convolution module according to the embodiment of the present invention; Figure 3 Schematic diagram of the convolution sampling position according to the embodiment of the present invention; Figure 4 Schematic structural diagram of the deformable convolution module integrating attention according to the embodiment of the present invention; Figure 5 Schematic structural diagram of the residual shrinkage unit according to the embodiment of the invention; Figure 6 Schematic structural diagram of the attention residual shrinkage unit according to the embodiment of the present invention; Figure 7 Schematic diagram of dilated convolutions with different dilation coefficients according to the embodiment of the present invention; Figure 8 Schematic structural diagram of the improved lightweight edge segmentation model according to the embodiment of the present invention; Figure 9 Flow chart of edge refinement according to the embodiment of the present invention; Figure 10 Flow chart of area filtering according to the embodiment of the present invention; Figure 11 Schematic diagram of the result of post-processing the edge image according to the embodiment of the present invention; Figure 12 Schematic diagram of extracting the gravel characteristic parameters according to the embodiment of the present invention; Figure 13 Schematic diagram of the qualitative analysis of the example images of the gravel pile data set of each algorithm model according to the embodiment of the present invention; Figure 14 Schematic diagram of the qualitative analysis of low-quality images according to the embodiment of the present invention; Figure 15 Curves of the training losses of each algorithm according to the embodiment of the present invention; Figure 16 Curves of the training accuracies of each algorithm during training according to the embodiment of the present invention. Detailed Embodiments

[0020] To make the purpose, technical solution and advantages of this application more clear and understandable, the following examples are given with reference to the accompanying drawings to further elaborate on this application in detail.

[0021] An embodiment of this application discloses a method for identifying the particle size of crushed rocks based on machine vision, and its process is as Figure 1 shown, including the following steps: S1. Obtain the image of the crushed rock pile in the quarry. The primary task of segmenting the crushed rock pile image is to prepare appropriate training data. The crushed rock pile images come from different scenarios, including construction sites, mines, etc., and are taken under different lighting conditions, angles, and backgrounds to obtain a diverse dataset.

[0022] S2. Since the crushed rocks in the crushed rock pile image are severely stacked and adhered, their shapes and sizes are different and the textures are complex and chaotic after being crushed. Coupled with the harsh image acquisition environment and uneven lighting, the contrast between the crushed rocks and the image background is poor. All these factors may make it difficult for the segmentation result to meet the ideal requirements. When directly using the traditional lightweight edge segmentation model (Lightweight Dense CNN for Edge Detection, LDC) to perform edge segmentation on the crushed rock pile image, there are a large number of pseudo-edges and noise points in the segmented result image, and some detail information is lost.

[0023] Based on this, this embodiment proposes an improved lightweight edge segmentation model. This model introduces a deformable convolution module with fused attention in the backbone network to accurately learn the detailed features of the crushed rocks in the crushed rock pile. Considering that the spatial acuity will decrease due to the expansion of the receptive field, dilated convolution is used to replace the conventional pooling layer to achieve the expansion of the receptive field and eliminate the interference of pseudo-edges without affecting the image resolution. At the same time, before the fusion stage of the four intermediate edge maps, the attention residual shrinkage unit is used to solve the problem of noise interference and improve the feature recognition ability. Therefore, the improved lightweight edge segmentation model aims to avoid as many false edges and noise interferences as possible on the premise of achieving high-precision determination of the crushed rock shape, thereby enhancing the edge segmentation effect of the crushed rock pile image.

[0024] As Figure 2As shown, the core components of deformable convolution include an offset convolution layer, a standard convolution layer, a batch normalization layer, and an activation layer. The use of the offset convolution layer enables it to better adapt to local changes in the input feature map. The standard convolution layer performs convolution operations on the adjusted feature map, which can more effectively extract the features of rocks, improving the accuracy and robustness of the features. The batch normalization layer can reduce the internal covariate shift of the feature map, making the feature extraction process more stable and improving the generalization ability of the model, especially when dealing with rock images under different lighting and background conditions. The non-linear transformation of the activation layer can enhance the model's ability to express rock features, enabling it to better capture the complex shapes and texture features of rocks. Since the standard convolution kernel has poor adaptability and generalization performance when dealing with gravel with complex shapes and cannot effectively extract the detailed features of gravel, deformable convolution introduces a learnable directional offset in the convolution kernel. The introduction of this offset can give the convolution kernel a large variable range during training, making the convolution area close to the shape of the target and more accurately extracting the target feature information. Deformable convolution introduces an offset on the basis of standard convolution to achieve dynamic adjustment of the convolution kernel. The sampling points of the convolution kernel will adaptively change with the progress of learning, so as to be able to learn the information of gravel images with different sizes and shapes.

[0025] As Figure 3 shown, it shows the difference in the sampling areas of standard convolution and deformable convolution. Taking a convolution kernel with a size of 7×7 as an example, Figure 3 in (a) represents standard convolution, Figure 3 in (b) represents deformable convolution, and the offset is obtained by applying the offset convolution layer. The offset convolution layer should have the same spatial resolution and padding rate as the standard convolution layer and the same number of channels as the input feature map to ensure that each channel has corresponding offset information while extracting features from the offset feature map. Figure 3 The standard convolution in (a) is evenly distributed according to the predefined network pattern. Such a sampling method has good effects when dealing with images with regular structures, but may not be able to effectively capture these detailed features when dealing with complex geometric shapes and irregularly distributed features. Figure 3 The sampling positions of the deformable convolution in (b) can be dynamically adjusted, which can flexibly adapt to local changes in the input image, thus better capturing the edge, corner, and texture features of rocks.

[0026] Since deformable convolution needs to calculate the offset additionally, it may introduce irrelevant context information, resulting in insufficiently close connections between context information. As Figure 4As shown, to overcome this defect, in this embodiment, a hybrid attention mechanism is introduced in parallel in the deformable convolution module, enabling the model to search for features globally and focus on key information. The hybrid attention module is a lightweight module that can improve the model performance with only a few parameters. Specifically, the hybrid attention module extracts features in the channel region by applying one-dimensional convolution, eliminating the adverse effects of dimensionality reduction operations on channel learning. In terms of the spatial region, global max pooling and average pooling techniques are adopted. After concatenation, dimensionality reduction, normalization, and dot product with channel output features are performed to obtain the overall weight distribution. The application of global max pooling and global average pooling helps to extract the most representative and global features in the rock image, which is crucial for understanding the overall characteristics and distribution patterns of the rock. Using the concatenation technique, the feature vector can contain local and global information of the rock image, helping to more comprehensively describe the features of the rock. After applying dimensionality reduction and normalization operations, the number of model parameters is effectively reduced, improving the stability, convergence speed, and generalization performance of the model. The deformable convolution module with fused attention can not only enhance the adaptability and flexibility of convolution but also enhance the ability of convolution to extract context information. The deformable convolution module with fused attention can be embedded into any existing convolutional neural network without modifying the network structure or training process.

[0027] In the harsh working environment of quarries, due to the diffusion of dust, uneven light distribution, and densely distributed holes on the surface of crushed stones, the images of crushed stone piles are often filled with a large amount of noise and redundant information, which will have an adverse impact on edge segmentation. The Residual Shrinkage Building Unit (RSBU) aims to eliminate the adverse effects of noise on the model, and its basic structure is as Figure 5As shown. The core of the residual shrinkage unit lies in soft thresholding. The soft thresholding function is a non-linear mapping that can set the feature values within a certain range to zero, thereby eliminating redundant or noise-related features. Specifically, the convolutional neural network first learns and masters a set of tiny positive thresholds, and then the soft threshold function sets the input data to zero when the absolute value of the weight is less than the threshold, and shrinks it towards zero when it is greater, fitting the unimportant redundant information near zero, that is, soft thresholding. This method helps to highlight the key features in the rock image while removing irrelevant information and improving the recognition accuracy. Theoretically, since the contribution of noise to edge segmentation is small, the corresponding weight values are low and thus set to 0, reducing the interference of noise on the model. The channel attention mechanism allows the network to adaptively emphasize or suppress different feature channels, thereby enhancing the most representative features in the rock image. This mechanism enables the network to pay more attention to the features that contribute to rock recognition while ignoring those unimportant features by learning the importance weights of each channel. The residual shrinkage unit mainly includes the soft threshold method and the channel attention mechanism. It first receives the input data, obtains the preliminary feature representation through feature extraction and residual learning, then refines the features according to the set threshold, and removes redundant information through soft threshold processing. Specifically expressed as:

[0028]

[0029]

[0030] Among them, represents the output value with the number of channels being ; represents the input value with the number of channels being ; represents the threshold with the number of channels being ; is the scale factor; is the output of the fully connected layer; is the height of the input feature map; is the width of the input feature map.

[0031] By using the residual shrinkage unit, a threshold can be adaptively set for each feature during the feature extraction stage, and this threshold can be continuously optimized through learning. The soft threshold method sets the features between the feature values to 0, thereby achieving the goal of reducing data noise.

[0032] In view of the fact that the two fully connected layers involve dimensionality reduction operations when learning the soft threshold and have low propagation efficiency, in order to improve the efficiency of the neural network and reduce the consumption of computing resources, this embodiment has made improvements on the basis of the original residual shrinkage unit structure, such asFigure 6 As shown, one-dimensional convolution is used as an alternative to implement the function of the network learning soft threshold. Such an improvement can better utilize the interaction between multiple different feature layers and effectively avoid the influence brought by dimensional reduction. In addition, according to its function, it can be regarded as a special attention mechanism, but it only focuses on the weight situation in the channel domain and does not consider the weight distribution in the spatial domain, that is, the importance of different spatial positions, which may lead to the loss of some information or the interference of noise. Therefore, in this embodiment, an attention mechanism with spatial weights is introduced at the identity mapping of the residual shrinkage unit, so that the improved attention residual shrinkage unit can take into account the feature recognition in both the channel domain and the spatial domain, and only slightly increase the parameters and computational complexity. It significantly improves the effect of gravel feature extraction and recognition, enhances key features, assigns higher weights to the key feature regions in the gravel image, enables the network to better capture the significant features of gravel, and also suppresses the interference of irrelevant information and irrelevant background and noise information in the feature extraction of gravel. At the same time, there are significant improvements in both the accuracy and robustness of recognition.

[0033] Due to the various shapes and complex textures of gravel, models with excellent detection results for normal images all show the phenomenon of incomplete recognition of the gravel edge in the gravel pile image and are seriously interfered by pseudo-edges. Given that the max pooling operation only retains the maximum value in each region, and it will change the positional relationship of the feature map, resulting in the model losing some valuable feature information and reducing its adaptability to translational invariance. The application of the max pooling operation may have certain negative impacts, leading to misjudging non-edge points in the gravel pile image as edge points, thus increasing the probability of pseudo-edges appearing. It should be noted that the main goal of the pooling layer is to reduce the dimension, and its calculation process is basically similar to that of the convolutional layer. Since the pooling layer error will be directly fed back to the convolutional layer, this may affect the ability of the edge segmentation module to learn features, especially for the learning of high-level semantic features.

[0034] As the convolution kernel size and network depth continue to increase, the receptive field will also increase accordingly, which helps to extract image features. However, this will lead to a rapid increase in the number of parameters and computational complexity, bringing a great burden to the calculation. The dilated convolution can be used to solve this problem. As Figure 7 shown, the dilated convolution adjusts the interval between the internal elements of the convolution kernel by introducing a dilation coefficient, thereby expanding the receptive field without increasing the number of parameters of the network model. The size of the dilated convolution kernel is expressed as:

[0035] where is the final convolution kernel size, is the kernel size, is the dilation coefficient.

[0036] Therefore, in this embodiment, a dilated convolutional layer is used as an alternative to the pooling layer to achieve a similar dimensionality reduction effect as the pooling layer. Moreover, by adjusting the dilation rate, the convolutional kernel can move on the input image with a larger stride, thereby expanding the receptive field and enabling the network to capture a larger range of context information without sacrificing resolution. Compared with the traditional pooling layer, dilated convolution can extract features without reducing the resolution of the feature map. This is crucial for maintaining the integrity of the gravel detail features, as these detail features are very important for the accurate identification and classification of gravel. Dilated convolution can be combined with different dilation rates to achieve multi-scale feature extraction. This enables the network to capture both the local details and global features of the gravel simultaneously, improving the accuracy and robustness of the identification. During the training process, by inversely modifying the weight parameters, the network can better learn the features under a larger receptive field to improve the detection accuracy of details such as image edges.

[0037] To solve the difficulties existing in the gravel edge segmentation process, in this embodiment, an improved lightweight edge segmentation model with faster convergence and higher accuracy is proposed based on the traditional LDC model. The overall structure after improvement is as Figure 8 shown. Generally speaking, the improvement of the LDC model includes the following aspects: Add a deformable convolutional module with fused attention, denoted as "AD". It is added before the second and third modules and shares information weights with the deep convolution through skip connections.

[0038] Add an attention residual shrinkage unit, denoted as "ARSBU". An attention residual shrinkage unit is introduced before the feature map fusion in the context-aware fusion block to eliminate the influence of noisy or redundant data, enabling the model to obtain a more accurate gravel feature fusion map.

[0039] Replace the max pooling layer with a dilated convolutional layer, denoted as "DC". The pooling layers in the second and third modules of the LDC model are replaced by dilated convolutions with a stride of 2, a convolutional kernel size of 3×3, and dilation coefficients of 2 and 3 respectively.

[0040] The network architecture of the obtained improved lightweight edge segmentation model is as follows: Input layer: The input layer receives preprocessed image data, usually a normalized color image. The size of the input image is adjusted according to the task requirements.

[0041] Convolutional layer and pooling layer: The network extracts image features through multiple convolutional layers. After each layer of convolution, a pooling layer follows immediately to gradually reduce the spatial dimension of the feature map.

[0042] Dense Connection Module: The core feature of the network lies in its dense connection structure. The output of each convolutional layer serves not only as the input to the next convolutional layer but also connects to the outputs of all previous layers. This can effectively avoid information loss and enable the network to make full use of the features learned by each layer, thereby improving the segmentation accuracy. The output of each layer is passed to the downstream module to form dense feature fusion.

[0043] Lightweight Convolution Module: The network adopts lightweight technologies such as depthwise separable convolution, splitting the traditional convolution operation into two independent convolution processes (depth convolution and pointwise convolution). This can significantly reduce the number of parameters and improve the computational efficiency, enabling the network to run on devices with limited resources.

[0044] Decoder Module: The decoder part contains upsampling operations, aiming to restore the low-resolution feature map to a segmentation map of the same size as the input image. Through upsampling, the network can generate accurate pixel-level predictions to distinguish the rubble pile and background regions.

[0045] Output Layer: The final output is a segmentation mask of the same size as the input image, and the value of each pixel indicates whether the pixel belongs to the rubble pile area or the background area. The output layer uses the softmax activation function to calculate the class probabilities.

[0046] S3. Use the improved lightweight edge segmentation model obtained from S2 to segment the rubble pile image, distinguish the rubble edge from the background in the rubble pile image, and obtain the rubble edge map. Specifically, it includes the following steps: S31. Data Preparation: Prepare the data for the obtained rubble pile image dataset, mainly including data annotation, data preprocessing, and data augmentation.

[0047] Data Annotation: The rubble pile image needs to be accurately annotated at the pixel level, that is, the specific contour of the rubble pile is annotated. This is achieved by generating a binary mask (Mask) image, where the rubble pile area is labeled as 1 and other background areas are labeled as 0. Through this annotation, the improved lightweight edge segmentation model can learn how to distinguish the rubble pile from other background areas in the image.

[0048] Data Preprocessing: To ensure that the improved lightweight edge segmentation model can learn efficiently, all images need to be standardized, including unifying the image size, normalizing, and denoising.

[0049] Data Augmentation: Increase the diversity of training data and reduce model overfitting. In this embodiment, the following data augmentation methods are carried out: Simple Spatial Geometric Transformations: Mainly through scaling, translation, rotation, and mirroring, etc.

[0050] Color transformation: Use data augmentation techniques based on color space such as Gamma transformation or histogram equalization to adjust brightness, enhance contrast, etc., and improve the effect of data augmentation.

[0051] Adding noise: Salt-and-pepper noise is added to some muckpile images to improve the generalization ability of the model.

[0052] Random cropping: Randomly crop the size of the gravel pile image to 352×352 pixels.

[0053] S32. Model training: Divide the dataset obtained in S31 into a training set and a test set, and use the data in the training set to train the improved lightweight edge segmentation model.

[0054] The Dice coefficient loss is selected as the loss function. The Dice coefficient is a commonly used evaluation metric in image segmentation, and its value ranges from 0 to 1. The closer it is to 1, the better the segmentation effect. The Dice coefficient loss can effectively address the problem of background imbalance in muckpile images by optimizing the overlap degree of the segmentation results, especially in the case where the target area is small.

[0055] Mini-batch gradient descent is adopted during training, combined with optimization methods such as momentum and learning rate decay. For each image in the training dataset, the loss is calculated through forward propagation, and the network weights are adjusted through backpropagation to minimize the loss function.

[0056] S33. Testing and evaluation: Use the trained model to segment the images in the test set. The performance of the model is evaluated by calculating the difference between the segmentation result and the ground truth mask during the testing process. The evaluation metrics include: Pixel accuracy: Measures the proportion of pixels correctly predicted by the network. It is applicable to the case where the distribution of the background and the target area is relatively uniform.

[0057] IoU: IoU represents the ratio of the intersection to the union of the predicted region and the ground truth region. The larger the IoU, the better the prediction effect.

[0058] Dice coefficient: The Dice coefficient measures the segmentation accuracy, and a higher value indicates that the prediction effect is closer to the ground truth annotation.

[0059] S34. Practical application: Apply the improved lightweight edge segmentation model to the actual muckpile monitoring task. Through the accurate segmentation of muckpile images, automated stacking monitoring can be achieved, and anomalies in the stacked materials can be identified in a timely manner.

[0060] S4. Use the table lookup method to refine the edges of the gravel edge map and perform area filtering in combination with an area filter to refine the edges and eliminate interfering edges.

[0061] To accurately calculate the characteristic parameters of crushed stones, first use an improved lightweight edge segmentation model to segment the crushed stone pile image, distinguish the edges of the crushed stones from the background in the image, and finally obtain the crushed stone edge map. However, due to the severe stacking and complex and chaotic texture of the crushed stones after crushing, the edges in the output result of the lightweight edge segmentation model are still relatively thick, and it is inevitable that there are interference edges in the internal area of the crushed stones that are difficult to remove. To solve the above problems, it is necessary to post-process the edge segmentation result. In this embodiment, a table lookup method is used to extract the edge skeleton, and the method of combining an area filter is used to achieve the effect of refining the edge and eliminating the interference edge.

[0062] The table lookup method is a method for refining binary images. First, the neighborhood situation of each edge pixel point is counted, and the whole image is traversed and divided into two different categories, one is the skeleton point and the other is the boundary point. Finally, the skeleton points are retained and the boundary points are deleted, so as to achieve the purpose of refining the edge. The specific process is as Figure 9 shown. Perform binary processing on the crushed stone edge map to obtain a binary image, and count each edge pixel point and its neighborhood in the binary image; traverse the whole image to divide the edge pixel points into skeleton points and boundary points, retain the skeleton points and delete the boundary points to obtain the edge skeleton map.

[0063] For the problem that there are still interference edges in some internal areas of the crushed stones, if morphological operations are used, some non-negligible effects may be introduced. Therefore, to eliminate larger interference edges or noise points in the image, this embodiment uses the method of area filtering. Its main idea is to set a threshold according to the number of pixels contained in each connected domain in the image, and delete the connected domains smaller than the threshold, so as to achieve the purpose of removing irrelevant small areas. The specific process is as Figure 10 shown. Invert the edge skeleton map obtained by the table lookup method, and count the area of all connected domains in the inverted image and the number of connected domains ; set the threshold , if the area of the connected domain is less than this threshold, that is , then this connected domain is inverted, and the whole image is inverted, and finally all edge branches in the edge skeleton are removed.

[0064] Edge refinement and area filtering were performed on the crushed stone edge map output based on the improved lightweight edge segmentation model, and the results are as Figure 11 shown, Figure 11 in which (a) is the edge image, Figure 11 in which (b) is the image after edge refinement, Figure 11 in which (c) is the image after area filtering. Through a series of post-processing of segmentation, the edge skeleton of the crushed stone pile image was successfully extracted and the isolated edge branches were removed, achieving a more ideal result.

[0065] S5. Extract the outline of the crushed stone edge from the crushed stone edge map, calculate the area and perimeter of the crushed stone, and finally obtain the particle size distribution of the crushed stone.

[0066] Figure 12 The statistical process of the crushed stone characteristic parameters of three groups of crushed stone piles is shown. Figure 12 In (a) is the image of the crushed stone pile. Figure 12 In (b) is the crushed stone contour map. Figure 12 In (c) is the connected region of the crushed stone. First, based on the crushed stone edge map, use the findContours function in OpenCV to accurately find the crushed stone contours in the image, so that each crushed stone forms an independent connected region; then, use the minAreaRect function to obtain the parameters of the minimum bounding rectangle of each crushed stone; finally, calculate the area of each connected region through the contourArea function, and calculate the perimeter of each connected region through the arcLength function. The size of the minimum bounding rectangle reflects the actual size of the crushed stone, where the length and width of the rectangle correspond to the length and width of the crushed stone respectively. The area of the crushed stone can be obtained by calculating the pixel area inside the connected region of the crushed stone, and the perimeter of the crushed stone is determined by measuring the pixel length of the boundary contour of the connected region. The diameter of the crushed stone can be calculated from the area and perimeter. According to the pixel size of the above-mentioned obtained parameters, convert them into the actual size of the crushed stone according to the proportionality coefficient, and finally count the number of crushed stones in each particle size level to obtain the particle size distribution of the crushed stone pile.

[0067] In a specific embodiment, in order to evaluate the improved LDC model, the improved LDC model and other current mainstream image edge segmentation models were trained in the self-built crushed stone pile dataset, and the corresponding weight parameters were saved. In the algorithms involved, BDCN-B2 refers to the first two blocks of BDCN. The purpose of this embodiment is to select the mainstream lightweight edge segmentation method for comparative research. Except for HED and DexiNed, the number of parameters of all models is controlled within 1MB. In order to ensure the fairness of the comparison results, after ensuring that the training and learning processes of all models have achieved good convergence effects, comparative analysis is carried out from both qualitative and quantitative aspects.

[0068] Qualitative analysis. Figure 13 The visual comparison effects of different edge segmentation network models after training on the self-built crushed stone pile image dataset are shown. Figure 13 In (a) are the image IDs, which are Image-1, Image-2, and Image-3 respectively. Figure 13 In (b) is the original image. Figure 13 In (c) is the label image. Figure 13 In (d) is the image detected by the HED algorithm. Figure 13 In (e) is the image detected by the PiDiNet algorithm.Figure 13 (f) is the image detected by TIN algorithm. Figure 13 (g) is the image detected by the BDCN-B2 algorithm. Figure 13 The image in (h) is detected by the DexiNed algorithm. Figure 13 (i) is the image detected by the LDC algorithm. Figure 13 (j) is the image detected by the improved LDC algorithm. As can be seen from the figure, the edge positioning detected by the HED, PiDiNet, TIN and BDCN-B2 algorithms is relatively accurate, but they are seriously affected by noise, the obtained edges are discontinuous, and there is a phenomenon of misidentifying the surface texture of the gravel as the edge of the gravel, and the overall visual effect is relatively stiff; while DexiNed and LDC are accurate in detecting the edges of large pieces of gravel with complex textures, and are less affected by noise, but the edge detection effect is not good in places with low contrast at the edge of the gravel, and the details are not rich enough. In contrast, the improved LDC algorithm proposed in this application can obtain clearer and continuous edges, has the ability to eliminate noise and accurately locate significant edges, and effectively reduces the generation of pseudo-edges. In places where gravel overlaps severely, the algorithm of this embodiment can accurately identify the edge of the gravel, and the predicted edge map is rich in details, has a good human eye perception, and is closer to the label map of the image.

[0069] Figure 14 The comparative results of the prediction of the LDC algorithm and the improved LDC algorithm proposed in this application on low-quality images with harsh environments, high dust content, and insufficient lighting are demonstrated. Figure 14 (a) is the original image. Figure 14 (b) is the image predicted by the LDC algorithm. Figure 14 (c) is the image predicted by the improved LDC algorithm. Under undesirable conditions such as insufficient image illumination, low contrast, and blurred edges, the edges detected by the traditional LDC have a large number of false detections and missed detections, the edge images are relatively rough, and the edge recognition of low background contrast is incomplete, and its performance is poor under harsh conditions. In contrast, the improved LDC algorithm proposed in the present application is more sensitive to fine edges than the LDC algorithm model, and can detect finer gravel pieces and more accurately capture edge information in gravel pile images, thereby obtaining clearer and more accurate edges. In general, the improved LDC algorithm proposed in the present application performs better in the detection of gravel edges in gravel piles, and can better perform the task of gravel edge segmentation in gravel piles.

[0070] Quantitative comparison. To objectively measure the effectiveness of the edge detection results of each model, in this embodiment, the same operations were performed on the gravel pile images during the specific tests of different models, and the same evaluation code was used to calculate all edge segmentation methods. As shown in Table 1, the algorithm proposed in this application was quantitatively compared with other algorithms for the edge detection effect of Image-1 to Image3. The results show that compared with the HED and DexiNed edge detection algorithms with a parameter quantity exceeding 1M, the performance of the improved LDC algorithm proposed in this application overall exceeded that of the HED algorithm, and the AP increased by 3.1%. Among the lightweight algorithm models, the three evaluation index values of the algorithm proposed in this application reached the maximum. Compared with the traditional LDC model, only a small number of parameters were added, and significant improvements in the performance of key indicators such as ODS, OIS, and AP were achieved. This indicates that the algorithm model proposed in this application effectively improves the accuracy and stability of edge detection while maintaining the model complexity and computational efficiency, and the edge pixels detected by the model have a high degree of coincidence with the manually annotated edge images.

[0071] Table 1 Objective quantitative comparison results of example images in the gravel pile dataset

[0072] In a specific embodiment, in order to better analyze the effectiveness of different improvement modules on the edge detection results, ablation experiments were further carried out on the self-built gravel pile image test set. The experimental results are shown in Table 2. In the table, "√" represents the use of the corresponding improvement module, and "×" represents the non-use of the corresponding improvement module. All algorithms in the table used the same data augmentation and evaluation index code during the experiment. Among them, AD represents the deformable convolutional module with fused attention, ARSBU represents the attention residual shrinkage unit, and DC refers to the use of dilated convolution instead of max pooling. By comparing Experiment 4 and Experiment 5, the introduction of the deformable convolutional module with fused attention only sacrificed 125KB of parameter quantity, but the evaluation indexes ODS and OIS increased by 2.8% and 3.1% respectively, and the AP increased by 2.9%, and the performance was significantly improved. Comparing the results of Experiment 3 and Experiment 5, the use of the attention residual shrinkage unit improved the algorithm performance, where the ODS increased by 2.5%, the OIS increased by 2.6%, and the AP increased by 1.6%. In addition, the results of Experiment 2 and Experiment 5 show that replacing max pooling with dilated convolution can also improve the algorithm performance to a certain extent. Generally speaking, the ODS of the algorithm proposed in this application increased by 3.4%, the OIS increased by 3.6%, and the AP increased by 2.8%.

[0073] As Figure 15 and Figure 16As shown in the figure, after 20 cycles of training, each algorithm model basically converges. By comparing the change trends of the loss functions and accuracies of different algorithm models in the ablation experiment, it can be seen that the algorithm proposed in this application has a relatively fast convergence speed, and its accuracy in the initial stage of training is much higher than that of other algorithms. The traditional LDC model finally converges to the largest loss value, and its accuracy is also much lower than that of other algorithms. AD+DC performs well in terms of training loss, but its accuracy is not as good as the algorithm proposed in this application.

[0074] Table 2 Results of the ablation experiment of the improved module

[0075] In a specific embodiment, by using the method of analyzing the cumulative distribution function and the frequency distribution function, the characteristic parameters of three groups of gravel piles, namely a, b, and c, are statistically analyzed in detail, and the particle size of the gravel is calculated. The particle size is divided into 7 different particle size grades, and the particle size distribution data statistically analyzed by artificial screening and the method proposed in this application are compared and analyzed in order to comprehensively evaluate the characteristics of the gravel and deeply understand its particle size characteristics. The comparison results are shown in Tables 3-5.

[0076] Table 3 Comparison of particle size distributions of gravel pile a

[0077] Table 4 Comparison of particle size distributions of gravel pile b

[0078] Table 5 Comparison of particle size distributions of gravel pile c

[0079] As can be seen from Tables 3-5, when using the method proposed in this application for particle size statistics, the average error of gravel pile a in each particle size grade interval is 1.10%, and the average errors of gravel pile b and gravel pile c in each particle size grade are 2.43% and 2.25% respectively. At the low particle size grade, the detection results of the three groups of gravel piles are slightly larger than the screening results, there is a certain over-segmentation phenomenon, and the error is relatively large. However, as the particle size grade increases, the detection results become more and more accurate, and the error in each particle size grade remains below 5%. Generally speaking, the ability demonstrated by the method proposed in this application in these three groups of experiments is satisfactory, which is not only reflected in the high-precision particle size distribution statistics, but also can maintain accuracy when dealing with larger particle size grades, providing reliable data support for practical applications.

[0080] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for identifying the particle size of broken rocks based on machine vision, characterized in that: The following steps are involved: S1, obtaining an image of a gravel pile in a quarry; S2. Based on the lightweight edge segmentation model, the deformable convolution module with fused attention and the attention residual shrinkage unit are combined, and the maximum pooling layer is replaced by the hole convolution layer to obtain an improved lightweight edge segmentation model; S3, using the improved lightweight edge segmentation model obtained in S2 to segment the gravel pile image, distinguishing gravel edges from the background in the gravel pile image, and obtaining a gravel edge map; S4, using a table lookup method to refine the edge of the gravel edge map, and combining it with an area filter to perform area filtering to refine the edge and eliminate interference edges; S5. Extract the edge contour of the gravel from the gravel edge map, calculate the area and perimeter of the gravel, and finally obtain the particle size distribution of the gravel.

2. The method for identifying the particle size of broken rocks based on machine vision according to claim 1 is characterized in that: The rubble pile images come from different scenes including construction sites and mines, and are taken under different lighting, angles and backgrounds.

3. The method for identifying the particle size of broken rocks based on machine vision according to claim 2 is characterized in that: The attention-fused deformable convolution module is composed of a deformable convolution module in parallel with a hybrid attention module.

4. The method for identifying the particle size of broken rocks based on machine vision according to claim 3 is characterized in that: The attention residual shrinkage unit uses one-dimensional convolution to replace the soft threshold of the residual shrinkage unit, and introduces an attention mechanism with spatial weights at the identity mapping of the residual shrinkage unit.

5. The method for identifying the particle size of broken rocks based on machine vision according to claim 4 is characterized in that: The structure of the improved lightweight edge segmentation model is as follows: A deformable convolution module with fused attention is added before the second and third modules of the lightweight edge segmentation model, an attention residual shrinkage unit is added before the context-aware fusion block of the lightweight edge segmentation model, and the maximum pooling layer of the lightweight edge segmentation model is replaced by a hole convolution layer.

6. The method for identifying the particle size of broken rocks based on machine vision according to claim 5 is characterized in that: The S3 comprises the following steps: S31, performing data annotation, data preprocessing and data enhancement on the acquired gravel pile image data set; S32, dividing the data set obtained in S31 into a training set and a test set, and using the data of the training set to train the improved lightweight edge segmentation model, and selecting the Dice coefficient loss as the loss function; S33. Use the trained model to segment the images in the test set, and evaluate the model performance by calculating the difference between the segmentation result and the true mask. The evaluation indicators include pixel accuracy, IoU and Dice coefficient. S34. Apply the improved lightweight edge segmentation model to the actual gravel pile monitoring task.

7. The method for identifying the particle size of broken rocks based on machine vision according to claim 6 is characterized in that: The method of using the table lookup method to refine the edge of the gravel edge map comprises the following steps: Binarize the gravel edge image to obtain a binary image, and count each edge pixel point and its neighborhood in the binary image; Traverse the entire image and divide the edge pixels into skeleton points and boundary points, retain the skeleton points and delete the boundary points to obtain the edge skeleton map.

8. The method for identifying the particle size of broken rocks based on machine vision according to claim 7, characterized in that: The area filtering performed in combination with the area filter comprises the following steps: The edge skeleton image obtained by the table lookup method is inverted, and the area of ​​all connected domains of the inverted image is counted. and the number of connected domains ; Setting the Threshold , if the area of ​​the connected domain is less than the threshold, that is , then this connected domain is inverted, and finally all edge branches in the edge skeleton are removed.

9. The method for identifying the particle size of broken rocks based on machine vision according to claim 8, characterized in that: The S5 comprises the following steps: Based on the gravel edge map, the findContours function of OpenCV is used to accurately find the gravel contours in the image, so that each gravel forms an independent connected domain; Use the minAreaRect function to obtain the minimum enclosing rectangle parameters of each gravel; The area of ​​each connected domain is calculated by the contourArea function, the perimeter of each connected domain is calculated by the arcLength function, and the diameter of the connected domain is calculated by the area and perimeter of the connected domain; According to the pixel size of the acquired area, perimeter and diameter, the actual size of the gravel is converted according to the proportional coefficient, and the number of gravel of each particle size is counted to obtain the particle size distribution of the gravel pile.

Citation Information

Patent Citations

  • Conveyor belt ore granularity detection method based on edge response fusion algorithm

    CN113240663A

  • Vehicle position estimation method fused with filter network and computer readable medium

    CN116086476A

  • Identity recognition method based on improved residual shrinkage network

    CN116305048A

  • Image semantic segmentation model and segmentation method

    CN116468740A

  • Multi-source data fusion unmanned aerial vehicle forestry remote sensing safety target identification method and system

    CN118506214A

Cited By

  • Ore image edge segmentation method and system based on image recognition

    CN120563845A

  • Deep well block falling while drilling real-time detection method based on DA-YOLO

    CN122115810A

  • Deep well while-drilling block shedding real-time detection method based on DA-YOLO

    CN122115810B