Method, Readable Storage Medium, and Terminal Device for Online Detection of Aggregates

Through the improved Cascade Mask-RCNN model and feature extraction network backbone, combined with ResNeSt and ResNet network structure, the problems of low accuracy and high leakage detection in aggregate detection are solved, and more efficient aggregate detection and analysis are achieved.

CN114648507BActive Publication Date: 2025-05-27POWERCHINA HUADONG ENG CORP LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210297285.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-05-27
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

The prior art has problems in the detection of aggregates with low detection accuracy, high leakage detection rate and sensitivity to complex lighting environments.

Method used

The improved Cascade Mask-RCNN model is adopted, combined with ResNeSt and ResNet network structures, and features are extracted through feature extraction network backbone, and data enhancement and depth labeling are performed to improve the accuracy of aggregate image segmentation.

Benefits of technology

It significantly improves the accuracy of aggregate detection, reduces the leakage detection rate of aggregate, and can better adapt to complex lighting environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648507B_ABST
    Figure CN114648507B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for on-line detection of aggregates, comprising the steps of: collecting a plurality of aggregate images to form a data set and dividing it into a training set and a test set, performing aggregate contour depth annotation and data augmentation on the aggregate images of the training set, then preprocessing the aggregate images, training the preprocessed aggregate images within the Pytorch framework using a pre-constructed improved Cascade Mask-RCNN model, validating using the aggregate images of the test set to obtain an aggregate image segmentation and analysis processing model; deploying the obtained aggregate image segmentation and analysis processing model in an aggregate on-line detection system, using this system to perform on-line detection and analysis of aggregates, and outputting the analysis results. The present invention also discloses a readable storage medium and a terminal device using this method. The method of the present invention reduces the missed detection of aggregates and improves the accuracy of aggregate detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for on-line detection of aggregate, a readable storage medium and a terminal device. Background Art

[0002] Sand and gravel aggregate is a general term for materials such as crushed stones, block stones, and dressed stones in water conservancy projects, and is one of the main building materials for structures such as concrete. Nowadays, in commercial construction enterprises, the detection of sand and gravel aggregate mostly adopts the traditional manual detection method. However, manual detection is affected by human factors, resulting in low detection efficiency. With the continuous improvement of industrial automation level, automated detection will gradually replace manual detection. The technologies adopted by some domestic research institutions and enterprises at the present stage can be divided into the following several types:

[0003] (1) Method based on traditional image processing: 1) Using image processing technology to extract the aspect ratio coefficient of coarse aggregate, so as to achieve the purpose of detecting sand and gravel aggregate; 2) Using digital image processing technology to segment rock particles and fine seams, extract the contour information of sand and gravel aggregate, and measure the aggregate; 3) Using image processing technology, first perform binarization in the first step, adopt the method of threshold segmentation to distinguish foreground objects and background, and here the global threshold segmentation method is used. In the second step, perform image denoising and then perform opening operation (first erosion, then dilation). In the third step, observe the dilated image, and it can be found that the aggregate is segmented into independent closed regions. At this time, use the corresponding function to find the outer contour. Sometimes there will be some incorrect detections, and certain conditions can be added for judgment.

[0004] (2) Method based on deep learning: Nowadays, there are many deep learning algorithms based on image segmentation, such as: traditional algorithms commonly used for region segmentation, watershed algorithm, region growing algorithm. The region-based and semantic segmentation algorithms are the main research directions at present, such as the fully convolutional network of the FCN series, as well as the u-net series and deeplab series commonly used in medical images.

[0005] The above methods generally have the following disadvantages in the process of practical application:

[0006] (1) Disadvantages of traditional image processing: Digital image processing technology has high requirements for the lighting environment. In the actual detection process, the surrounding environment is relatively complex, and the illumination degree of the light source may change with the unstable voltage, resulting in unstable light source; (2) Disadvantages of deep learning method: It is necessary to balance detection accuracy and detection speed. The detection accuracy of some occluded and overlapping aggregates is low, and it cannot well segment the instances of sand and gravel aggregate. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a method, a readable storage medium and a terminal device for on-line detection of aggregate, so as to improve the accuracy of aggregate detection and reduce the missed detection of aggregate.

[0008] The present invention is implemented as follows. A method for on-line detection of aggregate is provided, including the following steps:

[0009] Step 1: Use an on-line image acquisition device to pre-collect a variety of aggregate images to form a data set.

[0010] Step 2: The data set is divided into a training set and a test set according to a ratio of 7:3. The aggregate images in the training set are marked with the depth of the aggregate contour, and the marked aggregate images are subjected to data augmentation processing.

[0011] Step 3: Preprocess the aggregate images after data augmentation processing. The preprocessing includes performing histogram equalization processing on the aggregate images.

[0012] Step 4: Construct an improved Cascade Mask-RCNN model. The features of the input image to the model are extracted by the feature extraction network backbone. Among them, the feature extraction network backbone is combined by stacking Bottle block1 and Bottle block2. The Bottle block1 is the block structure of the ResNeSt network, and the Bottle block2 is the block structure of the ResNet network; the feature extraction network backbone consists of five stages. The 0th stage includes a convolutional layer, a batch normalization layer and a max pooling layer. The 1st stage and the 4th stage both include a Bottle block1 and two Bottle block2s. The 2nd stage includes a Bottle block1 and three Bottle block2s. The 3rd stage includes a Bottle block1 and five Bottle block2s.

[0013] Step 5: Use the improved Cascade Mask-RCNN model obtained in Step 4 to train the aggregate images in the training set after preprocessing in Step 3 within the Pytorch framework, and use the aggregate images in the test set for verification to obtain an aggregate image segmentation and analysis processing model.

[0014] Step 6: Deploy the obtained aggregate image segmentation and analysis processing model in the aggregate on-line detection system, use the system to perform on-line detection and analysis of the aggregate, and output the analysis results.

[0015] Further, in Step 1, the multiple aggregate images refer to the aggregate images obtained by setting different camera heights, camera focal lengths, different ambient light intensities, and shooting from different angles of the aggregates.

[0016] Further, in Step 2, the steps of performing aggregate contour depth annotation on the aggregate images in the training set include: performing estimated annotation on the deep aggregates according to the existing bounding boxes of the deep aggregates, and manually adding occlusion contour annotation to the occluded aggregates.

[0017] Further, in Step 2, the data augmentation processing includes processing the labeled pictures by using methods such as translation, rotation, or adding noise.

[0018] Further, in Step 3, the preprocessing further includes performing normalization processing on the image after histogram equalization processing.

[0019] Further, in Step 6, the analysis results include: the particle size, volume, and particle shape parameters of the aggregates, as well as the aggregate gradation parameters and the content of needle-like and flaky particles.

[0020] Further, in Step 6, the output of the analysis results includes transmitting the analysis results to the cloud platform and displaying them.

[0021] The present invention is implemented in this way. There is also provided a readable storage medium storing multiple instructions suitable for being loaded by a processor to execute the method for online detection of aggregates as described above.

[0022] The present invention is implemented in this way. There is also provided a terminal device including a processor and a memory, where the memory stores multiple instructions, and the processor loads the instructions to execute the method for online detection of aggregates as described above.

[0023] Compared with the prior art, for a method, a readable storage medium, and a terminal device for online detection of aggregates according to the present invention, first, enough data is collected to construct a data set, the data is labeled, and then data augmentation is performed on the data to ensure the balance of the number of data. Then, a network model is constructed for training and verification, and finally, it is deployed to a hardware device. The present invention selects ResNeSt50 as the backbone network for feature extraction of the model, optimizes the ResNeSt50 algorithm to improve the aggregate image segmentation method, and thus obtains an improved aggregate image segmentation and analysis processing model. Deploying this processing model in an aggregate online detection system improves the accuracy of aggregate online detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic diagram of the principle flow of the online detection method of the present invention;

[0025] Figure 2 It is the block diagram of the core model structure of Cascade Mask-RCNN;

[0026] Figure 3 It is the block diagram of the backbone structure of the feature extraction network of the improved Cascade Mask-RCNN of the present invention;

[0027] Figure 4 It is Figure 3 The block diagram of Bottle block1 in;

[0028] Figure 5 The block diagram of a base array in Bottle block1;

[0029] Figure 6 It is Figure 3 The block diagram of Bottle block2 in; Specific implementation manners

[0030] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0031] Please refer to Figure 1 As shown, a preferred embodiment of a method for on-line detection of aggregates of the present invention includes the following steps:

[0032] Step 1, Use an on-line image acquisition device to pre-collect a variety of aggregate images to form a data set.

[0033] The variety of aggregate images refer to aggregate images obtained by setting different camera heights, camera focal lengths, different environmental light intensities and shooting from different angles of the aggregates.

[0034] Step 2, The data set is divided into a training set and a test set according to a ratio of 7:3. The aggregate images in the training set are marked with the aggregate contour depth, and the marked aggregate images are subjected to data augmentation processing.

[0035] The steps of marking the aggregate contour depth of the aggregate images in the training set include: making a preliminary mark for the deep aggregates according to the existing frames of the deep aggregates, and artificially adding occlusion contour marks for the occluded aggregates.

[0036] The data augmentation processing includes processing the marked pictures by means of translation, rotation or adding noise.

[0037] Step 3: Preprocess the aggregate images after data augmentation processing. The preprocessing includes performing histogram equalization processing on the aggregate images.

[0038] The preprocessing further includes performing normalization processing on the images after histogram equalization processing.

[0039] Step 4: Construct an improved Cascade Mask-RCNN model. Extract features from the images input into the model through the feature extraction network backbone. Among them, the feature extraction network backbone is composed of the stacking combination of Bottle block1 and Bottle block2. The Bottle block1 is the block structure of the ResNeSt network, and the Bottle block2 is the block structure of the ResNet network. The feature extraction network backbone consists of five stages. The 0th stage includes a convolutional layer, a batch normalization layer, and a max pooling layer. The 1st stage and the 4th stage both include one Bottle block1 and two Bottle block2s. The 2nd stage includes one Bottle block1 and three Bottle block2s. The 3rd stage includes one Bottle block1 and five Bottle block2s.

[0040] Step 5: Use the improved Cascade Mask-RCNN model obtained in Step 4 to train the aggregate images of the training set preprocessed in Step 3 within the Pytorch framework, and use the aggregate images of the test set for verification to obtain an aggregate image segmentation and analysis processing model.

[0041] Step 6: Deploy the obtained aggregate image segmentation and analysis processing model in the aggregate online detection system, use this system to perform online detection and analysis on the aggregates, and output the analysis results.

[0042] The analysis results include: the particle size, volume, and particle shape parameters of the aggregates, as well as the aggregate gradation parameters and the content of flaky and elongated particles. The output of the analysis results includes transmitting the analysis results to the cloud platform and displaying them.

[0043] The following further illustrates a method for online detection of aggregates according to the present invention through specific embodiments.

[0044] Embodiment 1

[0045] Please refer to Figure 1 As shown, a method for online detection of aggregates according to the present invention includes the following steps:

[0046] Step 1: Image acquisition and dataset construction

[0047] Aggregate image acquisition is carried out based on the experimental platform built in the laboratory, or it is fixed on the conveyor belts of several selected actual aggregate production lines for aggregate image acquisition to form a dataset. The multiple aggregate images refer to the aggregate images obtained by setting different camera heights, camera focal lengths, different ambient light intensities, and shooting from different angles of the aggregate, so as to collect a sufficient number of aggregate size images of various types. After collecting a certain number of pictures, a dataset is formed.

[0048] Step 2: Dataset division

[0049] The dataset is divided into a training set and a test set in a ratio of 7:3.

[0050] Step 3: Data annotation and data augmentation

[0051] Deep annotation is performed on the aggregate image data of the training set, and the labelme software is used for annotation.

[0052] When annotating each piece of sand and gravel aggregate, the image needs to be enlarged for annotation to prevent gaps from appearing during annotation.

[0053] The so-called deep annotation means that not only obvious shallow aggregates need to be annotated during annotation, but also estimated annotation should be carried out according to the existing bounding boxes of deep aggregates. Deep annotation improves the accuracy of the model in detecting aggregates and reduces the missed detection of aggregates.

[0054] In the collected aggregate images, some aggregates may be occluded. During annotation, it is necessary to estimate according to the contour edges of the aggregate and artificially add occluded contour annotations according to the prediction to ensure that each aggregate is completely annotated.

[0055] The annotated images are matched with the json files for data augmentation. The annotated pictures are processed such as translation, rotation, adding noise, etc., and it is ensured that the json files also change with the changes of the pictures.

[0056] The data augmentation makes the test data as close as possible to the real data, otherwise there may be an overestimation of the performance of the model.

[0057] Step 4: Data preprocessing

[0058] For the aggregate images taken on-site, due to the influence of external light, the aggregate images will be bright and dark. Therefore, histogram equalization processing is performed on the aggregate images of the training set.

[0059] Histogram equalization processing is a method to enhance image contrast. Its main idea is to make the histogram distribution of this image more uniform, and the obtained image will be clearer than the original image.

[0060] Histogram equalization is a transformation based on the probability density function of a random variable. It broadens the gray levels with more pixels in the image and reduces those with fewer pixels. First, count the gray level r in the image k appearance probability:

[0061]

[0062] where L is the gray level of the image, n is the total number of pixels in the image, and n k is the number of pixels with gray level r k in the image. The transformation function is:

[0063]

[0064] Each pixel value with gray level r k in the image is mapped to s k .

[0065] After histogram equalization processing on the aggregate images in the training set, then normalize the pixel values in the aggregate images (min-max normalization), which will have better training performance and faster convergence speed. Its principle is to perform a linear transformation on the original data and map the values to the range [0, 1]. The min-max normalization processing formula is as follows:

[0066] For each pixel x 1 , x 2 ,......, x n in the image, perform the transformation:

[0067]

[0068] Then the new data y 1 , y 2 ,......, y n ∈[0, 1] and is dimensionless.

[0069] Step 5: Build the algorithm model

[0070] The core model structure of Cascade Mask-RCNN is as Figure 2 shown. After extracting features from the input image through the feature extraction network backbone, optimize the prediction results through three detection networks in serial cascade based on different IOU thresholds.

[0071] As Figure 2 shown, Pool represents the ROI Align module for calibrating the image, H is the detection network, C represents the classification result, S represents the segmentation result, B represents the detection box result, and backbone is the ResNet50 network.

[0072] As Figure 3 shown, the backbone feature extraction network adopts the block structure (Bottle block1) of ResNeSt (Split-Attention Networks), combines the block structure (Bottle block2) of ResNet in the original model, and stacks them into a feature extraction network. The backbone of the feature extraction network consists of five stages. The 0th stage (Stage 0) includes a convolutional layer (CONV), a batch normalization (BN) layer, and a max pooling layer (MAXPOOL). The 1st stage (Stage 1) and the 4th stage (Stage 4) both include one Bottle block1 and two Bottle block2s. The 2nd stage (Stage 2) includes one Bottle block1 and three Bottle block2s. The 3rd stage (Stage 3) includes one Bottle block1 and five Bottle block2s. Each block adjusts the number of channels according to different inputs.

[0073] Among them, Bottle block1 adopts the ResNeSt block structure, with the cardinality group K = 2 and the number of units r = 2 in each cardinality group to construct the network.

[0074] The structural block diagram of Bottle block1 is as Figure 4 shown.

[0075] For a certain cardinality group (taking the cardinality group K as an example) in the structural block diagram of Bottle block1, its structure is as Figure 5 shown.

[0076] In Bottle block1, the input of each block is divided into K (K = 2) groups along the channel dimension. Each group selects weights according to the global context information. Each group of features is a cardinality group. At the same time, each cardinality group is further divided into r (r = 2) units, and the output V of each unit c k is:

[0077]

[0078] Where:

[0079]

[0080]

[0081]

[0082] In the formula, Represents the output weights of the softmax function, and the mapping function Consists of two fully connected layers plus the ReLu function, and determines the global context representation information s of each channel unit k of the weights Is obtained by global average pooling of the channel global statistical information. The representation of the k-th cardinality group is U j Is the j-th unit under the k-th cardinality group, and H, W, and C are the height, width, and number of channels of the output feature layer respectively.

[0083] Bottle block2 adopts the residual structure in the ResNet block. ResNet is composed of blocks stacked by convolutions with different numbers of channels. The basic framework of each block is as Figure 6 shown, and the shortcut is used to solve the problem of gradient disappearance when the network becomes deeper.

[0084] ResNeSt uses a modular architecture, applies channel-wise attention to different network branches, and captures cross features in the network. ResNeSt divides the input of each block along the channel dimension into K (K = 2) groups, and each group selects weights according to the global context information. Each group of features is a cardinality group, maintaining the modularity of the block units. Similarly, a structure similar to ResNet is created by stacking blocks. By concatenating the outputs of the K cardinality groups along the channel dimension V K and performing a 1×1 convolution to obtain V:[[]]

[0085] V = Concat{V 1 , V 2 , …, V K}

[0086] Using the shortcut of the residual structure, the final output is:[[]]

[0087] Y = V + X

[0088] Step 6: Model training and model verification

[0089] Use the improved Cascade Mask-RCNN model obtained in step 5 to perform model training and model verification on the aggregate images of the training set preprocessed in step 4 in the Pytorch framework. The optimizer Adam (Adaptive Moment Estimation) is used during the model training process. The optimizer Adam is an algorithm that combines the Momentum algorithm and the RMSProp algorithm, that is, it uses momentum to accumulate gradients, and at the same time makes the convergence speed faster and the fluctuation amplitude smaller, and performs bias correction. The learning rate lr is set to 0.0002, and the epoch is set to 12. The training set is input into the improved Cascade Mask-RCNN model for model training. The trained algorithm model is verified through the test set. If it does not meet the expectations, the model is adjusted and then the model training continues until it meets the expectations.

[0090] Step 7: Model deployment

[0091] Deploy the tested improved Cascade Mask R-CNN model in the aggregate online detection system.

[0092] Step 8: Online detection and analysis result output

[0093] The aggregate online detection system performs online detection on the aggregate. The images collected by the online image acquisition device are transmitted to the aggregate online detection system and then the segmentation results are calculated and output through the improved Cascade Mask R-CNN model. The aggregate online detection system transmits the segmentation results to the server backend, and the backend further calculates the analysis results such as the particle size, volume, and particle shape of the aggregate and uploads them to the cloud platform. The cloud platform interface is displayed in real time through the monitor, and the information on the cloud platform interface also includes aggregate gradation parameters and flaky and elongated particle content.

[0094] The present invention also discloses a readable storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the method for online detection of aggregates as described above.

[0095] The present invention also discloses a terminal device, including a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the method for online detection of aggregates as described above.

[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for on-line detection of aggregates, characterized in that, it includes the following steps: Step 1: Use an on-line image acquisition device to pre-collect a variety of aggregate images to form a data set; Step 2: The data set is divided into a training set and a test set according to a ratio of 7:

3. The aggregate images in the training set are labeled with the depth of the aggregate contour, and the labeled aggregate images are subjected to data augmentation processing; Step 3: Preprocess the aggregate images after data augmentation processing. The preprocessing includes performing histogram equalization processing on the aggregate images; Step 4: Build an improved Cascade Mask-RCNN model. The features of the input image of the model are extracted by the feature extraction network backbone. Among them, the feature extraction network backbone is combined by stacking Bottle block1 and Bottle block2. The Bottle block1 is the block structure of the ResNeSt network, and the Bottle block2 is the block structure of the ResNet network; the feature extraction network backbone consists of five stages. The 0th stage includes a convolutional layer, a batch normalization layer and a max pooling layer. The 1st stage and the 4th stage both include a Bottle block1 and two Bottle block2s. The 2nd stage includes a Bottle block1 and three Bottle block2s. The 3rd stage includes a Bottle block1 and five Bottle block2s; Step 5: Use the improved Cascade Mask-RCNN model obtained in Step 4 to train the aggregate images in the training set after preprocessing in Step 3 within the Pytorch framework, and use the aggregate images in the test set for verification to obtain an aggregate image segmentation and analysis processing model; Step 6: Deploy the obtained aggregate image segmentation and analysis processing model in the aggregate on-line detection system, use the system to perform on-line detection and analysis of the aggregates, and output the analysis results.

2. A method for on-line detection of aggregates according to claim 1, characterized in that, in Step 1, the variety of aggregate images refer to the aggregate images obtained by setting different camera heights, camera focal lengths, different environmental light intensities, and shooting from different angles of the aggregates.

3. A method for on-line detection of aggregates according to claim 1, characterized in that, in Step 2, the steps of labeling the depth of the aggregate contour of the aggregate images in the training set include: estimating and labeling the deep aggregates according to the existing bounding boxes of the deep aggregates, and manually adding occlusion contour labels to the occluded aggregates.

4. A method for on-line detection of aggregates according to claim 1, characterized in that, in Step 2, the data augmentation processing includes processing the labeled pictures by means of translation, rotation or adding noise.

5. A method for on-line detection of aggregates according to claim 1, characterized in that, In step three, the preprocessing further includes normalizing the image that has undergone histogram equalization processing.

6. A method for on-line detection of aggregates as described in claim 1, characterized in that in step six, the analysis results include: the particle size, volume and particle shape parameters of the aggregates, as well as the aggregate gradation parameters and the content of flaky and elongated particles.

7. A method for on-line detection of aggregates as described in claim 1, characterized in that in step six, the output of the analysis results includes transmitting the analysis results to the cloud platform and displaying them.

8. A readable storage medium storing multiple instructions, characterized in that the instructions are adapted to be loaded by a processor to execute the method for on-line detection of aggregates according to any one of claims 1 to 7.

9. A terminal device comprising a processor and a memory, characterized in that the memory stores multiple instructions, and the processor loads the instructions to execute the method for on-line detection of aggregates according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Large aggregate concentration area detection method and device and network model training method thereof

    CN111797826A

  • Sanding pipe loosening fault image recognition method and system based on deep learning

    CN112418253A