Rock fracture identification method based on improved Deeplabv3 + model

By using lightweight MobileNetV3 network and ASFF structure in the Deeplabv3+ model, the problems of high complexity and insufficient recognition accuracy in rock fracture recognition are solved, and a more efficient and accurate rock fracture recognition effect is achieved.

CN119942301APending Publication Date: 2025-05-06XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY +2

Patent Information

Application Number
CN202510015234.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The traditional Deeplabv3+ model has redundant feature extraction in rock fracture recognition, resulting in high complexity and slow convergence speed, making it difficult to achieve high-precision recognition under complex lighting conditions and blurred backgrounds.

Method used

By replacing the backbone network Xception to the lightweight MobileNetV3 and adding an ASFF (Adaptively Spatial Feature Fusion) structure, an improved Deeplabv3+ model is built to reduce the complexity of the model, improve the convergence speed, and enhance the recognition ability through the adaptive fusion of multi-scale features.

Benefits of technology

The accuracy and stability of rock fracture recognition was significantly improved, especially in complex lighting conditions and blurred backgrounds, the average crossover ratio (MIOU) of the model increased by 0.53%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942301A_ABST
    Figure CN119942301A_ABST
Patent Text Reader

Abstract

The invention discloses a rock fracture identification method based on an improved Deeplabv3 + model. The rock fracture identification method comprises the following steps: obtaining a rock fracture model; the method comprises the following steps: S1, carrying out image acquisition on a crack-containing rock sample to obtain a rock sample crack data set; s2, performing point labeling on fractures in the rock sample fracture data set by adopting a Labeme tool, delineating a sample fracture area, and establishing a fracture labeling image data set; s3, an improved Deeplabv3 + deep learning model is constructed by replacing a backbone network Xception with a MobileNetV3 structure and adding an ASFF structure, and the improved Deeplabv3 + deep learning model is constructed; therefore, the crack identification precision of the Deeplabv3 + model is enhanced; and S4, using an improved Deeplabv3 + deep learning model to train a crack labeling image data set. According to the method, the complexity of the model is reduced, the convergence speed is increased, adaptive fusion of multi-scale features is realized, the recognition capability of the model on complex and fine rock fractures is enhanced, and the rock fractures are efficiently and accurately recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and computer vision, and in particular to a rock crack recognition method based on an improved Deeplabv3+ model. Background Art

[0002] In geology and geotechnical engineering, the identification of rock cracks is of great significance for evaluating the mechanical properties of rock masses, predicting geological disasters, and guiding engineering practice. Traditional rock crack identification methods mainly rely on manual observation and measurement, which is not only time-consuming and labor-intensive, but also easily affected by subjective factors, making it difficult to ensure the accuracy and consistency of the identification results. With the rapid development of deep learning technology, rock crack identification methods based on deep learning have gradually emerged. Among them, the Deeplabv3+ model, as an advanced semantic segmentation model, performs well in image segmentation tasks and is widely used in the identification of road cracks, building cracks and other scenes.

[0003] However, due to the diverse morphology of rock cracks, their sizes, shapes and distributions vary, which requires the crack identification model to have strong feature extraction and generalization capabilities. However, the traditional Deeplabv3+ model may have redundancy in feature extraction, resulting in high model complexity and slow convergence speed. Therefore, there are still many challenges in directly using the Deeplabv3+ model for rock crack identification.

[0004] When processing such images, the traditional Deeplabv3+ model may be affected by factors such as lighting changes and noise interference, resulting in poor recognition results. However, rock crack recognition needs to be performed under complex lighting conditions and blurred backgrounds, which places higher requirements on the robustness and recognition accuracy of the model.

[0005] In order to overcome the above challenges, there is a degree thesis entitled "Research on Rock Fracture Segmentation in Field Outcrop Areas Based on Improved Deeplabv3+ Network" (DOI: 10.26995 / d.cnki.gdqsc.2023.000790) which aims to improve the feature extraction ability and recognition accuracy of the model by introducing the attention mechanism to address the shortcomings of the traditional Deeplabv3+ model. However, this approach has the disadvantages of increased computational complexity, large memory usage, increased parameters, and increased training difficulty.

[0006] At the same time, the existing invention patent "Method and system for identifying surface cracks in tunnel rock mass based on deep learning" (authorization announcement number: CN117854060B) improves recognition accuracy and robustness by optimizing model structure and training strategy. However, this method still faces the problem of limited generalization ability when processing rock crack images with extremely high complexity and diversity, resulting in insufficient crack recognition accuracy under extreme conditions.

[0007] In summary, although existing literature and patents have attempted to make targeted improvements to the traditional Deeplabv3+ model to improve the accuracy and efficiency of rock fracture identification, the complexity of rock fractures has caused the existing rock fracture identification work to still face many challenges. Therefore, future research work urgently needs to explore more efficient and accurate model structures and training strategies to improve the accuracy and stability of rock fracture identification. Summary of the invention

[0008] In order to overcome the above technical problems, the purpose of the present invention is to provide a rock fracture identification method based on an improved Deeplabv3+ model. This method reduces the complexity of the model and improves the convergence speed by replacing the backbone network Xception with a lightweight MobileNetV3. At the same time, by adding the ASFF (Adaptively Spatial Feature Fusion) structure, the adaptive fusion of multi-scale features is realized, and the model's recognition ability for complex and subtle rock fractures is enhanced, thereby efficiently and accurately identifying rock fractures.

[0009] In order to achieve the above object, the technical solution adopted by the present invention is:

[0010] A rock fracture identification method based on an improved Deeplabv3+ model comprises the following steps;

[0011] S1: Capture images of rock samples containing cracks to obtain rock sample crack data sets;

[0012] S2: Use the Labeme tool to mark the cracks in the rock sample crack dataset, delineate the sample crack area, and establish a crack annotation image dataset;

[0013] S3: By replacing the backbone network Xception with MobileNetV3 and adding the ASFF structure, an improved Deeplabv3+ deep learning model is constructed, thereby enhancing the crack recognition accuracy of the Deeplabv3+ model;

[0014] S4: Use the improved Deeplabv3+ deep learning model to train the crack annotation image dataset. The crack recognition accuracy is improved by 0.53%.

[0015] The specific method for obtaining the rock sample fracture data set in S1 is:

[0016] Image acquisition of rock samples containing cracks uses a high-resolution camera to ensure that the image is clear, detailed, and the crack morphology is completely visible. During the shooting process, the lighting conditions should be controlled to avoid overexposure or too dark. At the same time, shooting at different angles and distances should be used to obtain diverse image data and obtain a rock sample crack data set.

[0017] The specific method for establishing the crack annotation image dataset in S2 is:

[0018] First, the rock sample fracture dataset is imported into the Labeme annotation tool; fine point annotation is performed along the fracture edge in sequence;

[0019] Secondly, these points are connected to form a curve to accurately delineate the crack area of ​​the sample;

[0020] Then, the annotated curve is smoothed; this process is repeated until all the crack images are annotated;

[0021] Finally, these annotated images are integrated into a crack annotated image dataset as training samples for model training.

[0022] The specific method for constructing the improved Deeplabv3+ deep learning model in S3 is:

[0023] By replacing the backbone network Xception with MobileNetV3 and adding the ASFF structure, an improved Deeplabv3+ deep learning model is constructed.

[0024] First, the rock crack annotated image is input into the backbone network MobileNetV3, which is responsible for extracting feature map information in different channels from the rock crack annotated image, and then generating a high-dimensional feature image of the rock crack;

[0025] Secondly, the rock fracture high-dimensional feature image is used as the input of the ASFF structure. The ASFF structure extracts multi-scale features from the rock fracture high-dimensional feature image output by each convolutional layer of the backbone network MobileNetV3, and performs adaptive fusion based on spatial information.

[0026] Then, the adaptive weighted multi-scale rock fracture feature map after ASFF structure fusion is further fused with the multi-scale rock fracture feature map output by the ASPP module in DeeplabV3+;

[0027] Finally, the DeeplabV3+ model processes the fused multi-scale rock fracture feature map through the upsampling layer and the classification layer to generate the final rock fracture segmentation result.

[0028] The specific method for establishing the MobileNetV3 network in step S3 is:

[0029] S311: MobileNetV3 network uses deep separable convolution to reduce the amount of calculation and model parameters;

[0030] In each convolution layer of the MobileNetV3 network, deep convolution (DW) is performed first and then point-by-point convolution (PW). In the deep convolution stage, each input channel uses a separate 3×3 convolution kernel to perform a convolution operation on the rock sample crack annotation image, thereby extracting the initial spatial features formed by the pixel values ​​and their permutations and combinations in the original images of rock crack annotations of different channels; point-by-point convolution is used to combine these features. Point-by-point convolution (PW) uses a 1×1 convolution kernel to perform a convolution operation on the output of the deep convolution of different channels, that is, the initial spatial features formed by the pixel values ​​and their permutations and combinations in the original images of rock crack annotations, to generate a low-dimensional feature image of rock cracks (that is, a lower-dimensional feature representation obtained by some change or processing of the original image of rock crack annotations). For each deep separable convolution operation, the total number of rock crack convolution parameters Param num The convolution calculation cost of rock fractures is:

[0031] Param num =D k ×D k ×M+1×1×M (1)

[0032] Cost = D k ×D k ×M×size out ×size out (2)

[0033] Where D k is the convolution kernel size, M is the number of channels of the input rock crack annotation feature image of each convolution layer;

[0034] MobileNetV3 is further lightweighted using the width scaling factor α∈(0,1] and the rock crack annotation feature image resolution scaling factor ρ∈(0,1];

[0035] For the width scaling factor α, the number of channels of the input and output rock crack annotation feature images of each convolutional layer is scaled to α M and α N , scaled rock fracture parameter Param num And the calculation cost is shown as follows:

[0036] Param num =D k ×D k ×α M +1×1×α M ×α N (3)

[0037] Cost = D k ×D k ×αM ×size out ×size out +size out ×size out ×α M ×α N (4)

[0038] Each convolutional layer performs the above operation and takes the output of the upper convolutional layer as the input of the next layer. As each convolutional layer goes deeper, the feature dimension of the rock crack feature map increases to capture higher-level abstract features of rock cracks, thereby obtaining a high-dimensional feature image of rock cracks.

[0039] S312: MobileNetV3 network uses h-swish function; specifically:

[0040] First, the h-swish function is used instead of the Relu6 function to solve the problem of information loss when processing the low-dimensional feature image of rock sample cracks. The function is shown in the following formula:

[0041] swish[x]=x·sigmod(βx) (5)

[0042] use The function piecewise linearly simulates the sigmoid function and is replaced by the h-swish function:

[0043]

[0044] The specific method for establishing the ASFF structure in step S3 is:

[0045] S321: The ASFF structure transforms the high-dimensional feature images of rock cracks in different convolutional layers obtained from MobileNetV3 to the same size, making the feature map size of each layer the same;

[0046] S322: The ASFF structure adaptively fuses the high-dimensional feature images of rock cracks obtained by MobileNetV3 at different convolutional layers; the adaptive mixing formula is as follows:

[0047]

[0048] Among them, α, β, γ are weight values; Represents the high-dimensional feature maps of rock cracks in different convolutional layers of MobileNetV3 after being resized to the same size. The [channels, w, h] of the three x are the same; that is, the positions of the high-dimensional feature maps of rock cracks in different layers are multiplied by the weights at their respective positions, and finally the three results are added together to obtain the value of the position after fusion;

[0049] The weight formula for calculating the spatial position of rock fractures is:

[0050]

[0051] The weights are learned when calculating the spatial position of rock fractures, that is, Perform 1*1 convolution to generate three single-channel rock fracture feature maps, namely, three rock fracture initial spatial weight information matrices, denoted as three λ. The weights at position ij are processed by formula (2), namely, the sum of the weights at the same position of the three rock fracture weight maps is 1, and all are ∈[0.1]. The final weight map becomes the following formula:

[0052]

[0053] Among them, α, β, and γ are the actual rock fracture weights.

[0054] According to the learned adaptive weights, the high-dimensional feature images of rock fractures from different convolutional layers obtained from MobileNetV3 are weighted fused, and the fused multi-scale rock fracture feature map contains richer spatial information and contextual information; the fused multi-scale rock fracture feature map is sent to the Decoder module in the DeeplabV3+ model for further fusion with the rock sample fracture feature map output by ASPP; finally, the DeeplabV3+ model processes the fused multi-scale rock fracture feature map through the upsampling layer and the classification layer, restores the feature map to the resolution of the original rock fracture image, and performs pixel-level classification prediction on the output of the Decoder module through the classifier to generate the final rock fracture segmentation result.

[0055] The specific method of training the crack annotation image dataset with the improved Deeplabv3+ deep learning model in S4 is:

[0056] The improved Deeplabv3+ deep learning model is used to systematically train the crack annotated image dataset. Through a large number of iterative training, the model can fully learn the characteristics of the cracks. At the same time, the model is regularly evaluated using the validation set to monitor the training progress and performance of the model.

[0057] Beneficial effects of the present invention:

[0058] The present invention fully considers the advantages of MobileNetV3 and ASFF structures, and uses a lighter and more efficient MobileNetV3 network as a rock fracture feature extractor to accurately extract key information of rock fractures. At the same time, the addition of the ASFF structure enables the model to adaptively fuse rock fracture feature maps from different levels that contain rich detail information, thereby achieving high-precision capture of the subtle features of rock fractures and their spatial distribution. This series of improvements ensures the efficient operation of the model and can significantly improve the recognition accuracy of rock fractures, especially in complex scenes such as underexposure, dark environments, and subtle and blurred fractures. The model can still accurately identify fractures, which increases the average intersection-over-union (MIOU) of the model by 0.53%. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a schematic flow chart of a rock fracture identification method based on deep learning in the present invention.

[0060] Figure 2 Schematic diagram of some cracks in the present invention.

[0061] Figure 3 This is a schematic diagram of crack marking in the present invention.

[0062] Figure 4a This is a schematic diagram of the improved Deeplabv3+ deep learning model in the present invention.

[0063] Figure 4b Schematic diagram of the MobileNetV3 network structure in the embodiment.

[0064] Figure 4c Schematic diagram of the ASFF structure in the embodiment.

[0065] Figure 5a This is a schematic diagram of the training results of the improved Deeplabv3+ deep learning model.

[0066] Figure 5b This is a schematic diagram comparing the training results before and after the improvement of the Deeplabv3+ deep learning model.

[0067] Figure 5c This is a schematic diagram comparing the MIoU values ​​before and after the improvement of the Deeplabv3+ deep learning model. DETAILED DESCRIPTION

[0068] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0069] Example

[0070] Figure 1 A schematic flow chart of a rock fracture identification method based on deep learning provided by the present invention is shown;

[0071] like Figure 1 As shown, the present invention discloses a rock fracture identification method based on deep learning, comprising:

[0072] S1, collect images of rock samples containing cracks to obtain rock sample crack data sets, such as Figure 2 The following are some schematic diagrams of rock fractures, which cover rock fracture images collected under various shooting conditions, including: rock fracture images taken under normal lighting conditions; fracture images taken under overexposure and blurred conditions; images taken under conditions with clear fracture edges and less background interference; and images taken under conditions with subtle fractures and blurred images. Through these diverse image data, this application aims to build a comprehensive and representative rock fracture dataset to provide a solid foundation for subsequent analysis and identification work.

[0073] In this embodiment, the rock samples containing cracks are professionally photographed using a high-resolution camera to ensure that the images are clear, detailed, and the crack morphology is fully visible;

[0074] During the shooting process, control the lighting conditions to avoid overexposure or darkening, and consider shooting at different angles and distances to obtain diverse image data;

[0075] The captured images are preprocessed, such as denoising and contrast enhancement, to improve image quality and provide a good foundation for subsequent annotation and model training.

[0076] S2, use the Labele tool to mark the cracks, delineate the sample crack area, and establish a crack annotation image dataset;

[0077] First, the collected rock sample fracture datasets need to be systematically imported into the Labeme annotation tool. Labeme is an open source software designed for image annotation. It supports multiple image formats and provides a convenient interface for multiple annotation methods such as point annotation and polygon annotation. When importing images, ensure that the image format, resolution, and naming rules meet the project requirements for subsequent management and processing;

[0078] In Labeme, select the point marking tool and mark the cracks one by one. During the marking process, it is necessary to ensure that each marked point is closely aligned with the actual position of the crack to ensure the accuracy of the marking. Use the zoom function of Labeme to zoom in on the crack details so as to place the marked points more accurately.

[0079] After completing the point marking, the user connects these points to form one or more curves to accurately delineate the crack area. The position and number of the marked points are constantly adjusted to further optimize the shape of the curve so that it better fits the actual contour of the crack;

[0080] In order to eliminate the jagged and uneven phenomenon that may occur during point marking and curve connection, the marked curve is smoothed to eliminate the unnatural turns in the marking process, making the crack marking more smooth and natural;

[0081] For each crack image in the dataset, repeat the above steps of point labeling, curve connection and smoothing to ensure the labeling quality and consistency of the entire dataset;

[0082] After all the crack images are labeled, a comprehensive quality check is performed on the labeled images to ensure the accuracy and consistency of the annotations, including checking whether the labeled points fit the cracks tightly, whether the curves accurately delineate the crack areas, and whether the smoothing process is appropriate.

[0083] The quality-checked annotated images are integrated into a complete crack annotated image dataset. This dataset contains the original images and the corresponding annotation information, providing reliable training samples for subsequent model training. Figure 3 Schematic diagram for marking the cracks.

[0084] S3, by replacing the backbone network Xception with the lightweight MobileNetV3 and adding the ASFF structure, builds an improved Deeplabv3+ deep learning model, thereby enhancing the recognition accuracy of the Deeplabv3+ model. Figure 4aSchematic diagram of the improved Deeplabv3+ deep learning model in the present invention, which is mainly composed of an input image, a MobileNetV3 backbone network, an ASPP module, an ASFF module and a Deeplabv3+ head; in the encoder part, the rock fracture annotated image passes through the backbone network MobileNetV3 to obtain two rock fracture high-dimensional feature image layers, one of which enters the ASFF structure to adaptively fuse rock fracture features of different scales, and the obtained adaptive weighted multi-scale rock fracture feature map enters the 1×1 convolution of the decoder for channel compression; the other rock fracture high-dimensional feature image layer enters the ASPP module of the encoder; the ASPP module is used to capture multi-scale contextual information, and includes multiple parallel void data layers, which have different spatial void rates to capture high-dimensional features of rock fractures of different scales; ASP The P module includes a 1×1 convolution, three 3×3 hole convolutions with different ratios (6, 12, 18) and a global pooling operation of the image. The rock crack feature maps of different branches generated by these operations are integrated through a 1×1 convolution to obtain richer rock crack feature information; in the decoder part, the upsampling layer restores the low-resolution rock crack feature map output by the encoder ASPP module to obtain a high-resolution rock crack feature map; the fusion feature layer splices the high-dimensional feature image of rock cracks from the encoder's backbone network MobileNetV3 with the upsampled high-resolution rock crack feature image to obtain a rock crack feature fusion image; the fused image is further refined and segmented through 3×3 convolution to restore the edge details of the rock cracks, and finally the rock crack feature fusion image is restored to the same resolution as the original image through the upsampling layer.

[0085] like Figure 4bAs shown in the figure, the network structure of MobileNetV3 is divided into three parts. The starting part uses a 3×3 convolution kernel to extract the features of the input rock crack image. This layer will be connected to the normalization layer and the h-swish activation layer; the middle part includes a block composed of multiple convolution layers, including a 1×1 convolution layer, a depth-separable convolution layer, an SE module, and a residual connection layer; the 1×1 convolution is used to change the number of channels to achieve dimensionality increase and dimensionality reduction of rock crack features; the depth-separable convolution layer is a 5×5 convolution kernel, which is used to decompose the standard convolution into depth-wise convolution and point-by-point convolution to reduce the amount of model calculation; the SE module is used to weight the importance of each channel, and the residual connection layer is used to promote the back propagation of the gradient; the last part includes a global pooling layer, a 1×1 convolution, and a Softmax layer; the global pooling layer performs global average pooling on the output of the last convolution layer to obtain the global features of each channel; the 1×1 convolution converts the output of the global average pooling into the required number of categories; the Softmax layer is used to convert the output of the convolution layer into a probability distribution. Starting from the input rock crack annotated image, the feature extraction is first performed through the backbone network MobileNetV3. The original model uses Xception as the backbone network. In order to improve the efficiency and performance of the model, it is now replaced with a lighter and more efficient MobileNetV3 as the new Backbone; MobileNetV3 uses deep separable convolution to greatly reduce the amount of calculation, and introduces the h-swish function to further improve the operation efficiency and recognition accuracy of the model; the network is responsible for extracting feature map information in different channels from the rock crack annotated image, and then generating a high-dimensional feature image of the rock crack; secondly, the high-dimensional feature image of the rock crack is used as the input of the ASFF structure and the ASPP structure respectively;

[0086] like Figure 4c As shown in the figure, the ASFF structure consists of two parts: same-size transformation and adaptive fusion; the ASFF structure extracts multi-scale features from the high-dimensional feature image of rock fractures output by each convolution layer of the backbone network MobileNetV3. Since the sizes of the rock fracture feature images obtained from different convolution layers are different, the feature images of different layers need to be transformed to the same size; if you choose to fuse at level 1, the transformation sizes of level 2 and level 3 are the same as the size of level 1, and they are transformed to the same size through upsampling and downsampling respectively; the high-dimensional feature maps of rock fractures of the same size are transformed and adaptively fused through ASFF1, 2, and 3 modules according to spatial information; then, the adaptive weighted multi-scale rock fracture feature map fused by the ASFF structure is further fused with the multi-scale rock fracture feature map output by the ASPP module in DeeplabV3+ through the convolution kernel; finally, the DeeplabV3+ model processes the fused multi-scale rock fracture feature map through the upsampling layer and the classification layer to generate the final rock fracture segmentation result.

[0087] S311: MobileNetV3 network uses deep separable convolution to reduce the amount of calculation and model parameters;

[0088] In each convolution layer of the MobileNetV3 network, deep convolution (DW) is performed first and then point-by-point convolution (PW). In the deep convolution stage, each input channel uses a separate 3×3 convolution kernel to perform a convolution operation on the rock sample crack annotation image, thereby extracting the initial spatial features formed by the pixel values ​​and their permutations and combinations in the original images of rock crack annotations of different channels; point-by-point convolution is used to combine these features. Point-by-point convolution (PW) uses a 1×1 convolution kernel to perform a convolution operation on the output of the deep convolution of different channels, that is, the initial spatial features formed by the pixel values ​​and their permutations and combinations in the original images of rock crack annotations, to generate a low-dimensional feature image of rock cracks (that is, a lower-dimensional feature representation obtained by some change or processing of the original image of rock crack annotations). For each deep separable convolution operation, the total number of rock crack convolution parameters Param num The convolution calculation cost of rock fractures is:

[0089] Param num =D k ×D k ×M+1×1×M (1)

[0090] Cost = D k ×D k ×M×size out ×size out (2)

[0091] Where D k is the convolution kernel size, M is the number of channels of the input rock crack annotation feature image of each convolution layer;

[0092] MobileNetV3 is further lightweighted using the width scaling factor α∈(0,1] and the rock crack annotation feature image resolution scaling factor ρ∈(0,1];

[0093] For the width scaling factor α, the number of channels of the input and output rock crack annotation feature images of each convolutional layer is scaled to α M and α N , scaled rock fracture parameter Param num And the calculation cost is shown as follows:

[0094] Param num =D k ×D k ×α M +1×1×α M×α N (3)

[0095] Cost = D k ×D k ×α M ×size out ×size out +size out ×size out ×α M ×α N (4)

[0096] Each convolutional layer performs the above operation and takes the output of the upper convolutional layer as the input of the next layer. As each convolutional layer goes deeper, the feature dimension of the rock crack feature map increases to capture higher-level abstract features of rock cracks, thereby obtaining a high-dimensional feature image of rock cracks.

[0097] S312: MobileNetV3 network uses h-swish function; specifically:

[0098] First, the h-swish function is used instead of the Relu6 function to solve the problem of information loss when processing the low-dimensional feature image of rock sample cracks. The function is shown in the following formula:

[0099] swish[x]=x·sigmod(βx) (5)

[0100] use The function piecewise linearly simulates the sigmoid function and is replaced by the h-swish function:

[0101]

[0102] The specific method for establishing the ASFF structure in step S3 is:

[0103] S321: The ASFF structure transforms the high-dimensional feature images of rock cracks in different convolutional layers obtained from MobileNetV3 to the same size, making the feature map size of each layer the same;

[0104] S322: The ASFF structure adaptively fuses the high-dimensional feature images of rock cracks obtained by MobileNetV3 at different convolutional layers; the adaptive mixing formula is as follows:

[0105]

[0106] Among them, α, β, β are weight values; Represents the high-dimensional feature maps of rock cracks in different convolutional layers of MobileNetV3 after being resized to the same size. The [channels, w, h] of the three x are the same; that is, the positions of the high-dimensional feature maps of rock cracks in different layers are multiplied by the weights at their respective positions, and finally the three results are added together to obtain the value of the position after fusion;

[0107] The weight formula for calculating the spatial position of rock fractures is:

[0108]

[0109] The weights are learned when calculating the spatial position of rock fractures, that is, Perform 1*1 convolution to generate three single-channel rock fracture feature maps, namely, three rock fracture initial spatial weight information matrices, denoted as three λ. The weights at position ij are processed by formula (2), namely, the sum of the weights at the same position of the three rock fracture weight maps is 1, and all are ∈[0.1]. The final weight map becomes the following formula:

[0110]

[0111] Among them, α, β, and γ are the actual rock fracture weights.

[0112] According to the learned adaptive weights, the high-dimensional feature images of rock fractures from different convolutional layers obtained from MobileNetV3 are weighted fused, and the fused multi-scale rock fracture feature map contains richer spatial information and contextual information; the fused multi-scale rock fracture feature map is sent to the Decoder module in the DeeplabV3+ model for further fusion with the rock sample fracture feature map output by ASPP; finally, the DeeplabV3+ model processes the fused multi-scale rock fracture feature map through the upsampling layer and the classification layer, restores the feature map to the resolution of the original rock fracture image, and performs pixel-level classification prediction on the output of the Decoder module through the classifier to generate the final rock fracture segmentation result.

[0113] S4, using the improved Deeplabv3+ deep learning model to train the fracture annotation image dataset to achieve accurate identification of rock fractures;

[0114] Use the improved Deeplabv3+ deep learning model to systematically train the crack annotation image dataset. During the training process, set the learning rate, batch size and other parameters reasonably to ensure effective learning of the model;

[0115] Through a large number of iterative training, the model can fully learn the characteristics of the cracks. After the training is completed, the model is evaluated using the test set, such as Figure 5aAs shown in the figure, it is a schematic diagram of the training results of the improved Deeplabv3+ deep learning model. Figure 5b Figure 1 is a schematic diagram comparing the training results before and after the improvement of the Deeplabv3+ deep learning model. It can be seen from the figure that the improved Deeplabv3+ deep learning model has a good recognition effect on each crack, and can better identify subtle cracks; the image No. 9 taken under exposed and blurred conditions and the crack No. 75 taken in the dark have missed recognition problems in the Deeplabv3+ deep learning model, and the cracks are not recognized in the thinner cracks, while the improved Deeplabv3+ deep learning model can better recognize the cracks; when the crack edges are clear and there is less interference, the Deeplabv3+ deep learning model before and after the improvement can both recognize the cracks well; when the cracks are fine and blurred, the cracks No. 60, No. 118 and No. 157 have crack recognition interruptions under the Deeplabv3+ deep learning model, and the improved Deeplabv3+ deep learning model can recognize coherent cracks with better recognition effect; finally, the MIoU value is used to judge the model prediction results. MIoU (Mean Intersection Over Unit) MIoU is a key indicator for evaluating the performance of image segmentation models, especially in semantic segmentation and instance segmentation tasks. It calculates the ratio of the intersection and union between the model's prediction results and the true labels, and then takes the average value for all categories. Therefore, the higher the MIoU value, the higher the overlap between the model's predicted area and the true area, that is, the more accurate the model's prediction results are; since MIoU is the average of all categories' IoU, it can comprehensively reflect the model's segmentation performance in each category. A high MIoU value indicates that the model not only performs well in certain categories, but also maintains good segmentation effects in all categories, reflecting the strong comprehensive performance of the model. Figure 5c Table 1 shows a schematic diagram of the MIoU value comparison of the model before and after improvement and a specific schematic diagram of the MIOU value comparison. It can be seen that the improved Deeplabv3+ deep learning model has a high MIoU value and the model recognition accuracy is 0.53% higher than the original model.

[0116] Table 1 Comparison of MIOU values ​​before and after improvement of Deeplabv3+ deep learning model

[0117] MIOU value before improvement Improved MIOU value 0 0 80.03092582019154 82.3748412297722 85.32891324595595 85.8105674415073 86.902350762355 88.9971851725255 87.13547727913618 89.42796484159902 88.86582140143764 89.92903949657119 89.67237072299248 90.60402900212331 89.24023712844092 90.79244554247774 90.27893293486878 91.00857478356006 89.89613453377301 90.88972223590781 90.43913074267039 91.2354215246412 90.04512887232559 91.2922636840268 90.6899649701711 91.454316391556 90.92266691395227 91.5065524198387 91.02085950278888 91.59235687764861 91.12170868195134 91.30265671309353 91.11419076734268 91.28360209385816 91.04445703044293 91.3791098231252 91.04473688743279 91.58213829078954 91.18938853949703 91.56584223197956 91.07462164191456 91.5943764467854

Claims

1. A rock fracture identification method based on an improved Deeplabv3+ model, characterized in that: The steps include: S1: Capture images of rock samples containing cracks to obtain rock sample crack data sets; S2: Use the Labeme tool to mark the cracks in the rock sample crack dataset, delineate the sample crack area, and establish a crack annotation image dataset; S3: Build an improved Deeplabv3+ deep learning model; S4: Use the improved Deeplabv3+ deep learning model to train the crack annotation image dataset.

2. The rock fracture identification method based on the improved Deeplabv3+ model according to claim 1, characterized in that: The specific method for obtaining the rock sample fracture data set in S1 is: Image acquisition of rock samples containing cracks uses a high-resolution camera to ensure that the image is clear, detailed, and the crack morphology is completely visible. During the shooting process, the lighting conditions should be controlled to avoid overexposure or too dark. At the same time, shooting at different angles and distances should be used to obtain diverse image data and obtain a rock sample crack data set.

3. The rock fracture identification method based on the improved Deeplabv3+ model according to claim 1 is characterized in that: The specific method for establishing the crack annotation image dataset in S2 is: First, the rock sample fracture dataset is imported into the Labeme annotation tool; fine point annotation is performed along the fracture edge in sequence; Secondly, these points are connected to form a curve to accurately delineate the crack area of ​​the sample; Subsequently, the marked curves are smoothed; Repeat this process until all crack images are labeled; Finally, these annotated images are integrated into a crack annotated image dataset as training samples for model training.

4. The rock fracture identification method based on the improved Deeplabv3+ model according to claim 3 is characterized in that: The specific method for constructing the improved Deeplabv3+ deep learning model in S3 is: First, the rock crack annotated image is input into the backbone network MobileNetV3, which is responsible for extracting feature map information in different channels from the rock crack annotated image, and then generating a high-dimensional feature image of the rock crack; Secondly, the rock fracture high-dimensional feature image is used as the input of the ASFF structure. The ASFF structure extracts multi-scale features from the rock fracture high-dimensional feature image output by each convolutional layer of the backbone network MobileNetV3, and performs adaptive fusion based on spatial information. Then, the adaptive weighted multi-scale rock fracture feature map after ASFF structure fusion is further fused with the multi-scale rock fracture feature map output by the ASPP module in DeeplabV3+; Finally, the DeeplabV3+ model processes the fused multi-scale rock fracture feature map through the upsampling layer and the classification layer to generate the final rock fracture segmentation result.

5. The rock fracture identification method based on the improved Deeplabv3+ model according to claim 4 is characterized in that: The specific method for establishing the MobileNetV3 network in step S3 is: S311: MobileNetV3 network uses deep separable convolution to reduce the amount of calculation and model parameters; In each convolution layer of the MobileNetV3 network, deep convolution (DW) is performed first and then point-by-point convolution (PW). In the deep convolution stage, each input channel uses a separate 3×3 convolution kernel to perform a convolution operation on the rock sample crack annotation image, thereby extracting the initial spatial features formed by the pixel values ​​and their permutations and combinations in the original images of rock crack annotations of different channels; point-by-point convolution is used to combine these features. Point-by-point convolution (PW) uses a 1×1 convolution kernel to perform a convolution operation on the output of the deep convolution of different channels, that is, the initial spatial features formed by the pixel values ​​and their permutations and combinations in the original images of rock crack annotations to generate a low-dimensional feature image of rock cracks. For each deep separable convolution operation, the total number of rock crack convolution parameters Param num The convolution calculation cost of rock fractures is: Stop num =D k ×D k ×M+1×1×M (1) Cost=D k ×D k ×M×size out ×size out (2) Where D k is the convolution kernel size, M is the number of channels of the input rock crack annotation feature image of each convolution layer; MobileNetV3 is further lightweighted using the width scaling factor α∈(0,1] and the rock crack annotation feature image resolution scaling factor ρ∈(0,1]; For the width scaling factor α, the number of channels of the input and output rock crack annotation feature images of each convolutional layer is scaled to α M and α N , scaled rock fracture parameter Param num And the calculation amount Cost is shown as follows: Stop num =D k ×D k ×α M +1×1×α M ×α N (3) Cost=D k ×D k ×α M ×size out ×size out +size out ×size out ×α M ×α N (4) Each convolutional layer performs the above operations and uses the output of the upper convolutional layer as the input of the next layer. By capturing higher-level abstract features of rock cracks, a high-dimensional feature image of rock cracks is obtained. S312: MobileNetV3 network uses h-swish function; The h-swish function is used to solve the problem of information loss when processing low-dimensional feature images of rock sample cracks. The function is shown as follows: swish[x]=x·sigmod(βx) (5) use Function piecewise linear simulation sigmoid function:

6. The rock fracture identification method based on the improved Deeplabv3+ model according to claim 4 is characterized in that: The specific method for establishing the ASFF structure in step S3 is: S321: The ASFF structure transforms the high-dimensional feature images of rock cracks in different convolutional layers obtained from MobileNetV3 to the same size, making the feature map size of each layer the same; S322: The ASFF structure adaptively fuses the high-dimensional feature images of rock cracks obtained by MobileNetV3 at different convolutional layers; the adaptive mixing formula is as follows: Among them, α, β, γ are weight values; Represents the high-dimensional feature maps of rock cracks in different convolutional layers of MobileNetV3 after being resized to the same size. The [channels, w, h] of the three x are the same; that is, the positions of the high-dimensional feature maps of rock cracks in different layers are multiplied by the weights at their respective positions, and finally the three results are added together to obtain the value of the position after fusion; The weight formula for calculating the spatial position of rock fractures is: For three Perform 1*1 convolution to generate three single-channel rock fracture feature maps, namely, three rock fracture initial spatial weight information matrices, denoted as three λ. The weights at position ij are processed by formula (2), namely, the sum of the weights at the same position of the three rock fracture weight maps is 1, and all are ∈[0.1]. The final weight map becomes the following formula: Among them, α, β, and γ are the real rock fracture weight values; According to the learned adaptive weights, the high-dimensional feature images of rock fractures from different convolutional layers obtained from MobileNetV3 are weightedly fused to obtain a fused multi-scale rock fracture feature map; the fused multi-scale rock fracture feature map is sent to the Decoder module in the DeeplabV3+ model for further fusion with the rock sample fracture feature map output by ASPP; finally, the DeeplabV3+ model processes the fused multi-scale rock fracture feature map through the upsampling layer and the classification layer, restores the feature map to the resolution of the original rock fracture image, and performs pixel-level classification prediction on the output of the Decoder module through the classifier to generate the final rock fracture segmentation result.

7. The rock fracture identification method based on the improved Deeplabv3+ model according to claim 6 is characterized in that: The specific method of training the crack annotation image dataset with the improved Deeplabv3+ deep learning model in S4 is: The improved Deeplabv3+ deep learning model is used to systematically train the crack annotated image dataset. Through a large number of iterative training, the model can fully learn the characteristics of the cracks. At the same time, the model is regularly evaluated using the validation set to monitor the training progress and performance of the model.

Citation Information

Patent Citations

  • Tunnel rock surface crack recognition method and system based on deep learning

    CN117854060B

Cited By

  • Shale blast CT image fracture extraction method based on DeeplabV < 3 + >

    CN121481996A