Crack image segmentation method based on ECA-AC-ResUnet network

By introducing a hollow residual convolution module, an ECA attention mechanism and a defect correction module into the ECA-AC-ResUnet network, combining the fusion loss function of cross entropy loss and Dice loss, the shortcomings of small crack treatments in road cracks, adaptability and continuity treatments in different scales in the prior art are solved, and a more accurate crack image segmentation effect is achieved.

CN120047452APending Publication Date: 2025-05-27HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510010105.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has shortcomings in the treatment of small cracks of road cracks, crack adaptability and crack continuity treatment of different scales, resulting in inaccurate segmentation results.

Method used

The fracture image segmentation method based on the ECA-AC-ResUnet network is adopted, and the model's ability to capture crack morphological information by introducing a hole residual convolution module in the encoder stage, and the ECA attention mechanism and defect correction module are introduced in the decoder stage, and combined with the fusion loss of the binary classification and the fusion loss function of Dice loss, the model's ability to capture crack morphological information is enhanced.

Benefits of technology

The accuracy of crack image segmentation is improved, the model's ability to identify small cracks and complex background cracks is enhanced, and the model's stable performance in various challenging situations is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047452A_ABST
    Figure CN120047452A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image segmentation, and discloses a crack image segmentation method based on an ECA-AC-ResUnet network, and the method comprises the steps: obtaining a road crack image data set; the method comprises the following steps: constructing an ECA-AC-ResUnet road crack segmentation network model, wherein the model is composed of an encoder part, a decoder part and a corresponding jump connection; the encoder part takes U-Net as a main network, and introduces a hole residual convolution module; the decoder part comprises four up-sampling modules, and corresponding jump connection performs splicing operation on a feature map after up-sampling operation and a feature map of a corresponding space size in the encoder part; an ECA attention mechanism module and a defect correction module are introduced into the decoder part; proposing a fusion loss function which is composed of a dichotomy cross entropy loss function and a Dice loss function; and training by using the data set in the step 1, and carrying out crack image segmentation by using the trained ECA-AC-ResUnet road crack segmentation network model. Compared with the prior art, the shape information of the crack can be more accurately captured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular to a crack image segmentation method based on an ECA-AC-ResUnet network. Background Art

[0002] Image segmentation technology is a basic technology in the field of computer vision. It is widely used in image processing, medical image analysis, remote sensing image processing, autonomous driving, transportation and other fields, and has great application value. The main goal of image segmentation is to divide an image into several regions or objects with separate features for subsequent morphological analysis, evaluation and monitoring tasks.

[0003] Roads are the cornerstone of the entire transportation network and play a key role in providing basic service activities, promoting social development and promoting people's well-being. The formation of road cracks is the result of the interaction of multiple factors, including traffic load, natural factors, soil conditions, construction quality, vehicle behavior, etc. In the field of road maintenance, the formation of cracks is an important factor affecting the service life and safety of roads. It not only causes damage to the road structure, but also affects the comfort and safety of vehicle driving. Therefore, accurate detection and segmentation of cracks has important practical significance. The main aspect of the current research is to more accurately locate the cracks in the crack images that have been collected.

[0004] Traditional crack detection methods mainly rely on manual inspection, which is not only inefficient but also prone to subjective errors. In recent years, crack detection methods based on image processing have gradually become mainstream. Detecting and segmenting cracks in an automated way can improve efficiency and accuracy. Common image segmentation methods include threshold segmentation and edge detection methods. The threshold segmentation method can simply divide the image into foreground and background, but the effect of threshold segmentation is not ideal when the crack morphology is complex, the lighting conditions change, or there is a lot of noise in the image. The edge detection method performs segmentation by detecting the edge information of the image. Commonly used edge detection operators include Sobel, Canny, etc. However, the edge detection method is prone to errors when dealing with fuzzy crack edges or interference information, resulting in inaccurate segmentation results.

[0005] In recent years, with the continuous development of neural networks, deep learning methods have made significant progress in the field of crack image segmentation. Convolutional neural networks (CNNs) are widely used in image segmentation tasks. Neural network models can learn a large amount of data training, automatically learn the feature information of cracks and backgrounds, and achieve pixel-level segmentation. At present, segmentation network models based on convolutional neural networks include U-Net, FCN (Fully Convolutional Network), SegNet and DeepLab, etc. These networks have achieved relatively excellent performance in image segmentation tasks. In particular, the U-Net structure has been widely used in crack segmentation tasks due to its ability to capture multi-scale features and detail information.

[0006] However, with the increasing use of roads, the types of road cracks are becoming more complex and diverse, so the difficulty of crack detection is gradually increasing. The current methods still have some problems and limitations. Strip cracks include horizontal cracks, longitudinal cracks, oblique cracks, block cracks, turtle-shaped cracks, etc. The sizes are not consistent, and the differences between different categories are also very large; there are problems such as illumination changes, noise interference and shadow shielding in the image, and there is also sample imbalance that leads to inaccurate crack segmentation; there are large differences in the crack image features of different road surfaces, all of which limit the versatility and accuracy of the current network model.

[0007] The main research method in the field of road crack segmentation is based on neural network models based on image processing and deep learning. Compared with other traditional convolutional neural networks, the U-Net network uses a data enhancement method for cutting the input image into pieces, so that the valuable information in each input image can be captured by the network as much as possible. Therefore, fewer training data sets are required, and the segmentation results are more accurate. It is widely used in the field of road crack segmentation. However, the model still has some problems and defects in the detailed processing of cracks, mainly the following key issues:

[0008] Insufficient ability to handle small cracks: Due to the various shapes of road cracks, especially small cracks, U-Net is prone to problems such as blurring and information loss when segmenting these details, ignoring the tiny features in the cracks, resulting in poor recognition of small cracks.

[0009] Insufficient adaptability to cracks of different scales: Road cracks vary greatly in shape and size. U-Net’s ability in multi-scale feature extraction is relatively limited, and it is unable to effectively distinguish and capture cracks of different scales, resulting in the segmentation results being unable to accurately reflect the actual crack morphology.

[0010] Insufficient continuity and integrity processing capabilities: Cracks often have a coherent linear or mesh shape, but the U-Net model may cause discontinuity or fracture in the crack area, resulting in the segmented crack shape being inconsistent with the actual crack shape, affecting the integrity of the segmentation effect. Summary of the invention

[0011] Purpose of the invention: In view of the problems in the background technology, the present invention discloses a crack image segmentation method based on the ECA-AC-ResUnet network, introduces a hole residual convolution module in the encoder stage, proposes an ECA attention mechanism and a defect correction module in the decoder stage, combines the fusion loss function of the cross entropy loss of the binary classification and the Dice loss, and introduces the auxiliary loss of each layer of upsampling module on the basis of the fusion loss. By weighted combination of the fusion loss and the sum of the auxiliary losses of each upsampling module, the model can more accurately capture the morphological information of the cracks in the learning process of multi-level features.

[0012] Technical solution: The present invention discloses a crack image segmentation method based on the ECA-AC-ResUnet network, comprising the following steps:

[0013] Step 1: Obtain a road crack image dataset, including transverse cracks, longitudinal cracks, and cracks with irregular shapes and different forms;

[0014] Step 2: Construct an ECA-AC-ResUnet road crack segmentation network model, which consists of an encoder part, a decoder part and corresponding jump connections; the encoder part uses U-Net as the backbone network, introduces a hole residual convolution module, and forms a hole residual convolution module AC-Res to replace the ordinary convolution block; the decoder part includes four upsampling modules, and the corresponding jump connections are performed by splicing the feature map after the upsampling operation with the feature map of the corresponding spatial size in the encoder part; the ECA attention mechanism module is introduced in the decoder part;

[0015] Step 3: Propose a fusion loss function, which is composed of the cross entropy loss function of the binary classification and the Dice loss function;

[0016] Step 4: Use the data set in step 1 for training, and use the trained ECA-AC-ResUnet road crack segmentation network model to perform crack image segmentation.

[0017] Furthermore, the upsampling module includes upsampling, skip connection and double convolution structure. The feature map after the upsampling operation is spliced ​​with the feature map of the corresponding spatial size of the encoder part. The skip connection ensures that the decoder part can directly obtain the feature details obtained by the encoder part and restore the spatial information lost in the image encoder stage; the spliced ​​feature map enters the double convolution structure for refinement; a defect correction module is also introduced in the upsampling module, which consists of a deep supervision mechanism and an adaptive refinement mechanism. The deep supervision mechanism generates an output feature map with semantic information and detail information in the four upsampling processes of the decoder part, and provides auxiliary losses for the upsampling features of each layer for supervision signals; the feature map of the upsampling output of each layer of the decoder part is further optimized through the adaptive refinement mechanism.

[0018] Furthermore, the dual convolution structure includes two consecutive convolution operations, each convolution is followed by a batch normalization layer and a ReLU activation layer, the dual convolution structure also embeds an ECA attention mechanism, and generates channel weights through one-dimensional convolution; the feature map enhanced by the ECA attention mechanism is further fused through the adaptive refinement mechanism in the defect correction module, and finally the fused feature map is output to the next upsampling module until the four upsampling modules of the decoder part are completed.

[0019] Furthermore, the deep supervision mechanism adds a 1×1 convolutional layer D on each upsampled and enhanced output feature map of the decoder part. i , generate the intermediate feature prediction result P i , and obtain the auxiliary loss supervision signal of the feature map; compare the generated prediction result with the output of the real pixel position, and use the IOU loss function to calculate and obtain the auxiliary loss L auxi , the auxiliary losses obtained from all upsampling layer outputs are weighted summed to obtain the loss L of the deep supervision mechanism deep ; Specifically:

[0020] P i =D i (F outi )

[0021]

[0022] Among them, F outi is the output feature map after each upsampling enhancement, ground_truth is the output of the real pixel position, F outi ∩ground_truth is the output of the positive sample pixel position in each upsampling layer, γ i is the weight coefficient of each auxiliary loss.

[0023] Furthermore, the specific process of the adaptive refinement mechanism is as follows:

[0024] First, the feature map F of each layer upsamples the output outi The refined convolutional layer is further processed, and the refined convolutional layer uses a hole residual convolution module to extract the detail information of the crack edge to obtain the refined feature map F aci ; Then, the refined feature map F aci The adaptively refined feature map F is generated by fusing the original upsampled output feature map refinei ; Finally, the adaptively refined feature map is input into the next upsampling layer to repeat the above operation; specifically:

[0025] F aci =C aci (F outi )

[0026] F refinei =F outi +C aci (F aci )

[0027] Among them, F refinei is the output feature map after adaptive refinement, C aci It is a refinement convolution layer.

[0028] Furthermore, the process of the ECA attention module is as follows:

[0029] Calculate global average pooling: In each upsampling process, the ECA attention module compresses the feature map into a single scalar value through global average pooling. For a given feature map F, the number of channels is designed to be C, the width is W, and the height is H. The global average pooling operation compresses the spatial information of the feature map into a global feature vector V for each channel, and obtains the global response of each channel:

[0030]

[0031] Where V∈R C×1 , the feature vector V can summarize the global information of each channel;

[0032] 1D convolution processing: The pooled feature vector V is passed to a 1D convolution operation with a convolution kernel size of 3, the convolution operation is applied on the channel dimension, and the interaction relationship between channels is learned through local connection and weight sharing to complete the adaptive channel weight W eca The specific formula for the generation of is:

[0033] W eca =Sigmoid(Conv1D(V,kernel_size=3))

[0034] Among them, kernel_size represents the convolution kernel size, Conv1D represents the 1D convolution operation, and Sigmoid represents the activation function;

[0035] Adaptive recalibration: After completing the 1D convolution, the generated attention weights are nonlinearly transformed through the Sigmoid activation function, and then multiplied with the original feature map to achieve adaptive feature recalibration. The adaptive feature recalibration process is:

[0036]

[0037] Among them, F′ is the enhanced feature map, Represents a channel-by-channel multiplication operation.

[0038] Furthermore, the fusion loss in step 4 is specifically:

[0039]

[0040] Among them, x i and i are the elements of the prediction and target labels, respectively. N is the total number of samples in the training data set. n is the total number of elements. smooth is the smoothing coefficient, which is set to a small positive number. σ(x i ) is x i Sigmoid function, α and β are the weights of the cross entropy loss function and Dice loss function of the two categories respectively.

[0041] Furthermore, the auxiliary loss of each upsampling module is introduced on the basis of the fusion loss, and the overall loss is determined as the weighted sum of the fusion loss and the auxiliary loss of each upsampling module, that is, the sum of the fusion loss and the loss of the deep supervision mechanism. The formula is:

[0042] L total =L combined +L deep

[0043] Among them, L total is the total loss, L combined is the fusion loss, L deep is the loss of the deep supervision mechanism.

[0044] Beneficial effects:

[0045] The present invention uses U-Net as the backbone network, adds a hole residual convolution module in the encoder stage, and adds an ECA attention mechanism and a defect correction module in the decoder stage. The accuracy of crack image segmentation is improved by the improved convolution module and attention mechanism module. The use of model jump connections allows the network to learn slight adjustments relative to the expected output, thereby improving the generalization ability of the model. In terms of dealing with crack integrity and continuity, a fusion loss function combining the cross entropy loss of binary classification and the Dice loss is proposed, which can maintain the overall coherence and integrity of the crack area while ensuring pixel-level classification accuracy. By introducing an auxiliary loss mechanism, the network can be gradually optimized at each upsampling layer, further improving the recognition ability of small cracks and complex background cracks, and ensuring the stable performance of the model in various challenging situations.

[0046] By introducing the ECA attention mechanism, the present invention can adaptively recalibrate the weights of channel features, thereby enhancing the network's attention to crack features. The ECA mechanism achieves efficient channel attention allocation through global average pooling and depth-separable convolution, allowing the network to focus on the crack area more accurately. Combining the AC-Res module of the dilated convolution and residual network, the model can expand the receptive field and capture the global features of the crack while avoiding the increase in computational complexity. These improvements work together to effectively improve the accuracy of crack image segmentation, especially when facing complex or long cracks, the segmentation results are more accurate.

[0047] In order to further expand the receptive field of the model and improve the backbone network's ability to extract crack features, the present invention introduces a void residual convolution module in the encoder stage to form a void residual convolution module (AC-Res module) to replace the ordinary convolution block. Without significantly increasing the computational complexity, the receptive field can be effectively improved, the model can capture the global information of the crack, and the gradient vanishing problem can be avoided.

[0048] In order to improve the model's extraction effect on the fine detail features of cracks, the present invention proposes an ECA attention mechanism and a defect correction module in the decoder stage. The ECA attention mechanism has adaptive characteristics and can dynamically adjust the model's attention details according to the characteristics of cracks in the input image. Through the deep supervision mechanism and adaptive refinement mechanism, the defect correction module can learn the multi-level details of the crack area on the feature maps of different scales in the decoder stage of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is the segmentation prediction graph of the improved model of the present invention.

[0050] Figure 2 It is the model workflow diagram of the present invention.

[0051] Figure 3 Flowchart of the ECA attention mechanism of the present invention.

[0052] Figure 4 Schematic diagram of the structure of the improved convolution block of the present invention.

[0053] Figure 5 Schematic diagram of the ECA-AC-ResUnet network structure of the present invention.

[0054] Figure 6 It is a schematic diagram of the structure of the defect correction module of the present invention. DETAILED DESCRIPTION

[0055] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0056] The present invention discloses a crack image segmentation method based on an ECA-AC-ResUnet network, which specifically comprises the following steps:

[0057] Step 1: The present invention selects a road crack image dataset for training and testing, including crack images of different forms such as transverse cracks, longitudinal cracks and irregular shapes, and contains a total of 1539 crack images for ECA-AC-ResUnet image segmentation network training and testing. The data is clear and has no obvious noise. Then, the training set, validation set and test set are distinguished at 7:2:1, and 1100 training sets, 300 validation sets and 139 test sets are obtained respectively. The training set, validation set and test set are stored in the corresponding folders train_imgs, val_imgs and test_imgs respectively.

[0058] The research target of this invention is road cracks, which belongs to the binary classification problem of cracks and their background in an image. Therefore, the background is defined as 0, which is set to black, and the crack is defined as 1, which is set to white. A binary mask image corresponding to the original crack image is obtained, and it can be ensured that each crack area has an accurate pixel-level label. The prepared binary mask image is saved according to the training set and the validation set, and stored in the train_masks folder and the val_masks folder respectively.

[0059] Step 2: Construct the ECA-AC-ResUnet road crack segmentation network model, which consists of an encoder part, a decoder part and corresponding jump connections, aiming to accurately identify the multi-scale features of road cracks.

[0060] 1. The core goal of the encoder module is to extract the multi-scale semantic information of the input crack image and generate a gradually downsampled feature map. The present invention uses U-Net as the backbone network and innovatively adopts the hole residual convolution module as the feature extraction module of the crack image to replace the original ordinary convolution block of the backbone network. The convolution module is improved as follows: Figure 4 As shown in the figure, the improvement results are specifically reflected in the following aspects:

[0061] 1. The dilated convolution expands the receptive field by introducing gaps in the convolution kernel, allowing the network to capture a wider range of contextual information without increasing the amount of computation and parameters.

[0062] 2. The dilated residual convolution combines the advantages of dilated convolution and residual connection, which can more effectively retain and transmit information, avoid the gradient vanishing problem, and accelerate the convergence speed of the network. This can not only extract richer image features, but also improve the training efficiency and performance of the network.

[0063] 3. Crack images usually contain texture and structural features of different scales. The atrous residual convolution module can fuse multi-scale information through different atrous rates to improve the detection and segmentation effects of cracks.

[0064] like Figure 5 As shown, the following is a detailed description of the encoder part:

[0065] The encoder extracts the crack image from the initial features, and uses the dilated residual convolution, batch normalization, and ReLU activation function to reduce the input image size by half of the original crack image, while increasing the number of feature channels to extract low-level feature information in the image. Then, the encoder passes through four dilated residual convolution modules in sequence, each of which contains dilated convolution layers, named layer1, layer2, layer3, and layer4. Therefore, a wider range of receptive fields can be obtained during the encoding process, and multi-scale global and local semantic information can be gradually extracted to capture a wider range of spatial context information, thereby effectively covering crack areas of different sizes.

[0066] The first layer of residual blocks (layer1): The input features of the crack image first enter the first layer of residual blocks, which is mainly used to extract the preliminary features of the image. The receptive field is expanded through the dilation rate = 2, and the detailed features of the crack edge are significantly highlighted. The feature map after the convolution operation is processed again through batch normalization and ReLU activation function to stabilize the feature distribution. After the first layer of residual blocks, the size of the image is reduced to 1 / 4 of the original input, but the number of channels is increased to 64 to retain richer detail information.

[0067] Second layer residual block (layer2): The second layer residual block further processes the feature map output by the first layer residual block, and its hole rate is set to 3 to enhance the diversity of crack features. The main task of this layer is to further expand the receptive field so that the model can perform pattern recognition in a larger spatial range, especially for target features with longer lengths such as cracks. The output feature map size of the second layer residual block continues to shrink to 1 / 8 of the original input, but the number of channels increases to 128, indicating that it has learned higher-level feature information.

[0068] The third layer residual block (layer3): By setting the dilated convolution with a dilated rate of 4, the third layer residual block can further expand the receptive field without reducing the resolution. At this stage, the number of channels is increased to 256, and the feature map size is reduced to 1 / 16 to ensure that more local and global features of the crack area are learned. This ensures that the model can capture the overall shape of the crack and identify small-scale detail areas.

[0069] The fourth layer residual block (layer4): The fourth layer residual block further increases the receptive field through convolution with a dilation rate of 5, allowing the model to capture long-distance dependencies between features. The main purpose of the design is to help the model learn the semantic context information of cracks more comprehensively to meet the detection needs of larger cracks. Finally, after the fourth layer residual block, the feature map size is finally reduced to 1 / 32 of the original image, and the number of channels of the feature map is increased to 512.

[0070] 2. The decoder part mainly includes four upsampling modules, named Up1, Up2, Up3 and Up4. The upsampling modules gradually restore the spatial size of the feature map and restore the image details by combining the feature information from the encoder. Each upsampling module includes upsampling (Up Sample), skip connection (Skip Connection) and double convolution structure (Double Conv).

[0071] The specific workflow of the decoder stage is as follows:

[0072] The first upsampling module Up1: The model upsamples the feature map from the last layer of the encoder, using a bilinear interpolation method to double the spatial size of the feature map, allowing the model to gradually restore the spatial resolution of the original image. The formula is:

[0073] F up1 =Bilinear(F 1 , scale_factor = 2)

[0074] Among them, F 1 is the input feature map, F up1is the feature map after the first upsampling, and scale_factor is the bilinear interpolation upsampling ratio.

[0075] Skip Connection: Since it is difficult to fully recover the details lost in the encoder layer-by-layer downsampling by upsampling alone, especially for small structures such as cracks, skip connections are designed in the model decoder. Skip connections pass the multi-scale semantic information generated by each stage of the encoder to the corresponding decoder stage through element-by-element addition, effectively fusing local and global features. Skip connections combine the low-level and high-level fine-grained features of the encoder, allowing the model to retain spatial details during the decoding process, further improving the accuracy of crack identification.

[0076] Specifically, the feature map F after the upsampling operation up1 The feature map F of the corresponding spatial size in the encoder enc1 The concatenation operation (Concat) is performed, which can retain more local detail information in the crack image. The jump connection ensures that the decoder can directly obtain the feature details obtained by the encoder and restore the spatial information lost during the image downsampling process. The feature map F after the jump connection cat1 The formula is:

[0077] F cat1 =Concat(F up1 ,F enc1 )

[0078] Double Convolution Structure (Double Conv): The double convolution structure contains two consecutive convolution operations (Conv2d), each of which includes a convolution layer, a batch normalization layer (BatchNormalization) and a ReLU activation layer. The double convolution structure further refines the feature map after fusion, deeply extracts the boundary feature information of the crack edge, enhances the nonlinear expression ability of the model, and improves the robustness of the model.

[0079] Specifically, the model concatenates the feature map F cat The double convolution structure is used to extract the characteristic details in the image. In addition, each convolution layer is embedded with the ECA attention mechanism to strengthen the characteristic expression of the crack area by dynamically adjusting the channel weights. The ECA module generates the channel weight matrix W eca1 , multiplying it with the convolution output channel by channel, further highlighting the model's ability to focus on the crack area.

[0080] Convolution output F of the double convolution structure dc1 It can be expressed as:

[0081] F dc1=RELU(BatchNorm(Conv(F cat ))

[0082] The weight matrix W of the ECA attention mechanism eca It can be expressed as:

[0083] W eca1 =Sigmoid(Conv1D(GlobalAvgPool(F dc1 ))

[0084] Enhanced feature output F out1 for:

[0085] F out1 =F dc1 ×W eca1

[0086] F ac1 =C ac1 (F out1 )

[0087] F refine1 =F out1 +C ac1 (F ac1 )

[0088] The enhanced output feature map is then further refined through the adaptive refinement mechanism of the defect correction module to generate a fused adaptive refinement feature map F refine1 , and input it to the next upsampling layer to continue the next upsampling operation.

[0089] The second upsampling module Up2: The adaptive refinement feature map F obtained in the first upsampling module refine1 Upsampling is performed to further double the spatial resolution of the output feature map. The feature map F after the second upsampling up2 The formula is:

[0090] F up2 =Bilinear(F refine1 , scale_factor = 2)

[0091] Skip connection: The feature map F after the second upsampling up2 and the corresponding layer feature map F in the encoder enc2 Concatenate. Feature map F after skip connection cat2 The formula is:

[0092] F cat2 =Concat(F up2 ,F enc2 )

[0093] Double convolution structure: After the feature information of the crack image is spliced, it is refined using the double convolution structure, and the ECA attention module is used to focus on the crack features. The attention weight can ensure that the crack area is effectively magnified. Specifically, the convolution output F of the double convolution structure is dc2 It can be expressed as:

[0094] F dc2 =RELU(BatchNorm(Conv(F cat2 ))

[0095] The weight matrix W of the ECA attention mechanism eca2 It can be expressed as:

[0096] W eca2 =Sigmoid(Conv1D(GlobalAvgPool(F dc2 ))

[0097] The further enhanced feature output F out2 for:

[0098] F out2 =F dc2 ×W eca2

[0099] F ac2 =C ac2 (F out2 )

[0100] F refine2 =F out2 +C ac2 (F ac2 )

[0101] The enhanced output feature map is then passed through the adaptive refinement part of the defect correction module to further extract detail information to generate the fused adaptive refinement feature map F refine2 , and input it into the third upsampling module to continue the next upsampling operation.

[0102] The third upsampling module Up3: The feature map F obtained in the second upsampling module refine2 After upsampling, the spatial size of the feature map is further increased, and it is restored to a size close to the input image. The feature map F after the third upsampling up3 The formula is:

[0103] F up3 =Bilinear(F refine2 , scale_factor = 2)

[0104] Skip connection: The feature map F after the third upsampling up3 and the corresponding layer feature map F in the encoder enc3 After splicing, the decoder can obtain more detailed information. Feature map F after skip connection cat3 The formula is:

[0105] F cat3 =Concat(F up3 ,F enc3 )

[0106] Double convolution structure: Specifically, the convolution output F of the double convolution structure dc3 It can be expressed as:

[0107] F dc3 =RELU(BatchNorm(Conv(F cat3 ))

[0108] The weight matrix W of the ECA attention mechanism eca3 It can be expressed as:

[0109] W eca3 =Sigmoid(Conv1D(GlobalAvgPool(F dc3 ))

[0110] The further enhanced feature output F out3 for:

[0111] F out3 =F dc3 ×W eca3

[0112] F ac3 =C ac3 (F out3 )

[0113] F refine3 =F out3 +C ac3 (F ac3 )

[0114] The enhanced output feature map is then passed through the adaptive refinement part of the defect correction module to further extract detail information to generate the fused adaptive refinement feature map F refine3 , and input into the fourth upsampling module to continue the next upsampling operation.

[0115] The fourth upsampling module Up4: the feature map F obtained in the third upsampling module refine3 Upsampling is performed to double the spatial resolution of the output feature map, and the resolution information of the original image is restored. The feature map F after the fourth upsamplingup4 The formula is:

[0116] F up4 =Bilinear(F refine3 , scale_factor = 2)

[0117] Skip connection: The feature map F after the fourth upsampling up4 and the corresponding layer feature map F in the encoder enc4 Splicing is performed to obtain complete details of the crack. Feature map F after jump connection cat4 The formula is:

[0118] F cat4 =Concat(F up4 ,F enc4 )

[0119] Double convolution structure: Specifically, the convolution output F of the double convolution structure dc4 It can be expressed as:

[0120] F dc4 =RELU(BatchNorm(Conv(F cat4 ))

[0121] The weight matrix W of the ECA attention mechanism eca4 It can be expressed as:

[0122] W eca4 =Sigmoid(Conv1D(GlobalAvgPool(F dc4 ))

[0123] The further enhanced feature output F out4 for:

[0124] F out4 =F dc4 ×W eca4

[0125] F ac4 =C ac4 (F out4 )

[0126] F refine4 =F out4 +C ac4 (F ac4 )

[0127] The enhanced output feature map is then passed through the adaptive refinement part of the defect correction module to further extract detail information to generate the fused adaptive refinement feature map F refine4 , output the feature map to the 1×1 convolution module.

[0128] 1×1 convolution module: 1×1 convolution can act on the channel at each position of the feature map without changing the spatial resolution, effectively realizing channel compression and category mapping. The rich features extracted in the decoder stage are integrated into a single-channel crack segmentation probability map, and the image thresholding is performed to obtain a binary segmentation mask image of the original image.

[0129] Specifically, the enhanced feature output map F refine4 The dimensionality reduction process is performed through 1×1 convolution to reduce the number of feature channels to the number of output categories specified by the original image to generate a segmentation mask image of the input crack image. The mask output can be expressed as:

[0130]

[0131] Among them, c represents the number of channels of the input feature map, and a single-channel segmentation mask is obtained after convolution operation.

[0132] For the decoder part of the network, in order to allow the model to learn the multi-level details of the crack area on the feature maps of different scales, and then supplement and correct the areas where the crack image output by the decoder may be missed or misjudged, the accuracy of segmentation is improved. The present invention innovatively introduces a defect correction module, which mainly consists of two parts: a deep supervision mechanism and an adaptive refinement mechanism. The contents of the two mechanisms are explained in detail below.

[0133] Deep supervision mechanism: The model generates output feature maps with semantic information and detail information in the four upsampling processes of the decoder through a deep supervision mechanism, and performs multi-level optimization learning on the crack area, significantly reducing missed detection and false detection in the segmentation process. Specifically, a 1×1 convolution layer D is added to each upsampled and enhanced output feature map of the decoder. i , generate the intermediate feature prediction result P i , and obtain the auxiliary loss supervision signal of the feature map. Then, compare the generated prediction result with the output of the real pixel position, and use the IOU loss function to calculate and obtain the auxiliary loss L auxi Finally, the auxiliary losses obtained from all upsampling layer outputs are weighted summed to obtain the loss L of the deep supervision mechanism. deep The main formula is:

[0134] P i =D i (F outi )

[0135] L auxi =IOU(P i ,ground_truth)

[0136]

[0137] Among them, F outi is the output feature map after each upsampling enhancement, ground_truth is the output of the real pixel position, γ i is the weight coefficient of each auxiliary loss.

[0138] Adaptive refinement mechanism: The model uses an adaptive refinement mechanism to further optimize the feature maps of each upsampled output layer of the decoder, making the edges of cracks and complex area features in the image more obvious and accurate.

[0139] Specifically, first, the feature map F of each layer upsampled output outi The refined convolutional layer is further processed, and the refined convolutional layer uses a hole residual convolution module to extract the detail information of the crack edge to obtain the refined feature map F aci ; Then, it is fused with the original upsampled output feature map to generate the adaptively refined feature map F refinei ; Finally, the adaptively refined feature map is input into the next upsampling layer to repeat the above operation. The optimization is performed step by step from the shallow layer to the deep layer, so that the edge features and complex details of the cracks can be more finely identified. The main formula is:

[0140] F aci =C aci (F outi )

[0141] F refinei =F outi +C aci (F aci )

[0142] Among them, F refinei is the output feature map after adaptive refinement, C aci It is a refinement convolution layer.

[0143] Step 3: To further extract and restore the local and global feature information in step 2, the decoder introduces the ECA attention mechanism module, which mainly generates an adaptive channel weight W for each channel. eca , improve the attention of each feature channel to the input crack image, so that the model can focus more on the recognizable areas and reduce the influence of redundant features. The processing flow of the ECA module is as follows:

[0144] Calculate global average pooling: In each upsampling process, the ECA attention module compresses the feature map into a single scalar value through global average pooling in order to obtain the global feature information of each channel. For a given feature map F, the number of channels is designed to be C, the width is W, and the height is H. The global average pooling operation compresses the spatial information of the feature map into a global feature vector V for each channel, and obtains the global response of each channel:

[0145]

[0146] Where V∈R C×1 , the feature vector V can summarize the global information of each channel, which can help the network identify salient areas and greatly reduce the computational complexity.

[0147] 1D convolution processing: The pooled feature vector V is passed to a 1D convolution operation with a convolution kernel size of 3. The 1D convolution can apply convolution operations on the channel dimension and learn the interaction relationship between channels through local connections and weight sharing to complete the adaptive channel weight W. eca The specific formula is:

[0148] W eca =Sigmoid(Conv1D(V,kernel_size=3))

[0149] Adaptive recalibration: After completing the 1D convolution, the generated attention weights are nonlinearly transformed through the Sigmoid activation function, and then multiplied with the original feature map to achieve adaptive feature recalibration. The process of adaptive feature recalibration is:

[0150]

[0151] Among them, F′ is the enhanced feature map, Represents a channel-by-channel multiplication operation.

[0152] The ECA attention mechanism introduces global feature information, allowing the network to focus on the crack area more effectively, reduce the interference of background noise on the segmentation results, and improve segmentation accuracy. Traditional attention mechanisms often require a lot of computing resources, while the ECA attention mechanism uses a lightweight design to significantly reduce the amount of computing while ensuring segmentation accuracy.

[0153] Step 4: The fusion loss function proposed in the present invention is composed of the binary cross entropy loss function and the Dice loss function. This fusion loss function can not only enable the network to accurately segment the cracks at the pixel level, but also maintain the overall continuity and integrity of the cracks. And on this basis, the auxiliary loss in the deep supervision mechanism is innovatively introduced. The auxiliary loss supervision signal can be output at each upsampling layer, and the auxiliary losses of the output results of these layers and the true labels are calculated respectively. For the output of each upsampling module of the decoder, the deep supervision mechanism will calculate the auxiliary loss L of each layer. auxi , the specific formula is:

[0154]

[0155] Among them, F outi is the output of each upsampling module. outi ∩ground_truth is the output of the positive sample pixel position in each upsampling layer.

[0156] The introduction of auxiliary loss not only makes the model more detailed in extracting local features, but also promotes the learning of the network at different scales, making the overall segmentation effect more stable and accurate, further improving the robustness and accuracy of crack segmentation.

[0157] Specifically, for the Dice loss function, we first need to calculate the similarity coefficient (DiceSimilarity Coefficient, DSC) between two samples. DSC is a set similarity measure. It is defined as:

[0158]

[0159] Where X is the prediction result, Y is the true label, |X| and |Y| are the total number of elements in X and Y respectively, and |X∩Y| is the intersection of X and Y.

[0160] The formula of Dice Loss is defined as:

[0161]

[0162] Among them, x i and i are the elements of the prediction and target labels respectively, n is the total number of elements, and smooth is the smoothing parameter, which is set to a small positive number.

[0163] Specifically, for the cross entropy loss function of the two categories, the BCEWithLogitsLoss formula is defined as:

[0164]

[0165] Where N is the total number of samples in the training data set, σ(x i ) is x i The Sigmoid function is

[0166] The mixed loss CombinedLoss formula is defined as:

[0167] L combined =α×L BCE +β×L Dice

[0168] Among them, L BCE is the cross entropy loss for binary classification, L Dice is the Dice loss, α and β are the weights of the two losses respectively. The complete formula of the fusion loss function is:

[0169]

[0170] Finally, the overall loss is determined to be the weighted sum of the fusion loss and the auxiliary loss of each layer of upsampling modules, that is, the sum of the fusion loss and the loss of the deep supervision mechanism. Through the innovation of the loss function, the model can gradually approach the target segmentation result at each layer. The overall loss function formula is:

[0171] L total =L combined +L deep

[0172] Among them, L total is the total loss, L combined is the fusion loss, L deep is the loss of the deep supervision mechanism.

[0173] Step 5: Use random seeds during training to ensure the repeatability of the results, thereby realizing the influence of fixed factors on the training results. Specifically, set the random seed equal to 55 to ensure the consistency of multiple training results. Use the Adam optimizer for model training, which has a faster convergence speed and strong robustness, which is conducive to dealing with the unbalanced data problem in the crack segmentation task. Set the initial value of the learning rate to 1e -6 , the batch size is 4, the image resolution is 448×448, and the training rounds are 60. The improved crack segmentation model in steps 2-4 is iteratively trained using the road crack dataset divided in step 1, and accurate crack segmentation is achieved at the pixel level through the optimization of the overall loss function combining the fusion loss and the auxiliary loss.

[0174] After each round of training, record the various evaluation indicators of the validation set, including but not limited to precision, recall, final loss function and F1 score, and save the best model weights according to the changes in the loss function coefficients. Finally, save the model's training weight file in .pt format and name it "ECA-AC-ResUnet_Model.pt".

[0175] Step 6: Use the model weight file trained in step 5 to perform the model testing process. The model implements the binary segmentation of cracks by loading the weight file and generates a binary mask image corresponding to the original image of the test set. At the same time, the detection performance of the model on different cracks is analyzed in detail to verify the accuracy and robustness of the model.

[0176] The experimental environment of the present invention is configured to use the Windows operating system, use the NVIDIA GeForce RTX 4060 graphics card, and use the PyTorch deep learning framework. The specific configuration is shown in Table 1.

[0177] Table 1 Experimental environment configuration

[0178]

[0179]

[0180] Figure 1 This is the segmentation prediction map of the improved model of the present invention. From the four annotated images, it can be seen that only the main trunk area of ​​the crack is annotated, but the boundary information is relatively rough and not clear enough, and there is a phenomenon of discontinuity of the cracks, and some small cracks cannot be annotated. From the four prediction maps, it can be clearly seen that the improved model has significant advantages in crack detection: for example, the model completely segments the crack cross structure, accurately captures the complete trunk of the crack, retains the overall shape of the crack during segmentation, has clear edges and smooth transitions, and is more accurate in identifying the edge area of ​​the crack. At the same time, the improved model is more accurate and coherent in segmenting small cracks, and can effectively identify different forms of road cracks, meeting the requirements for crack segmentation precision and accuracy.

[0181] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit of the present invention should be included in the protection scope of the present invention.

Claims

1. A crack image segmentation method based on ECA-AC-ResUnet network, characterized in that: The steps include: Step 1: Obtain a road crack image dataset, including transverse cracks, longitudinal cracks, and cracks with irregular shapes and different forms; Step 2: Construct an ECA-AC-ResUnet road crack segmentation network model, which consists of an encoder part, a decoder part and corresponding jump connections; the encoder part uses U-Net as the backbone network, introduces a hole residual convolution module, and forms a hole residual convolution module AC-Res to replace the ordinary convolution block; the decoder part includes four upsampling modules, and the corresponding jump connections are performed by splicing the feature map after the upsampling operation with the feature map of the corresponding spatial size in the encoder part; the ECA attention mechanism module is introduced in the decoder part; Step 3: Propose a fusion loss function, which is composed of the cross entropy loss function of the binary classification and the Dice loss function; Step 4: Use the data set in step 1 for training, and use the trained ECA-AC-ResUnet road crack segmentation network model to perform crack image segmentation.

2. The crack image segmentation method based on the ECA-AC-ResUnet network according to claim 1 is characterized in that: The upsampling module includes upsampling, skip connection and double convolution structure. The feature map after the upsampling operation is spliced ​​with the feature map of the corresponding spatial size of the encoder part. The skip connection ensures that the decoder part can directly obtain the feature details obtained by the encoder part and restore the spatial information lost in the image encoder stage; the spliced ​​feature map enters the double convolution structure for refinement processing; a defect correction module is also introduced in the upsampling module, which consists of a deep supervision mechanism and an adaptive refinement mechanism. The deep supervision mechanism generates an output feature map with semantic information and detail information in the four upsampling processes of the decoder part, and provides auxiliary losses for the upsampling features of each layer for supervision signals; the feature map of the upsampling output of each layer of the decoder part is further optimized through the adaptive refinement mechanism.

3. The crack image segmentation method based on the ECA-AC-ResUnet network according to claim 2 is characterized in that: The double convolution structure includes two consecutive convolution operations, each convolution is followed by a batch normalization layer and a ReLU activation layer. The double convolution structure also embeds the ECA attention mechanism and generates channel weights through one-dimensional convolution. The feature map enhanced by the ECA attention mechanism is further fused through the adaptive refinement mechanism in the defect correction module, and the fused feature map is finally output to the next upsampling module until the four upsampling modules of the decoder part are completed.

4. The crack image segmentation method based on the ECA-AC-ResUnet network according to claim 2 is characterized in that: The deep supervision mechanism adds a 1×1 convolutional layer D on each upsampled and enhanced output feature map of the decoder part. i , generate the intermediate feature prediction result P i , and obtain the auxiliary loss supervision signal of the feature map; compare the generated prediction result with the output of the real pixel position, and use the IOU loss function to calculate and obtain the auxiliary loss L auxi , the auxiliary losses obtained from all upsampling layer outputs are weighted summed to obtain the loss L of the deep supervision mechanism deep ; Specifically: P i =D i (F outi ) Among them, F outi is the output feature map after each upsampling enhancement, ground_truth is the output of the real pixel position, F outi ∩ground_truth is the output of the positive sample pixel position in each upsampling layer, γ i is the weight coefficient of each auxiliary loss.

5. The crack image segmentation method based on the ECA-AC-ResUnet network according to claim 2 is characterized in that: The specific process of the adaptive refinement mechanism is as follows: First, the feature map F of each layer upsamples the output outi The refined convolutional layer is further processed, and the refined convolutional layer uses a hole residual convolution module to extract the detail information of the crack edge to obtain the refined feature map F aci ; Then, the refined feature map F aci The adaptively refined feature map F is generated by fusing the original upsampled output feature map refinei ; Finally, the adaptively refined feature map is input into the next upsampling layer to repeat the above operation; specifically: F aci =C aci (F outi ) F refinei =F outi +C aci (F aci ) Among them, F refinei is the output feature map after adaptive refinement, C aci It is a refinement convolution layer.

6. The crack image segmentation method based on the ECA-AC-ResUnet network according to claim 3 is characterized in that: The process of the ECA attention module is as follows: Calculate global average pooling: In each upsampling process, the ECA attention module compresses the feature map into a single scalar value through global average pooling. For a given feature map F, the number of channels is designed to be C, the width is W, and the height is H. The global average pooling operation compresses the spatial information of the feature map into a global feature vector V for each channel, and obtains the global response of each channel: Where V∈R C×1 , the feature vector V can summarize the global information of each channel; 1D convolution processing: The pooled feature vector V is passed to a 1D convolution operation with a convolution kernel size of 3, the convolution operation is applied on the channel dimension, and the interaction relationship between channels is learned through local connection and weight sharing to complete the adaptive channel weight W eca The specific formula for the generation of is: W eca =Sigmoid(Conv1D(V,kernel_size=3)) Among them, kernel_size represents the convolution kernel size, Conv1D represents the 1D convolution operation, and Sigmoid represents the activation function; Adaptive recalibration: After completing the 1D convolution, the generated attention weights are nonlinearly transformed through the Sigmoid activation function, and then multiplied with the original feature map to achieve adaptive feature recalibration. The adaptive feature recalibration process is: Among them, F′ is the enhanced feature map, Represents a channel-by-channel multiplication operation.

7. The crack image segmentation method based on the ECA-AC-ResUnet network according to claim 1, characterized in that: The fusion loss in step 4 is specifically: Among them, x i and i are the elements of the prediction and target labels, respectively. N is the total number of samples in the training data set. n is the total number of elements. smooth is the smoothing coefficient, which is set to a small positive number. σ(x i ) is x i Sigmoid function, α and β are the weights of the cross entropy loss function and Dice loss function of the two categories respectively.

8. The crack image segmentation method based on the ECA-AC-ResUnet network according to claim 4 is characterized in that: On the basis of fusion loss, the auxiliary loss of each upsampling module is introduced, and the overall loss is determined as the weighted sum of fusion loss and auxiliary loss of each upsampling module, that is, the sum of fusion loss and loss of deep supervision mechanism. The formula is: L total =L combined +L deep Among them, L total is the total loss, L combined is the fusion loss, L deep is the loss of the deep supervision mechanism.

Citation Information

Cited By

  • Intelligent detection method and system for Mura defect in OLED display panel

    CN120543558A

  • Coal rock fracture intelligent extraction method based on improved U-Net

    CN120563993A

  • Image segmentation method, device and equipment and readable storage medium

    CN120655655A

  • Method and system for acquiring cracking parameters of wall painting cracks

    CN120953348A

  • A method and system for obtaining mural crack parameters

    CN120953348B