Concrete phase segmentation method based on lightweight deep learning and SE attention mechanism

By reducing the number of convolutional layer channels in the U-Net network and embedding the SE attention mechanism, the SE U-Net model was constructed, which solved the problem of low phase segmentation accuracy of concrete in the prior art, and achieved efficient and accurate phase segmentation effect.

CN119919429AActive Publication Date: 2025-05-02SOUTHEAST UNIV

Patent Information

Application Number
CN202510092362.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-02
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The prior art is difficult to deploy efficiently in a limited computing environment, and the problems of insufficient details recognition and unreasonable attention distribution in concrete images, resulting in low accuracy of concrete phase segmentation.

Method used

Based on the U-Net network architecture, the model is lightweighted by reducing the number of convolutional layers, and embedded the SE attention mechanism in the convolutional layers of the encoder and decoder to build the SE U-Net model to improve the weight allocation ability of the model to key features.

Benefits of technology

The accuracy of concrete phase segmentation was significantly improved, mIoU reached 88.5%, an increase of 9.7% compared with the U-Net model, while reducing the amount of model parameters and computing resource consumption, achieving efficient phase segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919429A_ABST
    Figure CN119919429A_ABST
Patent Text Reader

Abstract

The invention discloses a concrete phase segmentation method based on lightweight deep learning and an SE attention mechanism. The method comprises the following steps: acquiring an original concrete image by adopting X-ray tomography; in the pretreatment process, a semi-automatic labeling method is adopted for accurately labeling the phase of the concrete; dividing the data set into a training set, a verification set and a test set; on the basis of a U-Net network architecture, network lightweight is realized by adopting a strategy of reducing the number of channels of a convolutional layer, and an SE attention mechanism is embedded to construct an SE U-Net model; and finally performing model evaluation on the test set through model training and optimization. According to the SE U-Net model constructed through the method, in a concrete phase segmentation task, mIoU reaches 88.5% and is remarkably improved by 9.7% compared with a U-Net model, meanwhile, the model parameter quantity and computing resource consumption are reduced through lightweight design, and efficient and accurate concrete phase segmentation can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of concrete image semantic segmentation, and in particular relates to a concrete phase segmentation method based on lightweight deep learning and SE (Squeeze-and-Excitation) attention mechanism. Background Art

[0002] As a key material widely used in infrastructure construction, the quality and performance of concrete largely depend on the composition, properties and distribution of different phases. Accurately identifying and segmenting different phases in concrete images provides an important basis for concrete quality assessment, structural analysis and durability prediction. However, due to the high heterogeneity inside concrete, the grayscale values ​​of the various phases often overlap, making the automatic segmentation task extremely challenging.

[0003] Traditional methods for concrete phase segmentation mainly rely on image processing techniques such as manual recognition, threshold-based algorithms, and edge detection, which usually require setting parameters based on experience and specific rules. Such methods are not only time-consuming and labor-intensive, but also often have poor effects when dealing with complex heterogeneous features, and cannot effectively deal with noise and artifacts in the image, resulting in low segmentation accuracy.

[0004] In recent years, deep learning technology, especially the U-Net network architecture, has been gradually applied to fields such as building material image analysis due to its excellent performance in tasks such as medical image segmentation. With the symmetrical structure and jump connection of the encoder-decoder, U-Net can learn high-dimensional features through a deep network and achieve good segmentation performance in complex scenes or small sample data.

[0005] However, the U-Net model usually contains a large number of parameters and is difficult to deploy efficiently in an environment with limited computing power. In addition, for concrete images with high heterogeneity, relying solely on the U-Net architecture still faces problems such as insufficient detail recognition and unreasonable attention allocation. Therefore, there is an urgent need for a lightweight deep learning method that can improve segmentation accuracy and has an efficient attention mechanism. Summary of the invention

[0006] Purpose of the invention: In view of the above problems, the purpose of the present invention is to provide a method for concrete phase segmentation based on lightweight deep learning and SE attention mechanism. Based on the U-Net network architecture, the model is lightweight designed by reducing the number of channels in each convolutional layer, and the SE attention module is embedded in the convolutional layers of the encoder and decoder to effectively strengthen the model's weight allocation to key features and achieve high-precision segmentation of concrete phases.

[0007] Technical solution: The concrete phase segmentation method based on lightweight deep learning and SE attention mechanism described in the present invention comprises the following steps:

[0008] (1) Data acquisition: high-resolution images of the internal structure of concrete, i.e., original images, are obtained through X-ray tomography (X-CT) technology;

[0009] (2) Data preprocessing: The original images are preprocessed and the concrete phases are annotated using a semi-automatic annotation method to form an initial data set;

[0010] (3) Data segmentation: divide the initial data set into training set, validation set and test set;

[0011] (4) Lightweight model architecture design: Based on the U-Net network architecture, the number of model parameters is reduced by reducing the number of channels in each convolutional layer, and a lightweight U-Net model is constructed;

[0012] (5) SE attention mechanism integration: The SE (Squeeze-and-Excitation) attention mechanism is embedded in the convolutional layers of the encoder and decoder of the lightweight U-Net model to build the SE U-Net model. The model’s feature extraction and classification capabilities for different concrete phases are enhanced by using channel adaptive weighting.

[0013] (6) Model training: Use the training set obtained in step (3) to train the SE U-Net model, use the validation set obtained in step (3) to validate it, perform hyperparameter tuning based on the validation results, and save the optimal model obtained after training;

[0014] (7) Model testing: The test set obtained in step (3) is input into the optimal model for high-precision phase segmentation, the predicted mask is output, and evaluation indicators are given.

[0015] The concrete phase is one or more of aggregate, cement paste or pores.

[0016] Among them, in step (2), preprocessing includes rotation, cropping, and scaling to unify the image size; the semi-automatic labeling method is to perform preliminary labeling of concrete phases based on thresholds, supplemented by manual correction, and cover the concrete phases with masks respectively.

[0017] In step (3), the training set, validation set and test set are independent of each other; the training set is used for model training, the validation set is used for model tuning, and the test set is used for model evaluation.

[0018] Among them, in step (3), the initial data set is randomly divided into a training set, a validation set and a test set according to a preset ratio. The stratified sampling method is used in the division process to ensure that the data categories of the three subsets are evenly distributed. At the same time, a fixed random seed is set to ensure the repeatability of each random division result.

[0019] Among them, in step (4), the method of reducing the number of channels of each convolutional layer includes performing feature extraction and fusion with a smaller number of channels than the preset number of channels of the U-Net model in the convolution modules of each level of the encoder and the decoder, so as to reduce the overall model parameter amount and computing resource consumption.

[0020] Among them, in step (5), embedding the SE attention mechanism includes the following steps:

[0021] 1) Squeeze: Each channel of the input feature map is spatially compressed through global average pooling to form the global features of each channel. This step can be expressed by formula (1):

[0022]

[0023] Among them, H represents the height of the input feature map, W represents the width of the input feature map, and X c,h,w represents the pixel value of the cth channel at position (h,w) in the input feature map, z c Represents the global features of the cth channel after average pooling;

[0024] 2) Excitation: Use a two-layer fully connected network to transform the channel global features to learn and generate the adjustment weights of each channel. The channel global features are reduced in dimension by the first layer of the fully connected network, and the ReLU activation function is applied to add nonlinear processing. The second layer of the fully connected network expands the feature dimension back to the original number of channels and outputs the final channel weights through the Sigmoid activation function. This step can be expressed by formulas (2) to (4):

[0025] ReLU(x)=max(0,x) (2)

[0026]

[0027] s c =σ(W2·ReLU(W1·z c )) (4)

[0028] Among them, x represents the input value, ReLU is the activation function, σ represents the Sigmoid activation function, W1 represents the dimension reduction weight matrix, W2 represents the dimension increase weight matrix, and s c represents the weight of the cth channel;

[0029] 3) Scale (reweighting): The weight s of each channel adjusts the channel strength of the original input feature map through element-wise multiplication. This step can be expressed by formula (5):

[0030] Y c,h.w =s c×X c,h,w (5)

[0031] Among them, Y c,h.w is the output feature map after attention weighting, X c,h,w is the pixel value of the cth channel at position (h,w) in the input feature map.

[0032] In step (6), the model training uses categorical cross entropy as the loss function, which evaluates the difference between the model output probability distribution and the true mask distribution at the pixel level; the Adam optimizer is used to update the parameters, and the early stopping strategy and the best model preservation strategy are set based on the loss value of the validation set.

[0033] In step (6), data enhancement techniques are used on the training set data, including horizontal flipping, vertical flipping and random rotation, to improve the generalization ability of the model in complex scenarios.

[0034] Among them, in step (7), the evaluation indicators include Accuracy, Precision, Recall, F1-score and mIoU, all expressed in percentage. The higher the value, the better the performance of the model. The accuracy of the segmentation model is quantitatively evaluated by comprehensive analysis of the evaluation indicators; and the segmentation results are visually compared based on the original image, the real mask and the predicted mask.

[0035] The method for concrete phase segmentation based on lightweight deep learning and SE attention mechanism described in the present invention comprises the following steps: obtaining the original image of concrete by X-ray tomography technology; preprocessing the original image and accurately annotating the concrete phase by semi-automatic annotation method; dividing the obtained data set into training set, validation set and test set; on the basis of U-Net network architecture, adopting the strategy of reducing the number of channels of convolution layer to realize network lightweight, and embedding SE attention mechanism in convolution layer to adaptively enhance the response of important feature channels, constructing SE U-Net model; training the model by training set, and using validation set to tune the model, and finally evaluating the model on test set. The results show that the SE U-Net model constructed by the present invention can accurately classify in the concrete phase segmentation task, and the mIoU reaches 88.5%, which is significantly improved by 9.7% compared with the U-Net model. At the same time, the model parameters and computing resource consumption are reduced by lightweight design, and efficient concrete phase segmentation can be achieved.

[0036] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0037] (1) The present invention performs lightweight design based on the U-Net network architecture, significantly reducing the number of model parameters and computing resource consumption while ensuring segmentation accuracy, and has more advantages in hardware resource occupancy and running speed.

[0038] (2) The SE U-Net model constructed by the present invention embeds the SE attention module in the convolutional layers of the encoder and decoder, giving the model the ability to adaptively enhance key feature channels, significantly improving the segmentation accuracy of the model in complex environments with image grayscale overlap or large noise interference, and achieving accurate segmentation of concrete phases.

[0039] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0040] Figure 2 This is the segmentation result diagram of the SE U-Net model. DETAILED DESCRIPTION

[0041] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0042] Example 1

[0043] like Figure 1 As shown in the figure, the concrete phase segmentation method based on lightweight deep learning and SE attention mechanism includes the following steps:

[0044] (1) Data acquisition:

[0045] High-resolution images of the internal structure of concrete were obtained by X-ray tomography (X-CT). Specifically, the scanning direction was from top to bottom, the scanning interval was 0.05 mm, and the images with poor imaging quality at the top and bottom were removed. A total of 900 scanned images were obtained, and the images were numbered in sequence and stored in a designated directory. The above images are all original images.

[0046] (2) Data preprocessing:

[0047] 150 original images are selected from 900 original images by uniformly sampling, that is, one image is selected every five images starting from the first original image to avoid the high similarity between adjacent images affecting the generalization ability of the model; after converting these 150 original images into grayscale images, the image size is unified to 581×581 pixels through operations such as rotation, cropping, and scaling;

[0048] A semi-automatic annotation method is used to annotate the concrete phases. Specifically, the pores and aggregates in the concrete image are first annotated based on the threshold, where the pore threshold is set to 40 and the aggregate threshold is set to 130; then, in the superpixel mode of the MATLAB image segmentor, the preliminary annotation results are manually corrected to optimize the abnormal or edge areas to ensure the accuracy of the annotation; finally, the remaining area of ​​the image is annotated as cement paste, and the corresponding masks are covered for aggregates, pores, and cement paste, respectively, with the masks being 0, 1, and 2, to form a mask image;

[0049] In PyCharm, the pixel values ​​of the selected 150 original images were normalized and the pixel values ​​were divided by 255 to scale the pixel range to the interval [0,1]. The original image size was uniformly scaled to 256×256 pixels using bilinear interpolation. The mask image was scaled to 256×256 pixels using the nearest neighbor interpolation method to avoid category aliasing or blurred boundaries. The mask image was paired with the original image and stored to form an initial data set.

[0050] (3) Data segmentation:

[0051] The initial data set is divided into training set, validation set and test set in a ratio of 6:2:2. The stratified sampling method is used in the division process to ensure that the data categories of the three subsets are evenly distributed. At the same time, a fixed random seed is set to ensure the repeatability of each random division result. The training set, validation set and test set are independent of each other. The training set is used for model training, the validation set is used for model tuning, and the test set is used for model evaluation.

[0052] (4) Lightweight design of model architecture:

[0053] The U-Net model is constructed, and based on its network architecture, the number of channels in each convolutional layer is reduced to reduce the number of model parameters, and a lightweight U-Net model is constructed. The lightweight design reduces the computing resource consumption of the GPU / CPU. The way to reduce the number of channels in the convolutional layer includes the convolutional layers at all levels of the encoder and decoder. The specific architecture of the lightweight U-Net model is as follows:

[0054] Encoder: It is composed of 4 layers of "convolution + maximum pooling (2×2)". The convolution kernel size of each layer is 3×3 and is activated by the linear rectification function (ReLU). The number of channels increases gradually in each layer, but compared with U-Net, the initial number of channels is reduced from 64 to 32, and then reduced to 64, 128, and 256 in subsequent layers respectively;

[0055] The bottom layer of the network: retains 512 channels to capture the complex information of concrete phases;

[0056] Decoder: It is composed of 4 layers of "upper convolution + convolution", the convolution kernel size of each layer is 3×3, the number of channels of each convolution layer is 256, 128, 64, and 32 respectively, and is mapped to the specified number of categories 3 by 1×1 convolution at the end of the network;

[0057] (5) SE attention mechanism integration:

[0058] The SE attention mechanism is embedded in several convolutional layers of the encoder and decoder of the lightweight U-Net model to build the SE U-Net model. The model's feature extraction and classification capabilities for different phases of concrete are enhanced by using channel adaptive weighting. The SE attention mechanism is implemented through the following steps:

[0059] 1) Squeeze: Each channel of the input feature map is spatially compressed through global average pooling to form the global features of each channel. This step can be expressed by formula (1):

[0060]

[0061] Among them, H represents the height of the input feature map, W represents the width of the input feature map, and X c,h,w represents the pixel value of the cth channel at position (h,w) in the input feature map, z c Represents the global features of the cth channel after average pooling.

[0062] 2) Excitation: A two-layer fully connected network is used to transform the channel global features to learn and generate the adjustment weights of each channel. The channel global features are reduced in dimension by the first layer of fully connected network, and the ReLU activation function is applied to add nonlinear processing. The second layer of fully connected network expands the feature dimension back to the original number of channels and outputs the final channel weights through the Sigmoid activation function. This step can be expressed by formulas (2) to (4):

[0063] ReLU(x)=max(0,x) (2)

[0064]

[0065] s c =σ(W2·ReLU(W1·z c )) (4)

[0066] Among them, x represents the input value, ReLU is the activation function, σ represents the Sigmoid activation function, W1 represents the dimension reduction weight matrix, W2 represents the dimension increase weight matrix, and s c Represents the weight of the c-th channel.

[0067] 3) Scale: The weight s of each channel adjusts the channel strength of the original input feature map through element-wise multiplication. This step can be expressed by formula (5):

[0068] Y c,h.w =s c ×X c,h,w (5)

[0069] Among them, Y c,h.w is the output feature map after attention weighting, X c,h,w is the pixel value of the cth channel at position (h,w) in the input feature map.

[0070] (6) Model training:

[0071] The U-Net model and SE U-Net model are trained using the training set. Data enhancement techniques such as horizontal flipping, vertical flipping and random rotation are used during the training process to improve the generalization ability of the model in complex scenes. The classification cross entropy is used as the loss function. The loss function evaluates the difference between the model output probability distribution and the true mask distribution at the pixel level. Its calculation is shown in formula (6):

[0072]

[0073] Among them, L is the loss value, N is the number of samples, M is the total number of categories, and y ic is the indicator value of whether the i-th sample belongs to category c, is the probability that the model predicts that sample i belongs to category c.

[0074] The training parameters of the U-Net model and the SE U-Net model were set. Adam was used as the optimizer, BatchSize was set to 8, and Epochs was set to 100. The early stopping strategy was applied. When the loss value of the validation set no longer decreased after 20 consecutive rounds of training, the training was automatically stopped and the current optimal weights were saved to obtain the optimal models. The training processes of the U-Net model and the SE U-Net model were both accelerated by GPU. The U-Net model can usually complete training within 20 minutes, and the SE U-Net model can usually complete training within 10 minutes.

[0075] (7) Model testing:

[0076] The test set is input into the optimal model of U-Net and SE U-Net for high-precision object segmentation, and the predicted mask is output and the evaluation index is given. The evaluation index includes Accuracy, Precision, Recall, F1-score and mIoU. The calculation of the evaluation index is shown in formulas (7) to (11):

[0077]

[0078]

[0079]

[0080]

[0081]

[0082] Among them, TP is a true positive example, which means that the positive example is correctly classified as a positive example; FP is a false positive example, which means that the negative example is incorrectly classified as a positive example; FN is a false negative example, which means that the positive example is incorrectly classified as a negative example; TN is a true negative example, which means that the negative example is correctly classified as a negative example; TP i Indicates the number of pixels correctly predicted as category i; FP i FN represents the number of pixels that are actually not in category i but are mistakenly predicted as category i; i It represents the number of pixels that are actually in category i but are mistakenly predicted to be other categories; n is the total number of segmentation categories.

[0083] The evaluation indicators are all expressed in percentage form. The higher the value, the better the performance of the model. The accuracy of the U-Net model and the SE U-Net model is quantitatively evaluated by comprehensive analysis of the evaluation indicators. The segmentation results are visually compared based on the original image, the real mask and the predicted mask. The evaluation results of the U-Net model and the SE U-Net model constructed by the present invention are shown in Table 1. The visual comparison of the segmentation results of the SE U-Net model is shown in Table 1. Figure 2 shown.

[0084] Table 1 Evaluation results of U-Net and SE U-Net

[0085]

[0086] It can be seen from the above results that the SE U-Net model constructed by the present invention is significantly better than the U-Net model in all evaluation indicators in the concrete phase segmentation task, among which Accuracy, Precision, Recall and F1-score are respectively improved by 7.3%, 3%, 7.5% and 5.5% compared with U-Net, especially in the key indicator of mIoU, reaching 88.5%, which is a significant improvement of 9.7% compared with the U-Net model. The results show that compared with the U-Net model, the SE U-Net model has higher segmentation accuracy in the concrete phase segmentation task, can more accurately identify and distinguish the boundaries of different phases, and provide more accurate and reliable segmentation results.

Claims

1. A concrete phase segmentation method based on lightweight deep learning and SE attention mechanism, characterized in that: The following steps are involved: (1) Data acquisition: high-resolution images of the internal structure of concrete, i.e., original images, are obtained through X-ray tomography technology; (2) Data preprocessing: The original images are preprocessed and the concrete phases are annotated using a semi-automatic annotation method to form an initial data set; (3) Data segmentation: divide the initial data set into training set, validation set and test set; (4) Lightweight model architecture design: Based on the U-Net network architecture, the number of model parameters is reduced by reducing the number of channels in each convolutional layer, and a lightweight U-Net model is constructed; (5) SE attention mechanism integration: The SE attention mechanism is embedded in the convolutional layers of the encoder and decoder of the lightweight U-Net model to build a SE U-Net model. The model’s feature extraction and classification capabilities for different concrete phases are enhanced by using channel adaptive weighting. (6) Model training: Use the training set obtained in step (3) to train the SE U-Net model, use the validation set obtained in step (3) to validate it, perform hyperparameter tuning based on the validation results, and save the optimal model obtained after training; (7) Model testing: The test set obtained in step (3) is input into the optimal model for high-precision phase segmentation, the predicted mask is output, and evaluation indicators are given.

2. The method according to claim 1, characterized in that The physical phases of concrete are one or more of aggregate, cement paste or pores.

3. The method according to claim 1, characterized in that: In step (2), preprocessing includes rotation, cropping, and scaling to unify the image size; the semi-automatic labeling method is to perform preliminary labeling of concrete phases based on thresholds, supplemented by manual correction, and cover masks on the concrete phases respectively.

4. The method according to claim 1, characterized in that: In step (3), the training set, validation set, and test set are all independent of each other; the training set is used for model training, the validation set is used for model tuning, and the test set is used for model evaluation.

5. The method according to claim 1, characterized in that In step (3), the initial data set is randomly divided into a training set, a validation set, and a test set according to a preset ratio. A stratified sampling method is used in the division process to ensure that the data categories of the three subsets are evenly distributed. At the same time, a fixed random seed is set to ensure the repeatability of each random division result.

6. The method according to claim 1, characterized in that In step (4), the method of reducing the number of channels of each convolutional layer includes performing feature extraction and fusion with a smaller number of channels than the preset number of channels of the U-Net model in the convolution modules of each level of the encoder and the decoder, thereby reducing the overall model parameter amount and computing resource consumption.

7. The method according to claim 1, characterized in that In step (5), embedding the SE attention mechanism includes the following steps: 1) Squeeze: Each channel of the input feature map is spatially compressed through global average pooling to form the global features of each channel. This step can be expressed by formula (1): Among them, H represents the height of the input feature map, W represents the width of the input feature map, and X c,h,w represents the pixel value of the cth channel at position (h,w) in the input feature map, z c Represents the global features of the cth channel after average pooling; 2) Excitation: A two-layer fully connected network is used to transform the channel global features to learn and generate the adjustment weights of each channel. The channel global features are reduced in dimension by the first layer of fully connected network, and the ReLU activation function is applied to add nonlinear processing. The second layer of fully connected network expands the feature dimension back to the original number of channels and outputs the final channel weights through the Sigmoid activation function. This step can be expressed by formulas (2) to (4): ReLU(x)=max(0,x) (2) s c =σ(W2 ReLU(W1 z c )) (4) Among them, x represents the input value, ReLU is the activation function, σ represents the Sigmoid activation function, W1 represents the dimension reduction weight matrix, W2 represents the dimension increase weight matrix, s c represents the weight of the cth channel; 3) Scale: The weight s of each channel adjusts the channel strength of the original input feature map through element-wise multiplication. This step can be expressed by formula (5): AND c,h.w =s c ×X c,h,w (5) Among them, Y c,h.w is the output feature map after attention weighting, X c,h,w is the pixel value of the cth channel at position (h,w) in the input feature map.

8. The method according to claim 1, characterized in that: In step (6), the model training uses categorical cross entropy as the loss function, which evaluates the difference between the model output probability distribution and the true mask distribution at the pixel level; the Adam optimizer is used to update the parameters, and the early stopping strategy and the best model preservation strategy are set based on the loss value of the validation set.

9. The method according to claim 1, characterized in that: In step (6), data enhancement techniques are used on the training set data, including horizontal flipping, vertical flipping and random rotation, to improve the generalization ability of the model in complex scenarios.

10. The method according to claim 1, characterized in that In step (7), the evaluation indicators include Accuracy, Precision, Recall, F1-score and mIoU, which quantitatively evaluate the accuracy of the segmentation model and visually compare the segmentation results.

Citation Information

Patent Citations

  • Concrete crack segmentation method based on SCSEOCUnet

    CN111353396A

  • Attention mechanism improved U-net introduced concrete crack real-time detection method

    CN113284107A

  • Shale scanning electron microscope image segmentation method based on attention and U-Net

    CN116228797A

  • Skin melanin lesion image segmentation method based on U-Net enhanced multi-scale module and SE attention mechanism

    CN117576061A

  • Semantic segmentation method for quantifying bubble defects on surface of bare concrete based on YOLOv5 model

    CN118691822A

Cited By

  • Method and system for automatically identifying and segmenting zircon-quartz artifacts and quantitatively characterizing

    CN120564167A

  • SESUNet-based CT scanning image mineral intelligent segmentation method

    CN121505615A

  • A CT scan image mineral intelligent segmentation method based on SESUNet

    CN121505615B