Image semantic segmentation method based on feature complementation fusion model

By combining the feature complementarity fusion model of CNN and Transformer, the problem of poor performance in identifying large and small-scale defects in existing technologies has been solved, achieving efficient and accurate road defect detection, which is suitable for multi-scale defect identification in complex environments.

CN120913069APending Publication Date: 2025-11-07HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511016010.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing CNNs are not effective in identifying large-scale road defects (such as network cracks and large potholes), while Transformers are not effective in identifying small-scale defects (such as cracks and small potholes), which limits the accuracy and efficiency of road defect detection.

Method used

We propose an image semantic segmentation method based on a feature complement fusion model. Combining CNN and Transformer, we fuse multi-scale features through an improved MFFM fusion module and a hybrid encoder path. We also design a hybrid loss function with multi-scale weights and boundary awareness to improve the model's recognition ability.

Benefits of technology

It enables efficient identification of road defects at different scales, improves the accuracy and efficiency of detection, and can effectively identify defects such as cracks, network cracks and potholes in complex environments, while reducing the amount of computation and parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913069A_ABST
    Figure CN120913069A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of deep learning and image recognition, and discloses an image semantic segmentation method based on a feature complementation fusion model, the method is used for road disease recognition, the feature complementation fusion model comprises an encoder and a decoder, the encoder comprises a TransformerBlock module and a CNNBlock module, and the decoder comprises a CNNBlock module. The decoder comprises an upsampling Upsample Block module, a multi-scale feature fusion module, an output module and a context information recovery module. The method comprises the steps of constructing a road disease data set, preprocessing the road disease data set, labeling the data set, and constructing a feature complementary fusion model. Inputting the training set data into the feature complementation fusion model for training; and evaluating model performance by using test set data. According to the method, the problems of insufficient multi-scale feature adaptation and fuzzy boundary are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of deep learning and image recognition, and particularly relates to an image semantic segmentation method based on a feature complementary fusion model. BACKGROUND

[0002] Concrete pavement is a common choice for highway construction, characterized by long service life, good stability, and strong strength under proper maintenance. However, once damaged, concrete pavement can quickly deteriorate, with cracks, net cracks, and potholes being the most common forms of damage. If not promptly addressed, pavement cracks can lead to roadbed collapse, not only endangering the life of the road but also posing a significant risk to traffic safety. In addition, repairing the pavement incurs significant costs. Therefore, timely detection of pavement cracks at an early stage is crucial and has far-reaching implications for maintaining road health. This highlights the urgent need for early detection and intervention to mitigate these potential hazards. Initially, crack detection was heavily dependent on manual labor, which was fraught with challenges due to the subjectivity of individual judgment. Not only was it difficult to establish clear and consistent judgment criteria, but it also resulted in low efficiency, time consumption, and high labor costs.

[0003] However, in recent years, significant progress has been made in computer vision, leading to the gradual enhancement of automatic crack detection. Compared to traditional manual methods, automatic detection methods based on computer vision have many advantages, including improved efficiency, enhanced automation, and increased safety, making them the focus of extensive research. The mainstream architecture of existing road disease segmentation networks is mainly composed of CNN or Transformer, but CNNs show deficiencies in global modeling capabilities, and they perform poorly in identifying large-scale feature diseases such as net cracks and large potholes. While Transformers can capture long-range dependencies, they lack the ability to learn detailed boundaries and other information, and they perform poorly in identifying small-scale diseases such as cracks and small potholes. SUMMARY

[0004] To address the problems of poor recognition of large-scale diseases such as net cracks and large potholes by CNNs and poor recognition of small-scale diseases such as cracks and small potholes by Transformers in the prior art, the present application provides an image semantic segmentation method based on a feature complementary fusion model. This method uses an improved architecture that combines CNN and Transformer, and improves a MFFM fusion module to solve the problems of inadequate multi-scale feature adaptation and blurred boundaries.

[0005] To achieve the above-mentioned purposes, the present application is implemented through the following technical solutions:

[0006] The application is an image semantic segmentation method based on a feature complementary fusion model, characterized in that: the method is used for road disease identification, and the image semantic segmentation method is realized based on a feature complementary fusion model, and the image semantic segmentation method comprises the following steps:

[0007] Step 1, data acquisition and preprocessing: using a collection device to shoot ground images at different time periods and in different weather, obtaining an original image data set, and performing data preprocessing on the original image, specifically including: adjusting the size, random rotation, random cropping and normalization processing;

[0008] Step 2, labeling the original image data obtained in step 1: using Labelme software to label the road disease area on the original image data and converting the labeling result into a label image file, obtaining a road disease data set, and the road disease data set is divided into a training set and a test set;

[0009] Step 3, constructing a feature complementary fusion model (TCFUNet) to realize road disease identification: using Python to realize a semantic segmentation network model TCFUNet, which is different from existing models for road disease identification, the TCFUNet model researched and realized by the application is based on the idea of feature complementary fusion, and the purpose is to effectively identify different scale road diseases such as cracks, network cracks and pits;

[0010] Step 4, the training set in step 2 is used for training the feature complementary fusion model (TCFUNet), the error between the label value and the true value is calculated, the feature complementary fusion model (TCFUNet) weight parameter is updated by using the back propagation mechanism, so as to improve the performance of the model; the test set in step 2 is used for evaluating the performance of the feature complementary fusion model (TCFUNet) and verifying the generalization ability of the feature complementary fusion model (TCFUNet), so as to ensure the effectiveness and reliability of the model in actual application, and evaluation indexes including Dice score, mIou and the like are used to evaluate the performance of the model.

[0011] Further improvement of the application is that in step 3, the feature complementary fusion model (TCFUNet) comprises an encoder and a decoder, and the encoder and the decoder are connected through a feature fusion module;

[0012] The encoder comprises a Transformer encoder path (TEncPath) and a hybrid encoder path (HEncPath), and comprises 4 Transformer_Block modules and 5 CNN_Block modules in total;

[0013] The decoder comprises an up-sampling Upsample_Block module, three multi-scale feature fusion modules (MFFM), one output module OB, one side output module SOB, and a context information recovery module (CIR).

[0014] The further improvement of the present application is that the Transformer encoder path (TEncPath) comprises four Transformer_Block modules, each of which comprises an FFN and an MHSA, and directly uses a Transformer network for global relationship modeling.

[0015] The further improvement of the present application is that the hybrid encoder path (HEncPath) is a hybrid encoder of global long-range dependency and local representation, the hybrid encoder path (HEncPath) comprises a CNN module, the CNN module comprises five stage convolution modules and four down-sampling modules, each stage of the hybrid encoder path (HEncPath) is input from the previous convolution stage and the corresponding encoder path (TEncPath), the first stage comprises a convolution with a kernel size of 3, a BatchNorm, a maximum pooling and a convolution with a kernel size of 3; the new spatial resolution of the feature map after the first stage is half of the original resolution, this operation has two advantages: first, the model has fewer parameters, and second, it expands the overall receptive field, thereby helping to capture contextual features; fewer normalization layers and activation functions are used. The remaining four stages use two convolutions with a kernel size of 3 and BatchNorm, the down-sampling module uses a convolution with a kernel size of 1, BatchNorm and Hswish activation function.

[0016] The further improvement of the present application is that the up-sampling Upsample_Block module of the decoder comprises a point-wise convolution layer with a kernel size of 1, an H-Swish layer and a bilinear up-sampling layer for adjusting the spatial size;

[0017] The further improvement of the present application is that the multi-scale feature fusion module (MFFM) of the decoder comprises two CIR_Block modules and one HEncPath_Stage module, the CIR_Block module comprises a CIR_Block1 out5 , a CIR_Block2 out4 and an HEncPath_Stage3 outAlong the channel dimension splicing, then input to the convolutional block of depth-wise Conv, point-wise Conv and H-Swish activation function, using depth-wise Conv and point-wise Conv to reduce the parameter quantity and calculation quantity, realize the lightweight structure of the module, introduce the above lightweight structure into global average pooling and maximum average pooling, multilayer perception and softmax activation function, the improved MFFM effectively fuses the multi-scale information of different receptive fields and improves the model feature representation ability.

[0018] The further improvement of the application is that the output module OB of the decoder includes a transpose convolution layer with kernel_size=2 and a normal convolution layer with kernel_size=1, and the side output module SOB includes two convolution layers with kernel_size=3 and a side input block of LN_VT layer.

[0019] The further improvement of the application is that the context information recovery module (CIR) of the decoder restores the low-resolution feature map to the original resolution through upsampling and fuses it with the corresponding high-resolution feature map, helping the model better understand the local details and edge information of the image. First, a 3x3 convolution operation is used to reduce the number of input channels to half, then three branches are constructed in parallel, each branch contains different convolution layers; among them, the first branch, the second branch and the third branch use 5x5 deformable convolution layer, 7x7 deformable convolution layer and 3x3 deformable convolution layer respectively, and then multiply, fully connected layer, Swish activation function, fully connected layer, GELU activation function are used to process the input feature map, which is used to extract larger range of context features; the third branch uses a 3x3 deformable convolution layer to process the input feature map, which is used to extract local detail features; finally, the output feature maps of the three branches are merged after multiplication and fully connected layer to obtain feature maps of different levels of receptive field.

[0020] The further improvement of the application is that in step 1, the error between the label value and the true value is calculated, specifically: a hybrid loss function combining multi-scale weight and boundary perception is designed:

[0021] L MSWHL =αL Dice +βL Focal +γL Boundary

[0022]

[0023] Where: L Dice is the Dice loss, L Focal is the Focal loss variant, and L BoundaryFor the boundary perception loss, the scale weight coefficient A i For the pixel value of each disease connected region, alpha, beta, and gamma are the loss function weight coefficients, y i is the true label of the i-th pixel, p i is the probability that the i-th pixel predicted by the feature complementary fusion model (TCFUNet) is diseased, to enhance the sensitivity to small-scale diseases, N is the total number of pixels in the image, tau is a hyperparameter, B is a set of disease boundary pixels extracted by a Sobel operator, and |B| is the total number of boundary pixels.

[0024] The beneficial effects of the present application are:

[0025] The present application adopts an improved architecture combining CNN and Transformer, which not only has strong global perception ability, but also can capture low-level fine-grained information. This model can not only be used for road disease identification tasks, but also can be extended to other identification tasks.

[0026] The improved MFFM fusion module proposed in the present application allows information of different scales to flow in the network and effectively fuses multi-scale information of different receptive fields.

[0027] The present application also designs a hybrid loss function combining multi-scale weights and boundary perception to solve the problems of insufficient multi-scale feature adaptation and blurred boundaries.

[0028] The present application introduces a brand new original image dataset, named RoadDamage dataset, which aims to provide more diverse and challenging data to support research in the field of road diseases. This dataset contains 4680 high-quality, finely labeled images, covering different traffic scenarios and environmental conditions. The two notable features of the RoadDamage dataset are: first, the expressiveness in complex environments, including environmental interference such as leaves and tree shade, as well as adverse weather conditions such as night and rain; second, covering different types of road diseases such as cracks, net cracks, and potholes, solving the sample imbalance problem caused by the fact that the crack dataset is more common in public datasets, while the pothole dataset is extremely rare.

[0029] In summary, the present application designs a novel, reasonable, accurate and efficient model, which performs well in road disease identification tasks compared to existing models applied to identification tasks. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is the flowchart of the present application.

[0031] Figure 2 is the structure diagram of the feature complementary fusion model (TCFUNet) of the present application.

[0032] Figure 3 Figure 1 is a structural diagram of a Transformer_Block module of the present application.

[0033] Figure 4 Figure 2 is a structural diagram of a MFFM module of the present application.

[0034] Figure 5 Figure 3 is a structural diagram of a CIR module of the present application.

[0035] Figure 6 Figure 4 is a graph of a comparison experiment result of a disclosed data set of the present application.

[0036] Figure 7 Figure 5 is a graph of a comparison experiment result of a self-built data set of the present application. DETAILED DESCRIPTION

[0037] Embodiments of the present application will be described below with reference to the drawings. Many practical details will be described in the following description for the purpose of making the present application clear. However, it should be understood that these practical details are not intended to limit the present application. That is, in some embodiments of the present application, these practical details are not necessary. In addition, for the purpose of simplifying the drawings, some conventional structures and components will be shown in the drawings in a simplified schematic manner.

[0038] The present application is an image semantic segmentation method based on a feature complementary fusion model, which is used for road disease identification. The image semantic segmentation method is realized by a feature complementary fusion model.

[0039] As shown in Figure 2 , the feature complementary fusion model (TCFUNet) includes an encoder and a decoder, and the encoder and the decoder are connected through a feature fusion module. The encoder includes a Transformer encoder path (TEncPath) and a hybrid encoder path (HEncPath), which together include 4 Transformer_Block modules and 5 CNN_Block modules. The decoder includes an up-sampling Upsample_Block module, 3 multi-scale feature fusion modules (MFFM), 1 output module OB, 1 side output module SOB, and a context information recovery module (CIR). The model can capture high-level global semantic features and low-level fine-grained information at the same time. In addition, a series of efficient and lightweight designs are introduced, which effectively reduce the parameter quantity and computational quantity of the model on the premise of ensuring the accuracy without loss.

[0040] AsFigure 3 As shown, the Transformer encoder path (TEncPath) includes 4 Transformer_Block modules, each of which includes an FFN and an MHSA, for expanding the receptive field and gradually extracting features of the image; the TEncPath directly uses a Transformer network for global relationship modeling to locate significant objects, especially for importance weighting of large-scale features, improving the network's feature extraction ability and overall performance for large-scale diseases such as net cracks and pits.

[0041] Since the Transformer completely relies on self-attention to extract global context features, it lacks the ability to learn local features such as detail boundaries, and detail boundary features are often the most important features in disease cracks, therefore, the application introduces a hybrid encoder path (HEncPath) to enhance the local learning ability of the encoder. In order to fuse global and local features, each stage of the HEncPath is input from the previous convolution stage and the corresponding TEncPath stage, and the HEncPath introduces locality into the feature representation under the guidance of the global context of the TEncPath, so the hybrid encoder path (HEncPath) is a hybrid encoder of global long-range dependency and local representation, the hybrid encoder path (HEncPath) includes a CNN module, the CNN module includes 5 stage convolution modules and 4 down-sampling modules, each stage of the hybrid encoder path (HEncPath) is input from the previous convolution stage and the corresponding encoder path (TEncPath), the first stage includes a convolution with kernel_size=3, a BatchNorm, a max-pooling and a convolution with kernel_size=3; the new spatial resolution of the feature map after the first stage is half of the original resolution, this operation has two advantages: first, the model has fewer parameters, second, it expands the overall receptive field, thus helping to capture context features; fewer normalization layers and activation functions are used. The remaining 4 stages use two convolutions with kernel_size=3 and BatchNorm, the down-sampling module uses a convolution with kernel_size=1, BatchNorm and H-swish activation function.

[0042] As Figure 4As shown, the context information recovery module (CIR) of the decoder restores the low-resolution feature map to the original resolution through up-sampling and fuses it with the corresponding high-resolution feature map, helping the model better understand the local details and edge information of the image. The up-sampling Upsample_Block module of the decoder includes a point-wise convolution layer with kernel_size = 1, an H-Swish layer, and a bilinear up-sampling layer for adjusting the spatial size;

[0043] The CIR module can consider feature information of different scales at the same time through parallel convolution, and enrich the representation ability of the algorithm by merging different feature maps, thereby recovering the context information in semantic segmentation. In the CIR module, first, a 3x3 convolution operation is used to reduce the number of input channels to half, and then three branches are constructed in parallel, each of which contains different convolution layers; among them, the first branch and the second branch use a 5x5 deformable convolution layer and a 7x7 deformable convolution layer, respectively, and then process the input feature map through multiplication, a fully connected layer, a Swish activation function, a fully connected layer, and a GELU activation function, for extracting larger range of context features; the third branch processes the input feature map through a 3x3 deformable convolution layer, for extracting local detail features; finally, the output feature maps of the three branches are merged through multiplication and a fully connected layer to obtain feature maps of different levels of receptive field.

[0044] Specifically, the input low-level feature map x and the high-level feature map g connected by the encoder are respectively weighted by importance to obtain feature maps x_se and g_se, then an up-sampling layer is used to up-sample the feature map x_se to obtain a feature map d to match the size of the feature map g_se, and then the obtained feature map d is spliced with the feature map g_se to realize feature fusion between different levels.

[0045] As shown in Figure 4 , the multi-scale feature fusion module (MFFM) of the decoder is designed to allow different scale information to flow in the network and effectively fuse the extracted features. Unlike the pyramid pooling (ASPP), the input of MFFM comes from different modules, including two CIR_Block, and a HEncPath_Stage module, CIR_Block out5 , CIR_Block out4 and HEncPath_Stage3 outThe image is spliced along the channel dimension, and then input into a convolution block of a depth-wise convolution (Depth-wise Conv), a point-wise convolution (Point-wise Conv) and an H-Swish activation function, the depth-wise convolution and the point-wise convolution are used to reduce the parameter quantity and the calculation quantity, a lightweight structure of the module is realized, the lightweight structure is introduced into a parallel attention module PAM, the attention module PAM includes global average pooling and maximum average pooling, a multi-layer perception and a softmax activation function, the improved MFFM effectively fuses multi-scale information of different receptive fields and improves the feature representation capability of the model.

[0046] An output layer of the feature complementary fusion model (TCFUNet) is used to generate a final recognition result, and an output part includes two branches, a main branch and a side branch, the main branch is an output module OB, the output module OB includes a transpose convolution layer with a kernel size of 2 and a normal convolution layer with a kernel size of 1, and the side branch is a side output module SOB, including two convolution layers with a kernel size of 3 and a side input block of an LN_VT layer.

[0047] As shown in Figure 1 The image semantic segmentation method includes the following steps:

[0048] Step 1, data acquisition and preprocessing: using a collection device to take pictures of ground images at different time periods and in different weather, obtaining an original image data set, and performing data preprocessing on the original image, including: adjusting the size, random rotation, random cropping and normalization processing;

[0049] Step 2, labeling the original image data obtained in step 1: using LabelMe software to label the road disease area on the original image data and converting the labeling result into a label image file using an existing labeling tool LabelMe, obtaining a road disease data set, the road disease data set is divided into a training set and a test set; the step of converting the label image using the existing labeling tool LabelMe is: loading the image, drawing the area, naming the label, and saving the result.

[0050] Step 3, constructing a feature complementary fusion model (TCFUNet) to realize road disease recognition: using Python to realize a semantic segmentation network model TCFUNet, which is different from existing models for road disease recognition, the TCFUNet model researched and realized by the present application is based on the idea of feature complementary fusion, and the purpose is to effectively identify different scales of road diseases such as cracks, net cracks and potholes.

[0051] In this step, the road disease recognition is realized by the constructed feature complementary fusion model, the input is the road disease image, and the output is the pixel-level disease recognition result. The specific steps are as follows:

[0052] Step 1. Multi-scale feature extraction: global context features of the original image data are extracted through 4 Transformer_Block modules (including MHSA and FFN) of the Transformer encoder path (TEncPath), and local detail features are extracted through 5 CNN_Block modules (including 3x3 convolution, BatchNorm, and down-sampling). At the same time, the feature maps of each stage are spliced and fused with the corresponding layers of the hybrid encoder path (HEncPath), and the semantic information containing the global and local feature fusion of the disease is output.

[0053] Step 2: Feature fusion: global context features are extracted through deformable convolution of three branches, and after branch fusion, up-sampling is performed to the original resolution, and shallow features are spliced to restore disease edge details. The feature maps from different levels of the encoder and the decoder (such as out3, out4-CIR, and out5-CIR) are fused through the feature fusion module MFFM, and the calculation amount is reduced through depth separable convolution; global average pooling (GAP) and maximum pooling (GMP) are used to generate multi-scale attention weights; the sensitivity of the model to different size diseases such as cracks and pits is improved; and the enhanced multi-scale feature map is output after weighted fusion.

[0054] Step 3: Disease classification and output: the OB module restores the feature map to the input image size through transposed convolution (kernel_size=2) and ordinary convolution (kernel_size=1); the class probability of each pixel is generated using Softmax, and the output image is the recognized disease image. The SOB module is an auxiliary supervision branch that predicts the disease boundary through the intermediate layer feature, enhancing the sensitivity of the model to the edge.

[0055] Step 4, the training set in step 2 is used for training the feature complementary fusion model (TCFUNet), the error between the label value and the true value is calculated, and the feature complementary fusion model (TCFUNet) weight parameters are updated using the back propagation mechanism, so as to improve the performance of the model; the test set in step 2 is used to evaluate the performance of the feature complementary fusion model (TCFUNet) and verify the generalization ability of the feature complementary fusion model (TCFUNet), so as to ensure the effectiveness and reliability of the model in actual application. The evaluation indicators include Dice score, mIou, etc. to evaluate the performance of the model.

[0056] To achieve the present application, a dataset class class ConcreteDataset is defined for loading the dataset required for model training and preprocessing images and labels; specifically including: converting to PIL image format, adjusting size, normalizing. Secondly, create a data loader, use the data loader to pass the training data to the model. The training parameter settings of this experiment are as follows:

[0057] num_classes=2 binary classification

[0058] batch_size=8 batch size of data loader during training

[0059] lr=0.0001 initial learning rate of the model

[0060] weight_decay=0.0001 weight decay coefficient of the optimizer

[0061] α=0.5, β=0.3, γ=0.2 weight coefficients of the initial loss function

[0062] This embodiment is based on the experimental verification of the feature complementary fusion model on the Crack500, Pthole, and Roaddamage datasets, implemented using the image processing library in Python, including: adjusting the size, random cropping, and normalization. Then input these images into the TCFUnet model for training and validation to obtain the recognition result. The Crack500 dataset in this embodiment contains a total of 2752 images, all of which are crack images, of which 2252 images are divided into the training set and 500 images are divided into the test set; the Pthole dataset in this embodiment contains a total of 700 pit and groove images, of which 600 images are divided into the training set and 100 images are divided into the test set; the Roaddamage dataset in this embodiment contains a total of 4680 images, including crack, net crack, and pit and groove road diseases, of which 3744 images are divided into the training set and 936 images are divided into the test set.

[0063] Loss function dynamic weight calculation steps:

[0064] α=0.5, β=0.3, γ=0.2 are the weight coefficients of the initial loss function, and in each training cycle (epoch), the absolute change of each loss function value in the current cycle and the previous cycle is calculated:

[0065] ΔL Dice =|L Dice_x -L Dice_x-1 |

[0066] ΔL Focal =|L Focal_x -L Focal_x-1|

[0067] Delta L Boundary = |L Boundary_x - L Boundary_x-1 |

[0068] Based on the change rate of each loss function, the temporary weight is calculated according to the following formula:

[0069] alpha' = alpha0 x (1 + Delta L Dice / (Delta L Dice + Delta L Focal + Delta L Boundary + epsilon))

[0070] beta' = beta0 x (1 + Delta L Focal / (Delta L Dice + Delta L Focal + Delta L Boundary + epsilon))

[0071] gamma' = gamma0 x (1 + Delta L Boundary / (Delta L Dice + Delta L Focal + Delta L Boundary + epsilon))

[0072] Wherein, epsilon is a very small positive number (1e-7), used to avoid the denominator being 0.

[0073] Weight normalization:

[0074] alpha = alpha' / (alpha' + beta' + gamma')

[0075] beta = beta' / (alpha' + beta' + gamma')

[0076] gamma = gamma' / (alpha' + beta' + gamma')

[0077] The total loss is: L MSWHL = alpha L Dice + beta L Focal + gamma L Boundary .

[0078] In order to quantitatively evaluate the improved model recognition performance of the application, comparative experiments are carried out with the existing model respectively, and the following evaluation indexes are adopted: average intersection over union mIou, Dice score, parameter quantity, calculation amount. The comparison results are shown in Table 1 and Table 2 below.

[0079] Table 1

[0080] Model mIou Dice param Flops CrackSeU 81.21% 78.83% 7.7M 11.22G CT-crackseg 79.37% 77.98% 22.88M 22.88G TCFUNet 85.37% 84.03% 18M 19.83G ABiUnet 85.25% 82.96% 36M 34.31G

[0081] Table 2

[0082] Model mIou Dice param Flops CrackSeU 68.34% 68.15% 7.7M 14.64G CT-crackseg 69.38% 68.87% 22.88M 28.55G TCFUNet 73.74% 72.34% 18M 25.08G ABiUnet 72.17% 70.89% 36M 38.19G

[0083] Table 1 is the comparative experiment result of the public data set, Table 2 is the comparative experiment result of the Roaddamage data set, according to Table 1 and Table 2, and Figure 5 and Figure 6 The mIou and dice scores of the TCFUNet model on the crack data set are obviously better than those of the CNN CrackSeU model, the mIou and dice scores on the pit data set are obviously better than those of the CT-crackseg model of the Transformer, and compared with the combination of the CNN and the Transformer, the TCFUNet model effectively reduces the calculation amount while ensuring the accuracy without loss, and the above results show that the present application is superior to the comparative model.

[0084] In order to quantitatively evaluate the improved model identification performance of the present application, an ablation experiment is performed, and the following evaluation indexes are used: average intersection over union mIou, Dice score. The comparison results are shown in Table 3 below.

[0085] Table 3

[0086]

[0087] According to the ablation experiment results of the Roaddamage data set in Table 3, the mIou of the model with the CIR module is improved by 1.68% compared with the Baseline model, the Dice is improved by 0.83%, the mIou after adding the MFFM module is improved by 1.16%, the Dice is improved by 0.79%, the improved MSWHL loss function is improved by 0.62% compared with the original Bice loss function, and the Dice is improved by 0.89%.

[0088] Table 4 is the experimental result of different weight strategies on the Roaddamage data set.

[0089] Table 4

[0090]

[0091]

[0092] According to Table 4, the mIou and Dice of the dynamic weight are improved compared with the fixed initial weight, which shows that the present application can effectively improve the model performance.

[0093] In step 4, the binary cross-entropy loss function is used to measure the gap between the probability distribution of the model output and the true value, and specifically, a hybrid loss function combining multi-scale weight and boundary perception is designed, and the formula is as follows:

[0094] L MSWHL =αL Dice +βL Focal+ γL Boundary

[0095]

[0096] where L Dice is the improved Dice loss, introducing a scale weight coefficient A i is the pixel value of each disease connected region, y i is the real label of the i-th pixel, p i is the probability of the i-th pixel predicted by the model to be diseased, enhancing the sensitivity to small-scale diseases.

[0097]

[0098] where L Focal is the Focal loss variant, N is the total number of pixels in the image, y i is the real label of the i-th pixel, p i is the probability of the i-th pixel predicted by the model to be diseased, τ is a hyperparameter that controls the weight of difficult samples, solving the class imbalance problem.

[0099]

[0100] where L Boundary is the boundary-aware loss, B is the set of disease boundary pixels extracted by the Sobel operator, |B| is the total number of boundary pixels, y i is the real label of the i-th pixel, p i is the probability of the i-th pixel predicted by the model to be diseased.

[0101] The hybrid loss function combines multi-scale weight and boundary awareness to solve the problem of multi-scale feature adaptation and boundary ambiguity. Through experiments on different parameters, the weight coefficients of the loss function are determined.

[0102] The MFFM fusion module proposed in the application allows information of different scales to flow in the network, effectively fuses multi-scale information of different receptive fields, and improves the performance of the model in identifying different types of road diseases. At the same time, the model adopts a lightweight design, increases the performance of the model under the condition of reducing the parameter quantity and the calculation complexity. The model can not only be used for road disease identification tasks, but also can be popularized to tasks in other resource-limited environments, and is not limited to road disease segmentation networks. The CIR module proposed in the application can adaptively fit the complex morphology of road diseases, and is complementary to the multi-scale fusion of the MFFM module, further improving the joint modeling capability of the model for diseases of different scales. A multi-scale weighted hybrid loss (MSWHL) designed by the application combines multi-scale weight and boundary perception, solves the problems of insufficient multi-scale feature adaptation and fuzzy boundary.

[0103] The above merely describes the embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of claims of the present application.

Claims

1. An image semantic segmentation method based on a feature complementary fusion model, characterized in that: The method is used for road disease identification, and the image semantic segmentation method is realized by a feature complementary fusion model, and comprises the following steps: Step 1, data acquisition and preprocessing: using a collection device to shoot ground images at different time periods and in different weather, to obtain an original image data set, and to perform data preprocessing on the original image; Step 2, labeling the original image data obtained in step 1: labeling the road disease area on the original image data and converting the labeling result into a label image file, to obtain a road disease data set, which is divided into a training set and a test set; Step 3, constructing a feature complementary fusion model (TCFUNet) to realize road disease identification: identifying the road diseases of cracks, net cracks and pits based on the constructed feature complementary fusion model (TCFUNet); Step 4, the training set in step 2 is used for training the feature complementary fusion model (TCFUNet), to calculate the error between the label value and the true value, update the weight parameters of the feature complementary fusion model (TCFUNet), and the test set in step 2 is used for evaluating the performance of the feature complementary fusion model (TCFUNet) and verifying the generalization ability of the feature complementary fusion model (TCFUNet). 2.The image semantic segmentation method based on the feature-complementary fusion model according to claim 1, characterized in that: In step 3, the feature complementary fusion model (TCFUNet) comprises an encoder and a decoder, and the encoder and the decoder are connected through a feature fusion module; The encoder comprises a Transformer encoder path (TEncPath) and a hybrid encoder path (HEncPath), and comprises 4 Transformer_Block modules and 5 CNN_Block modules in total; The decoder comprises an up-sampling Upsample_Block module, 3 multi-scale feature fusion modules (MFFM), 1 side output module SOB, 1 output module OB, and further comprises a context information recovery module (CIR). 3.The image semantic segmentation method based on the feature-complementary fusion model according to claim 2, characterized in that: The Transformer encoder path (TEncPath) comprises 4 Transformer_Block modules, each Transformer_Block module comprises an FFN and an MHSA, and a global relationship is modeled directly using a Transformer network. 4.The image semantic segmentation method based on the feature-complementary fusion model according to claim 2, characterized in that: The hybrid encoder path (HEncPath) is a hybrid encoder of global long-range dependency and local representation, the hybrid encoder path (HEncPath) comprises a CNN module, the CNN module comprises five stage convolution modules and four down-sampling modules, each stage of the hybrid encoder path (HEncPath) inputs respectively come from the previous convolution stage and the corresponding encoder path (TEncPath), the first stage comprises a convolution with kernel_size=3, a BatchNorm, a max pooling and a convolution with kernel_size=3, the remaining four stages adopt two convolutions with kernel_size=3 and BatchNorm, and the down-sampling module adopts a convolution with kernel_size=1, a BatchNorm and an H-swish activation function.

5. The image semantic segmentation method based on the feature complementary fusion model according to claim 2, characterized in that: The up-sampling Upsample_Block module of the decoder comprises a point-wise convolution layer with kernel_size=1, an H-Swish layer and a bilinear up-sampling layer for adjusting the spatial size. 6.The image semantic segmentation method based on the feature-complementary fusion model according to claim 2, characterized in that: The multi-scale feature fusion module (MFFM) of the decoder includes two CIR_Block and one HEncPath_Stage module, CIR_Block out5 , CIR_Block out4 and HEncPath_Stage3 out are spliced along the channel dimension and then input into a convolution block of depth-wise convolution, point-wise convolution and H-Swish activation function, the depth-wise convolution and point-wise convolution are used to reduce the parameter quantity and calculation quantity, realize the lightweight structure of the module, and the above lightweight structure is introduced into global average pooling and maximum average pooling, multi-layer perception and softmax activation function.

7. The image semantic segmentation method based on the feature-complementary fusion model according to claim 2, characterized in that: The output module OB of the decoder comprises a transpose convolution layer with kernel_size=2 and a normal convolution layer with kernel_size=1, and the side output module SOB comprises two convolution layers with kernel_size=3 and a side input block of an LN_VT layer. 8.The method of claim 2, wherein the method further comprises: The context information recovery module (CIR) of the decoder restores the low-resolution feature map to the original resolution through up-sampling and fuses it with the corresponding high-resolution feature map, first reduces the channel number of the input to half through a 3*3 convolution operation, then constructs three branches in a parallel manner, wherein the first branch and the second branch use a 5*5 deformable convolution layer and a 7*7 deformable convolution layer respectively, and then process the input feature map through multiplication, a fully connected layer, a Swish activation function, a fully connected layer and a GELU activation function, the third branch processes the input feature map through a 3*3 deformable convolution layer to extract local detail features, finally, the output feature maps of the three branches are merged through multiplication and a fully connected layer to obtain feature maps of different levels of receptive field. 9.The method of claim 1, wherein: In step 1, the error between the label value and the true value is calculated, specifically: a hybrid loss function combining multi-scale weight and boundary perception is designed: L MSWHL = aL Dice + bL Focal + gL Boundary wherein: L Dice is the Dice loss, L Focal is the Focal loss variant, L Boundary is the boundary-aware loss, scale weight coefficient A i is the pixel value of each disease connected region, a, b, g are the loss function weight coefficients, y i is the i-th pixel real label, p i is the i-th pixel predicted by the feature complementary fusion model (TCFUNet) as a disease probability, enhances the sensitivity to small-scale diseases, N is the total number of pixels in the image, t is a hyperparameter, B is a set of disease boundary pixels extracted by a Sobel operator, and |B| is the total number of boundary pixels.