Double-branch embedded connection concrete crack image segmentation method and device

Through the DEC-Net model with dual-branch embedding and connecting, combined with CNN and Global Branch modules, the problem of difficulty in accurately identifying elongated concrete cracks in the prior art is solved, and high accuracy segmentation of crack images is achieved, taking into account local details and overall structure.

CN120259334APending Publication Date: 2025-07-04SHIJIAZHUANG TIEDAO UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411690278.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and deal with elongated concrete cracks, resulting in insufficient accuracy in crack segmentation and difficulty in taking into account local details and overall structure.

Method used

The DEC-Net model with dual-branch embedded connection is adopted, combined with the CNN module and the Global Branch module, and the global context information is extracted through the densely connected CSwin Transformer Block module, and the feature map is gradually upsampled and fusion through the decoder module to obtain feature maps of different scales, and finally the accurate segmentation of the crack image is achieved.

Benefits of technology

The segmentation accuracy of concrete crack images is improved, and the overall direction and local details of the crack can be better captured, so as to achieve effective identification and segmentation of slender cracks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259334A_ABST
    Figure CN120259334A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of image recognition, and provides a double-branch embedded connection concrete crack image segmentation method and device, and the method comprises the steps: obtaining a concrete crack image, inputting the concrete crack image into a preset DEC-Net network structure, and obtaining a segmentation result of the concrete crack image; wherein the DEC-Net network structure comprises a CNN (Convolutional Neural Network) module, a Global Branch module and a decoder module; the CNN module is used for performing 3 * 3 convolution on the concrete crack image to obtain a first-layer feature map; the Global Branch module is used for extracting features with global context information from the concrete crack image through the densely connected CSwin Transform Block module, and carrying out step-by-step up-sampling on the features to obtain multi-layer feature maps with different scales; and the decoder module is used for connecting the first-layer feature map and the multi-layer feature map, and performing two-layer convolution on a connection result to obtain a segmentation result of the concrete crack image. According to the invention, the accuracy of concrete crack image segmentation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a method and device for segmenting concrete crack images with a double-branch embedded connection. Background Art

[0002] In current civil engineering and water conservancy and hydropower projects, concrete structures account for a large proportion and are in a dominant position. Cracks are a common appearance quality defect in concrete structures. Under the combined action of internal and external factors, cracks often occur in concrete structures, which will have an adverse impact on the safety, applicability, and durability of the structures. Therefore, it is crucial for the safety of concrete structures to detect, evaluate, and take corresponding maintenance measures for concrete cracks.

[0003] However, in practical applications, there is a problem of how to accurately identify and process slender concrete cracks. Such cracks often have complex shapes and distributions, which pose a great challenge to image segmentation algorithms. Existing methods have certain limitations when dealing with concrete crack images, restricting the accuracy of crack segmentation and making it difficult to take into account both local details and overall structures simultaneously. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method and device for segmenting concrete crack images with a double-branch embedded connection to solve the problem in the prior art of difficultly accurately identifying and processing slender concrete cracks and improve the accuracy of concrete crack image segmentation.

[0005] The first aspect of the embodiments of the present invention provides a method for segmenting concrete crack images with a double-branch embedded connection, including:

[0006] Obtain a concrete crack image, and input the concrete crack image into a preset DEC-Net network structure to obtain a segmentation result of the concrete crack image;

[0007] Wherein, the DEC-Net network structure includes a CNN module, a Global Branch module, and a decoder module;

[0008] The CNN module is used to perform 3×3 convolution on the concrete crack image to obtain a first-layer feature map;

[0009] The Global Branch module is used to extract features with global context information from the concrete crack image through densely connected CSwin Transformer Block modules, and perform progressive upsampling on the features to obtain multi-layer feature maps of different scales;

[0010] The decoder module is used to connect the first-layer feature map and the multi-layer feature maps, and obtain the segmentation result of the concrete crack image by performing two-layer convolution on the connection result.

[0011] Optionally, the multi-layer feature maps of different scales include: a second-layer feature map, a third-layer feature map, a fourth-layer feature map, and a fifth-layer feature map;

[0012] The Global Branch module is used to:

[0013] Divide the concrete crack image into patches through the Patch embedding module;

[0014] Add position embeddings to the divided patches through the position embedding module and add them together;

[0015] Extract features with global context information from the concrete crack image through the densely connected CSwin Transformer Block module;

[0016] Gradually upsample the features to obtain the fifth-layer feature map, the fourth-layer feature map, the third-layer feature map, and the second-layer feature map in sequence.

[0017] Optionally, the decoder module is used to:

[0018] Connect the first-layer feature map and the second-layer feature map through the Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain the output feature map of the second layer;

[0019] After performing max pooling operation on the output feature map of the second layer, connect the output feature map of the second layer and the third-layer feature map through the Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain the output feature map of the third layer;

[0020] After performing max pooling operation on the output feature map of the third layer, connect the output feature map of the third layer and the fourth-layer feature map through the Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain the output feature map of the fourth layer;

[0021] After performing max pooling operation on the output feature map of the fourth layer, connect the output feature map of the fourth layer and the fifth-layer feature map through the Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain the output feature map of the fifth layer;

[0022] Connect the first-layer feature map and the output feature maps of the second, third, fourth, and fifth layers to obtain the connection result.

[0023] Optionally, before connecting the first-layer feature map with the output feature maps of the second layer, the third layer, the fourth layer, and the fifth layer, the decoder module is further configured to:

[0024] Upsample the output feature maps of the third layer, the fourth layer, and the fifth layer;

[0025] Among them, the output feature map of the third layer is subjected to one upsampling process, the output feature map of the fourth layer is subjected to two upsampling processes, and the output feature map of the fifth layer is subjected to three upsampling processes.

[0026] Optionally, the Embedding Connection module EC is configured to:

[0027] For two arbitrary input feature maps X1 and X2, perform average pooling operation, first convolution operation, and second convolution operation respectively to obtain X3 and X4;

[0028] Multiply X1 and X3 to get X5, and multiply X2 and X4 to get X6;

[0029] Add X1 and X2 to get X7;

[0030] Add X5, X6, and X7 to obtain the feature map X as the output.

[0031] Optionally, the first convolution operation includes a 1×1 convolution operation, a BN operation, and a ReLU operation, and the second convolution operation includes a 1×1 convolution operation, a BN operation, and a Sigmoid operation.

[0032] Optionally, the densely connected CSwin Transformer Block module is: densely and residually connect four CSwin Transformer Block modules.

[0033] The second aspect of the embodiments of the present invention provides a dual-branch embedding connection concrete crack image segmentation device, including:

[0034] An acquisition module, configured to acquire a concrete crack image;

[0035] A segmentation module, configured to input the concrete crack image into a preset DEC-Net network structure to obtain a segmentation result of the concrete crack image;

[0036] Among them, the DEC-Net network structure includes a CNN module, a Global Branch module, and a decoder module;

[0037] The CNN module is configured to perform 3×3 convolution on the concrete crack image to obtain a first-layer feature map;

[0038] The Global Branch module is used to extract features with global context information from the concrete crack image through densely connected CSwin Transformer Block modules, and gradually upsample the features to obtain multi-layer feature maps of different scales;

[0039] The decoder module is used to connect the first-layer feature map and the multi-layer feature maps, and obtain the segmentation result of the concrete crack image by performing two-layer convolution on the connection result.

[0040] Optionally, the multi-layer feature maps of different scales include: the second-layer feature map, the third-layer feature map, the fourth-layer feature map, and the fifth-layer feature map;

[0041] The Global Branch module is used for:

[0042] Partition the concrete crack image through the Patch embedding module;

[0043] Add position embeddings to the partitions through the position embedding module and sum them;

[0044] Extract features with global context information from the concrete crack image through densely connected CSwin Transformer Block modules;

[0045] Gradually upsample the features to sequentially obtain the fifth-layer feature map, the fourth-layer feature map, the third-layer feature map, and the second-layer feature map.

[0046] Optionally, the decoder module is used for:

[0047] Connect the first-layer feature map and the second-layer feature map through the Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain the output feature map of the second layer;

[0048] After performing max pooling operation on the output feature map of the second layer, connect the output feature map of the second layer and the third-layer feature map through the Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain the output feature map of the third layer;

[0049] After performing max pooling operation on the output feature map of the third layer, connect the output feature map of the third layer and the fourth-layer feature map through the Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain the output feature map of the fourth layer;

[0050] After performing a max pooling operation on the output feature map of the fourth layer, the output feature map of the fourth layer and the fifth layer feature map are connected through the embedding connection module EC, and a 3×3 convolution is performed on the connection result to obtain the output feature map of the fifth layer;

[0051] The output feature maps of the first layer, the second layer, the third layer, the fourth layer, and the fifth layer are connected to obtain the connection result.

[0052] Optionally, before connecting the output feature maps of the first layer, the second layer, the third layer, the fourth layer, and the fifth layer, the decoder module is further configured to:

[0053] Upsample the output feature maps of the third layer, the fourth layer, and the fifth layer;

[0054] Among them, the output feature map of the third layer is subjected to one upsampling process, the output feature map of the fourth layer is subjected to two upsampling processes, and the output feature map of the fifth layer is subjected to three upsampling processes.

[0055] Optionally, the embedding connection module EC is configured to:

[0056] For two arbitrary input feature maps X1 and X2, average pooling operations, a first convolution operation, and a second convolution operation are respectively performed to obtain X3 and X4;

[0057] Multiply X1 and X3 to obtain X5, and multiply X2 and X4 to obtain X6;

[0058] Add X1 and X2 to obtain X7;

[0059] Add X5, X6, and X7 to obtain the feature map X as the output.

[0060] Optionally, the first convolution operation includes a 1×1 convolution operation, a BN operation, and a ReLU operation, and the second convolution operation includes a 1×1 convolution operation, a BN operation, and a Sigmoid operation.

[0061] Optionally, the densely connected CSwin Transformer Block module is: densely connecting four CSwin Transformer Block modules.

[0062] A third aspect of the embodiments of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method in the above first aspect or any one implementation manner of the first aspect are implemented.

[0063] The fourth aspect of the embodiments of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method in the first aspect or any implementation manner of the first aspect are implemented.

[0064] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:

[0065] The embodiments of the present invention propose a concrete crack image segmentation model based on dual-branch embedding connection, namely the DEC-Net model. The DEC-Net model has a dual-branch encoder structure of CNN and Global Branch. The Global Branch introduces an improved CSwin Transformer Block module with dense connections, which can better capture the overall trend and continuity of crack features. The dense connections strengthen the transferability of crack features and realize feature reuse, and then multi-scale multi-layer feature maps with global context information are obtained through progressive upsampling. At the same time, the CNN takes the original image as input, passes through a 3×3 convolution block without downsampling, which not only strengthens the extraction of local detail information but also retains rich spatial detail information in the image. Finally, the decoder module connects the first-layer feature map and the multi-layer feature maps, and the connection result passes through two layers of convolution to obtain the segmentation result of the concrete crack image. The DEC-Net model of the present invention has excellent segmentation performance and can accurately segment concrete crack images. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0067] Figure 1 It is a schematic diagram of the network structure of the DEC-Net model provided by the embodiments of the present invention;

[0068] Figure 2 It is a schematic diagram of the structure of the ICS module provided by the embodiments of the present invention;

[0069] Figure 3 It is a schematic diagram of the structure of the EC module provided by the embodiments of the present invention;

[0070] Figure 4 It is a schematic diagram of the structure of the USM module provided by the embodiments of the present invention;

[0071] Figure 5It is a schematic diagram of the concrete crack image segmentation device with a double-branch embedded connection provided by an embodiment of the present invention;

[0072] Figure 6 It is a schematic diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0073] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0074] In order to illustrate the technical solutions described in the present invention, the following will be described through specific embodiments.

[0075] In practical applications, there is a problem of how to accurately identify and process slender concrete cracks. Such cracks often have complex shapes and distributions, which pose a significant challenge to image segmentation algorithms. Methods based on CNN have been widely used in the field of image processing, especially in capturing local features and spatial relationships of images. In the task of concrete crack image segmentation, CNN can effectively extract local information such as the texture and shape of cracks, which is crucial for identifying the existence and location of cracks. However, CNN also has an obvious limitation, that is, its receptive field is limited. The receptive field determines the range of images that CNN can process. When facing slender cracks, since the cracks may span a large image area, CNN often has difficulty capturing the overall trend and distribution of the cracks. This limitation is particularly obvious when dealing with slender cracks, which may lead to inaccurate segmentation results. To overcome this limitation of CNN, methods based on the Transformer series are introduced. The Transformer model can learn the global features of cracks from the entire image through the self-attention mechanism, which helps to capture the overall structure and distribution of the cracks. This global perspective makes the Transformer more advantageous when dealing with slender cracks and can better understand the overall trend and distribution of the cracks. However, the Transformer series methods are not a perfect solution either. In terms of dealing with local details, the Transformer is relatively weak and may not be able to accurately identify the subtle changes in the cracks. These subtle changes may be key indicators of crack width, depth, or other important features, which are of great significance for crack assessment and repair.

[0076] Therefore, whether it is a CNN-based method or a Transformer-based method, each has its own focus when dealing with concrete crack images, but both also have certain limitations. CNN is good at capturing local details and spatial relationships, but it is difficult to handle the overall structure and long-range dependencies; while Transformer is good at capturing the overall features of cracks from a global perspective, but it is slightly lacking in dealing with local details. This limitation restricts the accuracy of crack segmentation, making it difficult for us to take both local details and the overall structure into account at the same time. To further improve the accuracy of crack segmentation, it is necessary to comprehensively consider the advantages and limitations of these two methods and seek an effective fusion strategy. Based on this, the embodiment of the present invention proposes a concrete crack image segmentation model based on dual-branch embedding connection, namely the DEC-Net model.

[0077] Figure 1 It is a schematic diagram of the network structure of the DEC-Net model provided by the embodiment of the present invention, including: a CNN module, a Global Branch module, and a decoder module.

[0078] (1) The CNN module is used to perform 3×3 convolution (specifically 3×3 convolution + BN + Relu) on the concrete crack image to obtain the first-layer feature map. Here, CNN takes the original image as input and passes through a 3×3 convolution block. This process only changes the number of channels and does not undergo downsampling, which not only strengthens the extraction of local detail information but also retains rich spatial detail information in the image.

[0079] (2) The Global Branch module is used to extract features with global context information from the concrete crack image through densely connected CSwin Transformer Block modules (hereinafter, a CSwin Transformer Block module is simply referred to as an ICS module, and the module after densely connecting multiple CSwin Transformer Block modules is simply referred to as a DCIB module), and gradually upsample the features to obtain multi-layer feature maps of different scales.

[0080] The Global Branch uses the AM module to extract features with global context information from the concrete crack image. Then, the features with the obtained global context information are gradually upsampled to obtain multi-scale feature maps, which are the second to fifth layers respectively.

[0081] Specifically, as Figure 1As shown, the Global Branch module performs feature extraction using a Transformer structure. The AM module consists of a Patch Embedding module, a DCIB module, and a Position Embedding module. Bilinear represents bilinear, Linear Layer represents a linear layer, and MaxPool represents a max pooling layer.

[0082] The Global Branch module is used for:

[0083] Partitioning the concrete crack image through the Patch Embedding module;

[0084] Adding position embeddings to the partitions through the Position Embedding module and summing them;

[0085] Extracting features with global context information from the concrete crack image through the densely connected CSwin Transformer Block module;

[0086] Performing progressive upsampling on the features to obtain the fifth-layer feature map, the fourth-layer feature map, the third-layer feature map, and the second-layer feature map in sequence.

[0087] Exemplarily, the DCIB module performs dense residual connections on the linear mapping layers of four ICS modules (ICS1, ICS2, ICS3, ICS4). The structure of the ICS model is as Figure 2 shown. Such a structure allows each layer in the network to directly receive the outputs of all previous layers as inputs, thereby making full use of feature information and avoiding waste of features. The linear mapping layer reduces the input feature dimension to 32, which can save the computation of the model. Different ICS layers are tightly connected to maintain the representational ability with a lower feature dimension. The ICS layer retains the structure and window self-attention of the CSwin Transformer Bolck, that is, it retains the advantages of the CSwin Transformer Bolck. However, an EMLP module is proposed based on the MLP operation in the CSwin Transformer Bolck. The MLP has a strong non-linear expression ability; a 1×1 convolution can be regarded as a special linear transformation, which can achieve information fusion at the same position in different channels. By adding the 1×1 convolution and the MLP, the information in different channels can be integrated using the convolutional layer, which helps to improve the model's processing ability for multi-channel inputs. Here, LN, EMLP, and Cross-Shaped Window Self-Attention are all commonly used algorithm layers in neural networks, and this embodiment will not introduce them in detail here.

[0088] (3) The decoder module is used to connect the first-layer feature map and the multi-layer feature maps, and obtain the segmentation result of the concrete crack image by passing the connection result through two layers of convolution.

[0089] In this embodiment, an Embedded Connection Module EC is proposed to fuse the Global Branch and the CNN. The structure is shown in Figure 3 as follows.

[0090] Specifically, the features extracted by the Global Branch are gradually upsampled to obtain feature maps of different sizes, denoted as X1; the output feature map obtained from the previous layer is denoted as X2. First, the feature maps X1 and X2 from different sources are respectively subjected to the average pooling operation Avgpool. By averaging the pixel values within each concrete crack region, the spatial information of the input feature map is retained, enabling the DEC-Net network to better understand the spatial structure in the concrete crack image. The 1×1 convolution can perform feature selection and enhancement without changing the number of feature channels. BN accelerates the training and improves the stability of the model, helping to avoid overfitting. The loss functions ReLU and Sigmoid are used in combination with BN and the 1×1 convolution to further enhance the generalization ability of DEC-Net, helping to reduce the number of parameters and the computational amount.

[0091] After X1 undergoes the average pooling operation, the first convolution operation (1×1 convolution + BN + ReLU), and the second convolution operation (1×1 convolution + BN + Sigmoid), X3 is obtained.

[0092] After X2 undergoes the average pooling operation, the first convolution operation (1×1 convolution + BN + ReLU), and the second convolution operation (1×1 convolution + BN + Sigmoid), X4 is obtained.

[0093] Finally, X1 and X3 are multiplied to obtain X5, X2 and X4 are multiplied to obtain X6, X1 and X2 are added to obtain X7, and X5, X6, and X7 are added to obtain the feature map X as the output.

[0094] Here, referring to Figure 1 , before connecting the first-layer feature map with the output feature maps of the second, third, fourth, and fifth layers, the decoder module is also used to: perform upsampling USM on the output feature maps of the third, fourth, and fifth layers.

[0095] The USM module refers to the operation of gradually upsampling in the decoder. The structure is as shown in Figure 4As shown in the figure. The smallest unit is a combination of 3×3 convolution, BN, and ReLU. The concrete crack feature map with a size of 28×28 obtained from the fifth layer of the network is gradually restored to the feature map with the original size of 224×224 through three smallest units. This process is denoted as the USM3 module. The concrete crack feature map with a size of 56×56 obtained from the fourth layer of the network is gradually restored to the original size feature map through two smallest units. This process is denoted as the USM2 module. The feature map with a size of 112×112 obtained from the third layer of the network is restored to the original size feature map after passing through one smallest unit. This process is denoted as the USM1 module. The sizes of the feature maps obtained from the second and first layers of the network are the same as the original size and do not need to be restored. The output feature maps of these five layers are added together, and then the final segmentation map is obtained through two 3×3 convolutions.

[0096] An embodiment of the present invention proposes a concrete crack image segmentation model based on a dual-branch embedded connection, namely the DEC-Net model. The DEC-Net model has a dual-branch encoder structure of CNN and Global Branch. The Global Branch introduces an improved CSwin Transformer Block module with dense connections, which can better capture the overall trend and continuity of crack features. The dense connections strengthen the transmission of crack features and achieve feature reuse, and then multi-scale multi-layer feature maps with global context information are obtained through progressive upsampling. At the same time, CNN takes the original image as the input and passes through a 3×3 convolution block without downsampling, which not only strengthens the extraction of local detail information but also retains rich spatial detail information in the image. And an embedded connection module, namely the EC module, is proposed for fusing Global Branch and CNN. The purpose of this module is to better fuse the features with global context crack information of different scales obtained from the Global Branch and the local crack information obtained through CNN. The decoder part adopts a progressive upsampling method to directly restore the feature map of each layer to the size of the original image, connect the feature maps of each layer, and then obtain the final segmentation result through two layers of convolution.

[0097] Based on the above DEC-Net model, the process of concrete crack recognition is as follows:

[0098] Obtain a concrete crack image;

[0099] Input the concrete crack image into a preset DEC-Net network structure to obtain the segmentation result of the concrete crack image.

[0100] Through experiments, it is confirmed that the model has excellent segmentation performance.

[0101] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the sequence of execution. The execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0102] Figure 5 FIG. 4 is a schematic structural diagram of a concrete crack image segmentation device 50 with a double-branch embedded connection provided by an embodiment of the present invention, including:

[0103] An acquisition module 51, configured to acquire a concrete crack image.

[0104] A segmentation module 52, configured to input the concrete crack image into a preset DEC-Net network structure to obtain a segmentation result of the concrete crack image.

[0105] Wherein, the DEC-Net network structure includes a CNN module, a Global Branch module, and a decoder module;

[0106] The CNN module is configured to perform 3×3 convolution on the concrete crack image to obtain a first-layer feature map;

[0107] The Global Branch module is configured to extract features with global context information from the concrete crack image through a densely connected CSwin Transformer Block module, and perform progressive upsampling on the features to obtain multi-layer feature maps of different scales;

[0108] The decoder module is configured to connect the first-layer feature map and the multi-layer feature maps, and perform two-layer convolution on the connection result to obtain a segmentation result of the concrete crack image.

[0109] Optionally, the multi-layer feature maps of different scales include: a second-layer feature map, a third-layer feature map, a fourth-layer feature map, and a fifth-layer feature map;

[0110] The Global Branch module is configured to:

[0111] Block the concrete crack image through a Patch embedding module;

[0112] Add position embeddings to the blocks through a position embedding module and add them;

[0113] Extract features with global context information from the concrete crack image through a densely connected CSwin Transformer Block module;

[0114] Perform progressive upsampling on the features to sequentially obtain a fifth-layer feature map, a fourth-layer feature map, a third-layer feature map, and a second-layer feature map.

[0115] Optionally, the decoder module is used to:

[0116] Connect the first-layer feature map and the second-layer feature map through the embedding connection module EC, and perform a 3×3 convolution on the connection result to obtain the output feature map of the second layer;

[0117] After performing a max pooling operation on the output feature map of the second layer, connect the output feature map of the second layer and the third-layer feature map through the embedding connection module EC, and perform a 3×3 convolution on the connection result to obtain the output feature map of the third layer;

[0118] After performing a max pooling operation on the output feature map of the third layer, connect the output feature map of the third layer and the fourth-layer feature map through the embedding connection module EC, and perform a 3×3 convolution on the connection result to obtain the output feature map of the fourth layer;

[0119] After performing a max pooling operation on the output feature map of the fourth layer, connect the output feature map of the fourth layer and the fifth-layer feature map through the embedding connection module EC, and perform a 3×3 convolution on the connection result to obtain the output feature map of the fifth layer;

[0120] Connect the first-layer feature map and the output feature maps of the second, third, fourth, and fifth layers to obtain a connection result.

[0121] Optionally, before connecting the first-layer feature map and the output feature maps of the second, third, fourth, and fifth layers, the decoder module is further used to:

[0122] Upsample the output feature maps of the third, fourth, and fifth layers;

[0123] Among them, the output feature map of the third layer is subjected to one upsampling process, the output feature map of the fourth layer is subjected to two upsampling processes, and the output feature map of the fifth layer is subjected to three upsampling processes.

[0124] Optionally, the embedding connection module EC is used to:

[0125] For two arbitrary input feature maps X1 and X2, perform average pooling operations, a first convolution operation, and a second convolution operation respectively to obtain X3 and X4;

[0126] Multiply X1 and X3 to get X5, and multiply X2 and X4 to get X6;

[0127] Add X1 and X2 to get X7;

[0128] Add X5, X6, and X7 to obtain the feature map X as the output.

[0129] Optionally, the first convolution operation includes a 1×1 convolution operation, a BN operation, and a ReLU operation, and the second convolution operation includes a 1×1 convolution operation, a BN operation, and a Sigmoid operation.

[0130] Optionally, the densely connected CSwin Transformer Block module is: densely connecting four CSwin Transformer Block modules.

[0131] The embodiment of the present invention proposes a concrete crack image segmentation model based on dual-branch embedding connection, namely the DEC-Net model. The DEC-Net model has a dual-branch encoder structure of CNN and Global Branch. The Global Branch introduces an improved densely connected CSwin Transformer Block module, which can better capture the overall trend and continuity of crack features. The dense connection strengthens the transitivity of crack features and realizes feature reuse, and then obtains multi-scale multi-layer feature maps with global context information through step-by-step upsampling; at the same time, the CNN takes the original image as input and passes through a 3×3 convolution block without downsampling, which not only strengthens the extraction of local detail information but also retains rich spatial detail information in the image; and an embedding connection module for fusing Global Branch and CNN is proposed, namely the EC module. The purpose of this module is to better fuse the features with global context crack information of different scales obtained by the Global Branch and the local crack information obtained by the CNN. The decoder part adopts a step-by-step upsampling method to directly restore the feature map of each layer to the size of the original image, connect the feature maps of each layer, and then obtain the final segmentation result after two layers of convolution.

[0132] Figure 6 It is a schematic diagram of an electronic device 60 provided by an embodiment of the present invention. As Figure 6 shown, the electronic device 60 of this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61, such as a concrete crack image segmentation program. When the processor 61 executes the computer program 63, it implements the steps in each of the above-mentioned concrete crack image segmentation method embodiments. Alternatively, when the processor 61 executes the computer program 63, it implements the functions of each module in each of the above-mentioned device embodiments, such as Figure 5 the functions of the modules 51 to 52 shown.

[0133] Exemplarily, the computer program 63 can be divided into one or more modules / units, which are stored in the memory 62 and executed by the processor 61 to implement the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 63 in the electronic device 60. For example, the computer program 63 can be divided into an acquisition module 51 and a segmentation module 52 (modules in the virtual device), and the specific functions of each module are as follows:

[0134] The acquisition module 51 is used to acquire concrete crack images.

[0135] The segmentation module 52 is used to input the concrete crack image into a preset DEC-Net network structure to obtain the segmentation result of the concrete crack image.

[0136] The electronic device 60 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 6 merely examples of the electronic device 60, which do not constitute a limitation on the electronic device 60, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device 60 may further include input / output devices, network access devices, a bus, etc.

[0137] The so-called processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0138] The memory 62 may be an internal storage unit of the electronic device 60, such as a hard disk or memory of the electronic device 60. The memory 62 may also be an external storage device of the electronic device 60, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 60. Further, the memory 62 may also include both an internal storage unit and an external storage device of the electronic device 60. The memory 62 is used to store the computer program and other programs and data required by the electronic device 60. The memory 62 may also be used to temporarily store data that has been output or is to be output.

[0139] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0140] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0141] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0142] In the embodiments provided by the present invention, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0143] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0144] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0145] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0146] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for segmenting concrete crack images with a double-branch embedded connection, characterized in that, Including: Obtain a concrete crack image, input the concrete crack image into a preset DEC-Net network structure, and obtain a segmentation result of the concrete crack image; Among them, the DEC-Net network structure includes a CNN module, a Global Branch module, and a decoder module; The CNN module is used to perform 3×3 convolution on the concrete crack image to obtain a first-layer feature map; The Global Branch module is used to extract features with global context information from the concrete crack image through densely connected CSwin Transformer Block modules, and perform progressive upsampling on the features to obtain multi-layer feature maps of different scales; The decoder module is used to connect the first-layer feature map and the multi-layer feature maps, and perform two-layer convolution on the connection result to obtain the segmentation result of the concrete crack image.

2. The method for segmenting concrete crack images with double-branch embedded connection according to claim 1, wherein, The multi-layer feature maps of different scales include: a second-layer feature map, a third-layer feature map, a fourth-layer feature map, and a fifth-layer feature map; The Global Branch module is used for: Partition the concrete crack image through a Patch embedding module; Add position embeddings to the partitions through a position embedding module and sum them; Extract features with global context information from the concrete crack image through densely connected CSwin Transformer Block modules; Perform progressive upsampling on the features to sequentially obtain the fifth-layer feature map, the fourth-layer feature map, the third-layer feature map, and the second-layer feature map.

3. The method for segmenting concrete crack images with double-branch embedded connection according to claim 2, wherein, The decoder module is used for: Connect the first-layer feature map and the second-layer feature map through an Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain an output feature map of the second layer; After performing a max pooling operation on the output feature map of the second layer, connect the output feature map of the second layer and the third-layer feature map through an Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain an output feature map of the third layer; After performing a max pooling operation on the output feature map of the third layer, connect the output feature map of the third layer and the fourth-layer feature map through an Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain an output feature map of the fourth layer; After performing a max pooling operation on the output feature map of the fourth layer, connect the output feature map of the fourth layer and the fifth-layer feature map through an Embedding Connection module EC, and perform 3×3 convolution on the connection result to obtain an output feature map of the fifth layer; Connect the first-layer feature map and the output feature maps of the second, third, fourth, and fifth layers to obtain the connection result.

4. The method for segmenting concrete crack images with double-branch embedded connection according to claim 3, wherein Before connecting the first-layer feature map and the output feature maps of the second, third, fourth, and fifth layers, the decoder module is further used for: Perform upsampling on the output feature maps of the third, fourth, and fifth layers; Among them, the output feature map of the third layer is subjected to one upsampling process, the output feature map of the fourth layer is subjected to two upsampling processes, and the output feature map of the fifth layer is subjected to three upsampling processes.

5. The method for segmenting concrete crack images with double-branch embedded connection according to claim 3, characterized in that, The embedding connection module EC is used for: For two arbitrary input feature maps X1 and X2, perform average pooling operations, first convolution operations, and second convolution operations respectively to obtain X3 and X4; Multiply X1 and X3 to get X5, and multiply X2 and X4 to get X6; Add X1 and X2 to get X7; Add X5, X6, and X7 to obtain the feature map X as the output.

6. The method for segmenting concrete crack images with double-branch embedded connection according to claim 5, characterized in that, The first convolution operation includes a 1×1 convolution operation, a BN operation, and a ReLU operation, and the second convolution operation includes a 1×1 convolution operation, a BN operation, and a Sigmoid operation.

7. The method for segmenting concrete crack images with double-branch embedded connection according to claim 2, wherein, The densely connected CSwin Transformer Block module is: densely connecting four CSwin Transformer Block modules.

8. A concrete crack image segmentation device with a double-branch embedded connection, characterized in that, Including: An acquisition module for acquiring concrete crack images; A segmentation module for inputting the concrete crack image into a preset DEC-Net network structure to obtain a segmentation result of the concrete crack image; Among them, the DEC-Net network structure includes a CNN module, a Global Branch module, and a decoder module; The CNN module is used to perform 3×3 convolution on the concrete crack image to obtain a first-layer feature map; The Global Branch module is used to extract features with global context information from the concrete crack image through a densely connected CSwin Transformer Block module, and perform step-by-step upsampling on the features to obtain multi-layer feature maps of different scales; The decoder module is used to connect the first-layer feature map and the multi-layer feature maps, and perform two-layer convolution on the connection result to obtain the segmentation result of the concrete crack image.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.