A method and system for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on deep learning

The SegB2-Unet neural network, which combines Segformer-B2 with U-Net, solves the problems of low accuracy and efficiency in ECC crack detection, achieves high-precision crack identification and quantitative analysis, and is suitable for crack detection in infrastructure such as bridges and tunnels, improving the efficiency and accuracy of structural health monitoring.

CN120070833BActive Publication Date: 2025-09-30GUANGDONG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510535529.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-09-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Existing ECC crack detection methods have the disadvantages of low traditional detection efficiency and poor accuracy. Deep learning models are insufficient in capturing long-range dependencies and multi-scale features, and there is a lack of research on the identification of ECC tensile cracks, making it difficult to achieve high-precision crack identification and quantitative analysis.

Method used

The SegB2-Unet neural network, which combines Segformer-B2 with the improved U-Net, is used to construct a semantic segmentation model for high-toughness concrete cracks through multi-scale feature fusion and attention mechanism, which enables accurate identification and quantitative analysis of ECC cracks.

Benefits of technology

The accuracy and efficiency of ECC crack identification have been improved, and the length and width of cracks can be automatically identified. This makes it suitable for crack detection and maintenance of infrastructure such as bridges and tunnels, and improves the efficiency and accuracy of structural health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070833B_ABST
    Figure CN120070833B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of computer vision and high-performance concrete mechanical testing, and discloses a method and system for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method. The method comprises: constructing an image feature dataset based on the features of high-toughness concrete crack image samples; preprocessing the image feature dataset to obtain a preprocessed dataset; constructing a SegB2-Unet neural network based on a pretrained model Segformer-B2 and an improved U-net network, training the SegB2-Unet neural network using the preprocessed dataset to obtain a semantic segmentation model for high-toughness concrete cracks; and identifying high-toughness concrete tensile cracks based on the high-toughness concrete crack semantic segmentation model, extracting the crack length and width, and obtaining crack morphology information. The present invention is suitable for the precise identification and quantitative analysis of cracks in uniaxial tensile tests of high-toughness concrete.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and high-performance concrete mechanical testing, and specifically relates to a method and system for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method. Background Art

[0002] Engineered Cementitious Composite (ECC), a high-performance fiber-reinforced cementitious composite, has garnered significant attention for its unique pseudo-strain hardening behavior and outstanding multi-crack tensile properties, and is increasingly gaining importance in modern civil engineering. Compared to conventional concrete, ECC offers significant advantages in many structural applications due to its high tensile strain capacity, multi-crack behavior, and unique crack control capabilities. Under tensile stress, ECC develops multiple microcracks instead of a single wide crack. Strain can reach 3-8% under tension, and crack widths can be controlled to <100 μm before final failure. This multi-crack behavior enhances the material's ductility, effectively extending the service life of the structure while minimizing the impact of crack propagation on its integrity.

[0003] With the continuous deepening of ECC material research and continuous improvement in preparation technology, the cost of ECC has gradually decreased and its performance has continued to improve. Currently, ECC has been widely used in the repair and reinforcement of concrete beam structures and the construction of new engineering structures. Although ECC materials have excellent crack control capabilities, the occurrence and propagation of cracks remain key factors affecting their structural performance. The formation, distribution, width, and propagation of cracks have a direct impact on the performance of ECC material structures, especially in terms of corrosion, penetration, and durability. If these cracks are not identified and repaired in a timely manner, they may further expand and cause structural damage. Therefore, early detection and monitoring of ECC cracks are crucial, which can help to timely identify potential structural problems and repair them, thereby avoiding more serious damage.

[0004] The traditional crack detection method used in the early days was mainly based on manual inspection, which was not only time-consuming and labor-intensive, but the results were easily affected by the inspector's experience, making it difficult to achieve accurate detection. Therefore, the efficiency and accuracy were relatively low, and data loss or errors were prone to occur.

[0005] Later, nondestructive detection methods such as ultrasound, infrared thermal imaging, and electrical sensing emerged. While ultrasound and infrared thermal imaging can provide more detailed crack information, they require expensive equipment, complex operation, and have a limited measurement range. With the development of computer technology, image processing techniques based on computer vision have emerged. Images are captured using equipment and then preprocessed using computers using techniques such as edge detection, thresholding, clustering, binarization, histogram equalization, and denoising to achieve image segmentation and pixel detection. However, crack identification using image processing has certain limitations. For example, it requires high image quality. Issues such as blur, noise, or uneven lighting can lead to inaccurate crack identification. Furthermore, complex background environments and the diversity of crack morphology increase the difficulty of identification, making it difficult for algorithms to account for all conditions and requiring manual identification of some images.

[0006] With the rapid development of deep learning in recent years, the application of deep learning algorithms to concrete has shown great potential. A growing number of researchers have begun applying deep learning techniques to areas such as concrete structural health monitoring. Concrete crack detection and identification is a key research area. For ECC crack detection and identification, deep learning technology uses multi-layer neural networks to learn patterns and features from samples, enabling the identification and location of objects in images. By training deep neural network models, deep learning-based ECC crack identification methods can achieve automated crack detection with high efficiency, accuracy, and robustness, while avoiding interference from subjective factors. This not only significantly improves the efficiency and accuracy of concrete structure inspections, but also helps effectively ensure the safety and extend the service life of structures.

[0007] In existing research, from the perspective of methods, the semantic segmentation models currently used for ECC and other concrete cracks still have the following deficiencies: (1) The encoder receptive field of traditional convolutional neural networks (such as U-Net) is limited, making it difficult to capture the long-range dependence and multi-scale characteristics of ECC cracks; (2) The machine learning models currently used in this field still have insufficient feature extraction and limited segmentation accuracy when dealing with small cracks, and also consume a lot of time, which is not conducive to the application of the model in practical engineering; from the perspective of research objects, there are currently few studies on ECC uniaxial tensile crack recognition based on deep learning methods. Existing studies also tend to focus on the recognition of ordinary concrete cracks, and lack relevant research on ECC tensile crack recognition. Especially for ECC cracks, the cracking morphology and cracking mode of ECC cracks are different from those of ordinary concrete, and the durability and permeability of ECC structures depend on the crack width. Therefore, crack recognition is performed to provide parameter information about the crack length and width in ECC tension. Summary of the Invention

[0008] In order to solve the problems existing in the prior art, the present invention provides a method and system for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method, which is suitable for accurate identification and quantitative analysis of cracks in uniaxial tensile tests of high-toughness concrete.

[0009] To achieve the above object, the present invention provides the following solutions:

[0010] A method for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method, the method comprising:

[0011] An image feature dataset is constructed based on the features of high-toughness concrete crack image samples;

[0012] Preprocessing the image feature dataset to obtain a preprocessed dataset;

[0013] Constructing a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and training the SegB2-Unet neural network using the pre-processed dataset to obtain a semantic segmentation model of high-toughness concrete cracks;

[0014] Based on the high-toughness concrete crack semantic segmentation model, high-toughness concrete tensile cracks are identified, the length and width of the cracks are extracted, and the crack morphology information is obtained.

[0015] Preferably, the features of the high-toughness concrete crack image sample include: input features and output features;

[0016] The input features include: experimentally produced ECC crack images and manually annotated ECC crack images;

[0017] The output features include: crack semantic segmentation image, and numerical information of crack length and width.

[0018] Preferably, the method of preprocessing the image feature dataset to obtain the preprocessed dataset includes:

[0019] Performing maximum and minimum normalization processing on the image feature dataset to obtain a normalized image dataset;

[0020] The normalized data set is divided based on a random division method to obtain the preprocessed data set.

[0021] Preferably, the method of constructing a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network includes: using Segformer-B2 as an encoder, combining it with a U-Net decoder, and integrating multi-scale feature fusion and attention mechanism. The overall network structure is divided into an encoding part, a bottleneck part, and a decoding part.

[0022] Preferably, in the encoder part, Segformer-B2 is used as the feature extraction network, which first processes the input image through overlapping patch embedding, so that the image is divided into several patches with overlapping areas. Then, the image enters the four-layer Transformer block, each layer of which contains an efficient self-attention mechanism and a hybrid feedforward network to capture crack features of different scales.

[0023] Preferably, the bottleneck part is located between the encoder and the decoder, and is used to further extract high-level features. This part uses 3×3 convolution for feature extraction, and combines batch normalization and ReLU activation function. The bottleneck part also has a channel adjustment function, so that the deep-level features of the encoder adapt to the input requirements of the decoder.

[0024] Preferably, in the decoder part, the U-Net structure is used to gradually restore the resolution of the feature map to generate the crack segmentation result. The decoder consists of multiple upsampling modules, feature fusion modules and convolution operations, namely Conv3×3 +BatchNorm + ReLU. Skip connections are introduced in the decoder so that the high-resolution features of different scales of the encoder can be directly passed to the corresponding decoding layer. An attention module based on the CBAM idea is introduced in the decoder. Specifically, the module first adopts the channel attention mechanism to extract global information through adaptive average pooling and maximum pooling, and calculates the channel attention coefficient through the fully connected layer to weight the input features and enhance the model's attention to important channels. Subsequently, the spatial attention mechanism is used to fuse the spatial information of global average pooling and maximum pooling, and the spatial attention coefficient is calculated through the convolution layer and the Sigmoid activation function to improve the model's sensitivity to crack areas.

[0025] The present invention also provides a system for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method. The system is used in the above method, and the system includes: an acquisition module, a preprocessing module, a construction module, and an identification module;

[0026] The acquisition module is used to construct an image feature data set based on the features of the high-toughness concrete crack image samples;

[0027] The preprocessing module is used to preprocess the image feature dataset to obtain a preprocessed dataset;

[0028] The construction module is used to construct a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and train the SegB2-Unet neural network using the pre-processed data set to obtain a high-toughness concrete crack semantic segmentation model;

[0029] The recognition module is used to identify high-toughness concrete tensile cracks based on the high-toughness concrete crack semantic segmentation model, extract the length and width of the cracks, and obtain crack morphology information.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention proposes an end-to-end ECC tensile crack feature recognition model based on deep learning and multi-scale feature fusion modules. First, the deep learning model combining the Segformer-b2 model and the U-net is used to replace the traditional single network deep learning model, thereby improving the network's feature mining ability and semantic segmentation ability. Secondly, the multi-scale feature fusion and attention mechanism technology are used to further enhance the network's feature segmentation ability, that is, compared with the initial single network model, the time consumption is reduced and the segmentation accuracy is further improved. The present invention uses 466 image samples collected from the experiment to prove that the proposed SegB2-Unet method has good recognition ability in identifying cracks generated during the ECC tensile test, and the corresponding evaluation indicators PA, MPA, MIoU, Recall, Precision, and F1 Score are 99.46, 97.16, 94.50, 94.62, 94.37, and 94.41, respectively. This method is superior to single U-net, Segformer-B2 and other neural network deep learning methods. In summary, the proposed SegB2-Unet method achieves high-precision semantic segmentation of ECC cracks and automatic measurement of crack length and width. This method can be widely applied to crack detection and maintenance in infrastructure such as bridges, tunnels, and roads, providing an efficient and reliable technical solution for structural health monitoring in civil engineering. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 This is a schematic diagram of the dimensions of a test piece according to an embodiment of the present invention;

[0034] Figure 2Schematic diagram of an ECC crack image produced experimentally according to an embodiment of the present invention and an ECC crack image manually annotated;

[0035] Figure 3 This is a schematic diagram of the network architecture of the model according to the embodiment of the present invention;

[0036] Figure 4 Schematic diagram of an actual image according to an embodiment of the present invention and a prediction result according to the present invention;

[0037] Figure 5 This is a schematic diagram for comparing models of embodiments of the present invention;

[0038] Figure 6 This is a schematic diagram comparing the generalization performance of the model according to the embodiment of the present invention;

[0039] Figure 7 Schematic diagram of the performance of the model according to the embodiment of the present invention on the Hao et al. dataset;

[0040] Figure 8 Schematic diagram of the performance of the model according to an embodiment of the present invention on the Deepcrack dataset;

[0041] Figure 9 This is a schematic diagram of ablation experiment performance indicators according to an embodiment of the present invention;

[0042] Figure 10 This is a schematic diagram of ablation experiment results according to an embodiment of the present invention;

[0043] Figure 11 This is a schematic diagram showing the comparison before and after data enhancement according to an embodiment of the present invention;

[0044] Figure 12 The crack length analysis process of an embodiment of the present invention, wherein (a) is the original crack; (b) is the crack segment skeleton; (c) is the length identification result;

[0045] Figure 13 The crack width analysis process of the embodiment of the present invention, wherein (a) is the original crack; (b) is the crack segment skeleton; (c) is the width calculation;

[0046] Figure 14 This is an analysis of the local crack width of an embodiment of the present invention, where (a) is a local crack; (b) is the skeleton of the local crack segment; (c) is the local crack width;

[0047] Figure 15 This is a flow chart of a method for identifying and characterizing the behavior of multiple cracks in high-toughness concrete based on a deep learning method in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0050] Example 1

[0051] like Figure 15 As shown, the present invention provides a method for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method, the method comprising:

[0052] An image feature dataset is constructed based on the features of high-toughness concrete crack image samples;

[0053] Preprocessing the image feature dataset to obtain a preprocessed dataset;

[0054] Constructing a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and training the SegB2-Unet neural network using the pre-processed dataset to obtain a semantic segmentation model of high-toughness concrete cracks;

[0055] Based on the high-toughness concrete crack semantic segmentation model, high-toughness concrete tensile cracks are identified, the length and width of the cracks are extracted, and the crack morphology information is obtained.

[0056] In this example, the present invention proposes a deep learning model (SegB2-Unet) based on the combination of Segformer-B2 and U-Net for high-precision semantic segmentation of uniaxial tensile cracks in high-ductility concrete (ECC). This method combines the global feature extraction capabilities of the Transformer architecture with the high-resolution feature recovery capabilities of the U-Net architecture. Combined with multi-scale feature fusion and an attention mechanism, it effectively improves crack feature extraction, enabling precise crack segmentation and accurate identification of length and width. Furthermore, while ensuring recognition accuracy, this method optimizes computational efficiency and reduces computation time, making it suitable for large-scale crack identification tasks. Data processing, training, and analysis in this invention are based on the PyCharm 2021 integrated development environment, using Python 3.11 for programming. PyTorch 2.4.1 was selected as the deep learning framework. The workstation used was equipped with an NVIDIA GeForce RTX 3060 Ti GPU and a 13th Gen Intel(R) Core(TM) i5-13400F @ 2.50 GHz CPU, respectively. CUDA 12.4 is used as a parallel computing platform to accelerate training.

[0057] The specific implementation steps are as follows:

[0058] In this example, we propose a deep learning model (SegB2-Unet) combining Segformer-B2 and U-Net for semantic segmentation of uniaxial tension cracks in high-ductility concrete (ECC). This model combines the global feature extraction capabilities of the Transformer with the fine segmentation capabilities of a convolutional neural network (CNN) to achieve high-precision recognition of complex crack patterns and further extract crack length and width to support crack morphology analysis and structural health monitoring.

[0059] In this embodiment, the Segformer-B2 principle is as follows: Segformer is a semantic segmentation framework that combines the Transformer model with a lightweight multi-layer perceptron (MLP) decoder. Its encoder utilizes a hierarchical Transformer design, capable of extracting multi-scale features and avoiding positional encoding, thus maintaining stable performance even when the test resolution differs from the training resolution. This design ensures that Segformer possesses both efficient computation and powerful feature representation capabilities, enabling it to extract both global and local feature information in crack identification tasks, improving segmentation accuracy and generalization.

[0060] U-Net Principle: U-Net is a classic CNN architecture widely used in semantic segmentation tasks. Its structure consists of a symmetrical encoder (downsampling path) and decoder (upsampling path), forming a typical U-shaped architecture. By introducing skip connections between the encoder and decoder, U-Net effectively fuses features at different scales and enhances the ability to recover high-resolution spatial information. This design enables U-Net to accurately extract crack edges in crack semantic segmentation tasks, ensuring the completeness of small crack detection and reducing false or missed detections due to resolution loss.

[0061] SegB2-Unet model design: In the SegB2-Unet architecture, the encoder utilizes Segformer-B2, leveraging its powerful global feature extraction capabilities to ensure effective recognition of cracks of varying scales. The decoder employs an improved U-Net architecture, combining multi-scale feature fusion and an attention mechanism to further enhance the model's ability to identify subtle cracks. This design leverages the Transformer's strengths in global modeling while combining the CNN architecture's ability to recover local details. This enables the model to maintain high-precision crack segmentation even in complex backgrounds, and further extracts crack geometric parameters (length and width) to support crack morphology analysis and engineering applications.

[0062] Attention Mechanism: In the SegB2-Unet model proposed in this paper, an attention mechanism is introduced into the decoder to enhance the detection capability of crack areas, reduce background interference, and improve the accuracy of crack segmentation. Cracks are often characterized by being slender, non-uniform, and low-contrast, making it easy for traditional CNNs to miss important information during feature extraction. Through the attention mechanism, the model can dynamically enhance the expression of crack features, suppress irrelevant background information, and optimize crack segmentation. This paper adopts the CBAM (Convolutional Block Attention Module) mechanism, which uses two consecutive steps to adaptively weight the input features:

[0063] First comes channel attention, then spatial attention. In the channel attention stage, the module generates 1×1 feature descriptors in the spatial dimension through adaptive average pooling and adaptive maximum pooling. These two pooling methods capture the global average information and the most significant activation of the input features, respectively.

[0064] Next, the two descriptors pass through a shared fully connected network, which includes a dimensionality reduction layer (using a linear transformation to reduce the number of channels to 1 / reduction of the original number of channels), a ReLU activation function, and a linear layer to restore the number of channels to extract inter-channel dependencies. The outputs of the two fully connected layers are added together and normalized to the (0,1) range using Sigmoid activation to form refined channel attention weights. The original features are then weighted channel by channel to highlight key information.

[0065] Then, in the spatial attention stage, the module calculates the average and maximum values ​​of the channel-weighted feature maps along the channel dimension, generating two spatial description maps that reflect the overall activation level and the most significant response of the local region, respectively. These two description maps are concatenated along the channel dimension and fused through a convolutional layer (with a kernel size of 7 to ensure a sufficient receptive field). Finally, a sigmoid activation is performed to generate the spatial attention map.

[0066] Finally, the spatial attention map is used to weight the feature map pixel by pixel, effectively enhancing the features of key areas while suppressing background noise. This dual attention mechanism helps the model more accurately capture subtle and non-continuous features when handling tasks such as crack detection, significantly improving overall segmentation accuracy.

[0067] The present invention further optimizes the performance of SegB2-Unet in ECC crack recognition through spatial attention mechanism and channel attention mechanism, enabling it to achieve efficient and accurate semantic segmentation in complex crack environments, providing reliable technical support for concrete structure health monitoring.

[0068] Multi-Scale Feature Fusion: Cracks vary greatly in morphology and size. The width, length, and depth of different cracks can differ significantly within the same image. Relying solely on single-scale features for segmentation makes it difficult to achieve robust detection results. Therefore, this paper introduces a Multi-Scale Feature Fusion (MSFF) strategy into SegB2-Unet to simultaneously leverage local detail features and global semantic information to improve crack detection accuracy.

[0069] The proposed method uses a Segformer-B2 encoder to extract multi-scale features. The model extracts feature maps of different resolutions at different Transformer layers. These feature maps contain high-resolution local details (crack edges) and low-resolution global structural information (overall crack morphology).

[0070] To fully utilize multi-scale information, the model combines skip connections with a two-stage up-sampling mechanism during the decoding process to fully fuse low-resolution global features with high-resolution detail information, thereby enhancing the reconstruction of crack regions. Specifically, each decoding stage first performs preliminary upsampling through transposed convolution, then precisely resizes the feature maps through bilinear interpolation to ensure strict alignment with the corresponding encoder features. This is followed by concatenation and feature fusion. Furthermore, the incorporation of a channel-wise and spatial attention (CBAM) mechanism helps enhance the representation of crack regions while suppressing background noise and improving the robustness of crack detection. Experimental results demonstrate that this decoding strategy effectively improves the detection of cracks of different scales, avoiding the poor performance of single-scale models for cracks of specific sizes. It also enhances the integrity of crack morphology, ensuring that the segmentation results cover the entire crack region and reducing crack breakage or omissions, thereby improving the model's crack recognition capability. The experimental results of multi-scale feature fusion show that: it improves the detection ability of cracks of different scales and avoids the problem of poor performance of single-scale models under certain crack sizes; strengthens the integrity of crack morphology, ensures that the detection results cover the entire crack area, and avoids crack breakage or omission; improves the generalization ability of the model, adapts to different types of cracks and complex background environments, and improves its applicability in real engineering environments.

[0071] The SegB2-Unet proposed in this paper effectively improves the accuracy and robustness of crack identification by integrating a multi-scale feature fusion mechanism with skip connections and upsampling strategies. This method can adapt to cracks of varying sizes and shapes, making SegB2-Unet a promising candidate for broad application in ECC structural health monitoring, meeting the demand for automated crack identification in infrastructure such as bridges, roads, and tunnels.

[0072] In this example, the model was trained and tested using both public and experimental datasets to comprehensively evaluate the model's detection performance and generalization capabilities under various crack environments. A high-quality public dataset containing a variety of crack types was selected to verify the model's adaptability to varying crack morphologies, background conditions, and size variations. The experimental dataset, obtained from laboratory tests and containing images of uniaxial tensile cracks in high-toughness concrete (ECC), was used to further verify the model's reliability and stability in real-world engineering applications.

[0073] The cracks used in the experimental dataset were generated through uniaxial tension testing of dog-bone specimens made of ECC material. The tests were conducted in a controlled environment, with the cracks propagating gradually under standard loading conditions. This ensured that the crack formation process was consistent with actual engineering conditions, thereby enhancing the model's applicability and engineering value.

[0074] Through systematic testing on public and experimental datasets, the detection capabilities of the proposed model in different environments have been comprehensively analyzed and quantitatively evaluated using PA (pixel accuracy), mIoU (mean intersection over union), Recall, Precision, and F1 Score to ensure the reliability, stability, and generalization ability of the model.

[0075] Public Datasets: This study selected several publicly available concrete crack datasets to evaluate the model's generalization and robustness under various crack environments. These datasets contain crack images from diverse backgrounds, lighting conditions, and crack morphologies, encompassing smooth and rough surface cracks, shallow cracks, and deep cracks. This effectively improves the model's adaptability across diverse application scenarios. These datasets are professionally annotated and widely used in crack recognition research, providing high-quality supervision data for model training.

[0076] The model in this paper was trained and tested directly on publicly available datasets without any additional data processing to ensure fairness and objectivity in the comparative experiments. All publicly available datasets were independently used for model evaluation to verify the model's adaptability and generalization across different data sources.

[0077] Experimental dataset: In this paper, the experimental dataset consists of ECC concrete tensile test images actually obtained in the laboratory, which is mainly used to evaluate the performance of the model in real experimental crack identification scenarios.

[0078] Experimental specimens were prepared using high-ductility concrete (ECC) with a mix ratio consisting of cement, fly ash, silica fume, sand, water, a water-reducing agent, and fiber. After 28 days of standardized curing, specimens were formed. Specimen dimensions were based on JSCE standards. Uniaxial tensile tests were conducted on a universal testing machine at a loading rate of 0.5 mm / min. During the tests, crack formation and propagation were recorded and annotated using a high-resolution camera. Ultimately, a high-quality experimental dataset consisting of 466 specimens was constructed.

[0079] The design of this experimental dataset ensures the authenticity of the crack morphology, and its crack growth pattern is consistent with the crack propagation mechanism in actual engineering structures, so it is of great value for the engineering adaptability evaluation of the model. The introduction of the experimental dataset not only enables the model to have good generalization ability on the public dataset, but also maintains high recognition accuracy in real experimental crack recognition tasks. Figure 1 、 Figure 2 shown.

[0080] In this embodiment, data preprocessing: In order to further improve the generalization ability and stability of the model, the experimental dataset was preprocessed before training to enhance the model's detection ability under different crack morphologies and complex backgrounds.

[0081] First, we use image rotation (e.g. 、 ), horizontal flipping, and vertical flipping are used to expand the data samples, so that the model can adapt to the detection tasks of cracks in different directions and enhance the ability to recognize non-uniform crack growth patterns.

[0082] Furthermore, to speed up model training and enhance its ability to identify small-scale cracks, we split a single sample into multiple 352×480 small samples for training. This approach not only reduces the GPU computational burden but also ensures the model can learn more fine-grained crack information, optimizing crack edge detection and detail recovery capabilities.

[0083] In summary, the present invention adopts rigorous experimental design and data preprocessing methods to ensure the stability, robustness and efficient computing capability of the model under different environments, making it suitable for complex crack morphology detection and structural health monitoring tasks.

[0084] In this embodiment, the model architecture is as follows Figure 3 As shown in the figure, the proposed SegB2-Unet uses Segformer-B2 as the encoder, combined with a U-Net decoder, and integrates multi-scale feature fusion and an attention mechanism to enhance ECC crack detection. The overall network structure can be divided into the encoder, bottleneck, and decoder.

[0085] Encoder: In the encoder, Segformer-B2 is used as the feature extraction network. This network first processes the input image using Overlap Patch Embeddings, segmenting the image into several overlapping patches to minimize information loss. The image then enters a four-layer Transformer Block, each of which incorporates an efficient self-attention mechanism and a Mixed Feedforward Network (Mix-FFN). This architecture ensures strong global feature extraction and captures crack characteristics at different scales. During this process, Transformer blocks at different levels output feature maps of different scales, with sizes of (64, H / 4, W / 4), (128, H / 8, W / 8), (320, H / 16, W / 16), and (512, H / 32, W / 32), respectively, to ensure multi-scale crack perception.

[0086] Bottleneck layer: Located between the encoder and decoder, the bottleneck layer further extracts high-level features and reduces redundant information. This layer uses 3×3 convolutions (Conv3×3) for feature extraction, combined with batch normalization and ReLU activation functions to ensure effective information propagation and mitigate the vanishing gradient problem. The bottleneck layer also features channel resizing, expanding the number of channels from the 512 output channels of the Transformer block to 1024. Deep-level features are convolved to extract high-level features before returning to a 512-channel structure. This process not only integrates information at a deeper level, compressing and refining features, but also provides more refined high-level semantic information for the decoder stage and adapts to the decoder's input requirements.

[0087] Decoder: In the decoder, a progressive upsampling approach is used to restore the resolution of feature maps to produce highly accurate crack predictions. The decoder consists of a two-stage upsampling module combining multiple transposed convolutional layers (ConvTranspose2d) with bilinear interpolation, as well as a convolutional fusion layer (Conv2d) and convolution operations (Conv3×3 + BatchNorm + ReLU). Transposed convolutions are used for feature learning, while interpolation ensures precise size matching to optimize feature alignment via skip connections. This architecture not only restores crack boundary details but also improves the smoothness and continuity of the predictions. The transposed convolutional layer module performs inverse convolution with a learnable convolution kernel to achieve upsampling. The trainable convolution kernel allows for learning an upsampling strategy suitable for crack feature recovery. For crack detection, preserving texture details is crucial for edge recognition, so a learnable upsampling strategy is more beneficial for crack boundary recovery. While the decoder of this model still employs upsampling and skip connections, its overall design is not a simple symmetrical structure; rather, it is carefully tailored to fully accommodate the multi-scale features output by the MiT-B2 encoder. First, before entering the decoding stage, a bottleneck module processes the high-dimensional features of the final encoder layer, aiming to compress redundant information and extract more discriminative feature representations. This bottleneck module not only adjusts the number of channels but also effectively improves the compactness and discriminability of features through a combination of convolution, batch normalization, and nonlinear activation functions. Subsequently, the decoder employs a multi-stage upsampling strategy to gradually restore the spatial resolution of the feature maps. In each stage, a transposed convolution (ConvTranspose2d) module performs preliminary upsampling, followed by bilinear interpolation to precisely resize the feature maps. This not only increases the spatial size of the feature maps but also effectively transforms the features through trainable parameters. Because transposed convolution may not accurately restore the target size in some cases, bilinear interpolation is introduced to fine-tune the scale of the upsampled feature map to ensure strict alignment with the corresponding skip-connected feature map. This effective combination of transposed convolution and interpolation effectively alleviates the output size mismatch of the MiT-B2 encoder at multiple scales and ensures seamless information integration. Furthermore, to compensate for spatial detail that may be lost during information transfer in deep networks, skip connections are fully utilized at each stage, directly transferring high-resolution features from the encoder to the corresponding decoding layer, enhancing the ability to capture crack edges and other subtle details.Each skip connection is followed by a convolutional fusion layer (3×3 convolution, BatchNorm, and ReLU) to further fuse the upsampled features with the features passed from the encoder. This convolutional fusion layer not only alleviates the matching issues between different semantic levels but also refines the reconstructed feature information, ensuring smoothness and continuity in the final output, thereby achieving high-precision crack semantic segmentation. Furthermore, to further enhance feature representation and capture edge details, an attention module based on the CBAM concept is introduced in the final decoder stage. After being applied to the decoder output features, this module first uses a channel-wise attention branch to weight the importance of different channels, automatically adjusting the response of each channel's features to highlight key semantic information. Subsequently, a spatial attention branch carefully analyzes the spatial distribution of the feature map, enhancing regions with important local information (such as crack edges) while suppressing redundant or noisy information. This combined strategy further optimizes feature fusion, improving the smoothness and continuity of the segmentation results, and enhancing sensitivity to subtle crack features.

[0088] The Convolutional Block Attention Module (CBAM) attention mechanism consists of two parts: channel attention and spatial attention. The purpose of channel attention is to learn the importance of different channels and enhance the feature responses of important channels; the purpose of spatial attention is to learn the importance of different spatial locations in the feature map and highlight key areas, such as cracks. The CBAM attention mechanism formula is as follows:

[0089] in, Represents the Sigmoid function; Final features generated for the model and Ask the average and maximum values ​​of channel pooling respectively; The convolution kernel size is Convolution operation; Represents the output of the channel attention mechanism; Represents the output of the spatial attention mechanism.

[0090] Final output: Finally, in the final stage of the decoder, a 1×1 convolutional layer combined with a Sigmoid activation function is used to generate a crack segmentation map. In this segmentation map, the pixel value corresponding to the crack area is 1, while the pixel value of the background area is 0, ensuring that the model can maintain high-precision crack recognition capabilities at different scales. Overall, the SegB2-Unet model combines the global feature extraction advantages of Segformer-B2 with the excellent performance of the U-Net structure in fine-grained segmentation. By fusing multi-scale features and introducing the CBAM attention mechanism, the model significantly improves the accuracy and robustness of crack recognition. Experimental results show that the model performs well in the ECC uniaxial tensile crack recognition task, demonstrating its application potential in the fields of automated crack detection and structural health monitoring.

[0091] Among them, Skip Connections: In terms of Skip Connections, due to the symmetry of the U-Net structure, the high-resolution features extracted early by the encoder can be directly transferred to the corresponding level of the decoder through skip connections, thereby effectively compensating for feature loss and improving segmentation accuracy. In crack recognition tasks, cracks often appear slender or discontinuous. Ordinary deep networks may cause small cracks to be ignored or mis-segmented, while skip connections can retain shallow local detail information and pass it to the decoding stage, so that the spatial information of the cracks can be preserved. At the same time, this connection method can significantly improve the detection effect of crack edges, avoid the problem of blurred or incomplete edges, improve the overall segmentation accuracy, and reduce information loss, making the final segmentation result more accurate.

[0092] Convolutional Block Attention Module (CBAM): The attention module employs the CBAM (Convolutional Block Attention Module) mechanism, which adaptively weights input features in two sequential steps: first, channel-wise attention, and then, spatial attention. During the channel-wise attention stage, the module generates 1×1 feature descriptors in the spatial dimension through adaptive average pooling and adaptive max pooling, respectively capturing the global average information and the most significant activations of the input features. These two descriptors then pass through a shared fully connected network, which includes a dimensionality reduction layer (using a linear transformation to reduce the number of channels to 1 / reduction of the original number of channels), a ReLU activation function, and a linear layer to restore the number of channels to extract inter-channel dependencies. The outputs of the two fully connected layers are summed and normalized to the (0, 1) range using a sigmoid activation to form refined channel-wise attention weights. These weights are then applied to the original features channel by channel to highlight key information. Subsequently, in the spatial attention stage, the module calculates the average and maximum values ​​along the channel dimension for the channel-weighted feature map, generating two spatial descriptors reflecting the overall activation level and the most significant response of the local region, respectively. These two descriptors are concatenated along the channel dimension and fused through a convolutional layer (with a kernel size of 7 to ensure a sufficient receptive field). Finally, a sigmoid activation is used to generate the spatial attention map. Finally, the spatial attention map is used to perform pixel-by-pixel weighting on the feature map, effectively enhancing features in key areas while suppressing background noise. This dual attention mechanism helps the model more accurately capture subtle and non-continuous features when handling tasks such as crack detection, significantly improving overall segmentation accuracy.

[0093] In summary, the SegB2-Unet proposed in the present invention fully combines the advantages of Segformer-B2 and U-Net, giving the model a powerful crack recognition capability. As an encoder, Segformer-B2 provides global feature extraction capabilities, can effectively capture long-distance feature associations, and enhance the recognition capability of complex crack morphologies. The decoding part of the U-Net structure retains high-resolution spatial features through jump connections, ensuring that the detailed information of the crack area is not lost, thereby improving the recognition accuracy of the crack. In addition, the multi-scale feature fusion and attention mechanism introduced in the model further optimize the crack recognition effect, enabling the model to focus on the crack area, reduce background interference, and improve segmentation accuracy. Experimental results show that the structure exhibits superior performance in the ECC uniaxial tensile crack recognition task, can accurately segment the crack area, and provide reliable technical support for subsequent crack analysis, quantitative evaluation and structural health monitoring.

[0094] In this embodiment, the present invention adopts an image skeleton extraction method and a distance transformation method to extract crack length and width respectively.

[0095] In terms of length extraction, we first convert the binary segmentation image into a single-pixel wide skeleton. The topological center line of the image is retained through iterative erosion operations to generate a single-pixel wide skeleton. By observing the semantic segmentation results, we found that most cracks have multiple bifurcation points, which makes extraction more difficult. Therefore, we introduce a bifurcation point detection method to traverse the skeleton pixels and count the number of pixels at a certain pixel. 8 - The number of non-zero pixels in the neighborhood. If a pixel is a skeleton point and the number of neighborhood statistics 3, then the pixel will be determined as a bifurcation point. The formula is as follows:

[0096]

[0097] Among them, count represents the total number of statistics; Represents a binary image matrix, which only has two cases: 0 and 1; Indicates the coordinates of the pixel point to be judged in the binary image matrix; The coordinates in the matrix are Pixel value of The pixel to be judged currently The pixel value of .

[0098] After bifurcation point detection, the pixel values ​​at the bifurcation points are set to 0, disconnecting the connection and decomposing complex cracks into multiple independent segments. This prevents intersecting cracks from being misidentified as single cracks during subsequent connected domain analysis. After skeleton extraction and bifurcation point processing, the cracks in the image are segmented into multiple independent components. The resulting images of the independent crack segments are then subjected to connected domain analysis and labeled. Subsequently, all identified crack segments are subjected to area filtering. If the length of a crack segment is 0 or less than 1 mm after conversion, processing of that segment is skipped. This eliminates invalid crack segments, reduces the impact of noise, and reduces the subsequent computational effort.

[0099] After completing the marking of independent crack segments, the length of the cracks in the skeleton graph is calculated. First, the skeleton graph is modeled as a graph structure, with the skeleton pixels regarded as nodes of the graph. The distance between nodes is calculated based on the adjacency relationship. The distance between horizontally or vertically adjacent nodes is 1 pixel, and the distance between diagonally adjacent nodes is Pixels (about 1.414). The Dijkstra algorithm is used to find the longest path between the two endpoints in the skeleton graph. The length of this path is the length of the crack segment. Since this length is composed of pixels, we need to convert it into physical length. The corresponding relationship between image pixels and physical size in the present invention is roughly calculated to obtain Pixels. Therefore, the physical length can be obtained by dividing the pixel length by 57.0, that is:

[0100] in, Indicates physical length; Indicates pixel length.

[0101] To extract the width, similar to calculating the length, the image needs to be skeletonized and bifurcation points processed. Once this processing is complete, the segmented, independent skeleton crack segments are mapped to the cracks in the corresponding locations in the original binary image to calculate the crack width. Once mapped, each crack segment is extracted and a local distance transform is performed to obtain the half-width of the local area. The distance transform values ​​of the skeleton pixels of the local crack segment are then averaged and multiplied by 2 to obtain the average crack width. This calculation formula is as follows:

[0102] in, Represents crack segment skeleton pixels Euclidean distance to the nearest background pixel (i.e. half-width); Indicates the coordinates of the current skeleton pixel; Represents the coordinates of background pixels.

[0103] in, represents the half-width sum of all skeleton pixels in the crack segment; Represents the total number of skeleton pixels.

[0104] In this example, the present invention implements refined model training and optimization for the semantic segmentation of ECC uniaxial tension cracks, ensuring that SegB2-Unet can efficiently and stably complete the task. The entire training process includes key steps such as data preparation, training configuration, optimization strategy, and inference. Furthermore, the present invention not only segments crack regions but also provides parameter information on crack length and width, providing comprehensive data support for subsequent structural health monitoring.

[0105] In terms of data preparation, the dataset was first divided into a training set (80%), a validation set (10%), and a test set (10%) in a ratio of 8:1:1. This ensured that the model could learn crack characteristics from sufficient samples. Generalization was also assessed on the validation set. Data were normalized before input, with pixel values ​​normalized to the range [0, 1] to improve training stability. Furthermore, to further enhance the model's generalization, data augmentation techniques were applied during training, including rotation (±45° and ±90°) and flipping. Images were also segmented into 352×480 pixels for training to increase training speed and enhance the model's ability to recognize cracks in different orientations.

[0106] In terms of training configuration, this study used the Adam optimizer for weight updates, with an initial learning rate set to 0.0001. ReduceLROnPlateau was used as the learning rate scheduler, with a decay factor of 0.1 and a patience value of 10. During training, the learning rate was dynamically adjusted based on changes in the validation loss to improve the stability of model convergence and prevent overfitting. The loss functions used were Dice Loss and OHEM Focal Loss. Dice Loss mitigates data class imbalance by calculating the overlap between the predicted results and the true labels, effectively improving the detection capability of small target (crack) regions. Focal Loss further emphasizes the focus on difficult-to-classify samples and introduces an Online Hard Example Mining (OHEM) strategy, which selects only samples with high loss values ​​for gradient backpropagation, thereby optimizing the model's learning of complex crack regions. The final loss is the weighted sum of Dice Loss (weight 0.6) and OHEM Focal Loss (weight 0.4), the smoothing parameter of Dice Loss is 0.000001, the positive class weight factor (alpha) of OHEM FocalLoss is 0.25, the focus factor (gamma) is 2.0, the OHEM ratio (ohem_ratio) is 0.7, and the total number of training rounds (Epochs) is 100. After each round of training, the pixel accuracy (PA), mean intersection over union (mIoU), recall rate (Recall), accuracy (Precision) and F1Score are calculated to measure the performance of the model.

[0107] The formula is as follows:

[0108]

[0109]

[0110]

[0111] in, is the smoothing coefficient; The model predicts The value of pixels; is the true label The value of pixels; is the intersection of the predicted crack area and the actual crack area (i.e., the sum of correctly predicted crack pixels); is the sum of the predicted crack area and the actual crack area (a measure of the overall size of the crack area); is the class weight factor, which gives higher weight to the crack area (positive class) to alleviate the class imbalance problem; To focus on the factors, we can increase the attention to the samples that are difficult to classify; To dynamically adjust the weights of difficult and easy samples, The larger it is, the more attention the model pays to misclassified samples; and They are Dice Loss weight and FocalLoss weight respectively.

[0112] The combined loss function can optimize the overall crack morphology while enhancing the recognition ability of crack boundaries and fine cracks, making the model perform better in the ECC uniaxial tensile crack prediction task.

[0113] During inference, the crack image to be identified first undergoes preprocessing consistent with the training phase (primarily tensor transformations) before being fed into the trained SegB2-Unet model. For large images, this paper employs a sliding window strategy, dividing the image into fixed-sized blocks and performing predictions block by block. The model utilizes a multi-scale feature extraction module within each window, efficiently fusing features at all levels through skip connections and an attention mechanism to capture the subtle structure of the crack. During the decoding phase, spatial information is gradually restored, ultimately generating a binary segmentation mask (crack regions are labeled as 1 or 255, and background as 0).

[0114] In addition, in order to improve the prediction effect, the output results are also post-processed by non-local mean denoising, morphological dilation and median filtering to ensure that the segmentation results are smoother and more coherent.

[0115] Through the above-mentioned refined training and optimization strategies, the SegB2-Unet model of the present invention can achieve efficient, stable and accurate crack segmentation in the ECC uniaxial tensile crack identification task, while accurately identifying the length and width of the cracks, providing reliable data support and technical guarantees for subsequent structural health monitoring, crack development analysis and engineering maintenance.

[0116] In this embodiment, this project uses seven evaluation indicators to evaluate the semantic segmentation model and prediction model for ECC crack images, namely pixel accuracy (PA), mean pixel accuracy (MPA), mean intersection over union (MIoU), recall, precision, and F1 score.

[0117] The calculation of evaluation indicators of this model is based on the confusion matrix. Indicates that it was originally The class is predicted to be classes, namely true positives (TP) and true negatives (TN), Indicates that it was originally The class is predicted to be Class, namely false positive (FP) and false negative (FN), if the The class is positive, When, then Indicates TP, Indicates TN, represents FP, Indicates FN.

[0118] The specific calculation formula of the model evaluation index is as follows:

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125] Figure 4 The prediction results of the present invention are demonstrated on a large, high-resolution image (1760*3840). For this large, high-resolution image, the present invention employs a sliding window strategy for step-by-step inference of local regions. Subsequently, non-local means denoising, morphological dilation, and median filtering are sequentially introduced to post-process the prediction results, removing noise, connecting fractured regions, and reducing false detection rates. This method achieves efficient processing of large images while ensuring model robustness while also balancing accurate segmentation and preservation of edge details, demonstrating its advantages in crack detection applications.

[0126] Experimental results show that the present invention can efficiently process large-size images and effectively obtain accurate prediction results.

[0127] Figure 5This paper demonstrates the performance evaluation of the proposed SegB2-Unet model on a self-constructed dataset. As shown in the figure, a single Segformer-B2 achieves approximately 98.41% pixel accuracy (PA), 90.32% mean pixel accuracy (MPA), mean intersection over union (MIoU), recall, precision, and F1 score (F1 Score), respectively. In comparison, the U-Net achieves significant improvements on these metrics, reaching 99.06%, 96.13%, 92.19%, 92.82%, 91.41%, and 92.10%, respectively. Segformer-B0 and Segformer-B1 achieve lower performance on PA, MPA, MIoU, recall, precision, and F1 score. It is worth noting that while SegB2-Unet's PA (99.01%) is slightly lower than U-Net's 99.06%, SegB2-Unet's Recall (93.00%) is slightly higher than U-Net's. Its Precision and F1 Score reach 91.73% and 92.36%, respectively. Therefore, SegB2-Unet slightly outperforms U-Net in overall performance. These results demonstrate that SegB2-Unet achieves a higher recall rate for target regions while maintaining high overall performance, providing potential value for subsequent applications in scenarios sensitive to missed detection, aiming to achieve more comprehensive segmentation performance improvements.

[0128] Figure 6 、 Figure 7 、 Figure 8 This paper demonstrates the performance evaluation and prediction results of the proposed SegB2-Unet model on two public datasets. The experimental results show that the model achieves mean intersection-over-union (mIoU) values ​​of 0.858 and 0.856 on both datasets, respectively, demonstrating high segmentation accuracy and robustness in image segmentation tasks. Furthermore, the recall rate is 0.846 for the first dataset and 0.809 for the second, suggesting that the model achieves high crack detection completeness on the first dataset, but may experience some missed detections on the second dataset. The corresponding precision rates are 0.852 and 0.893, respectively. The higher precision on the second dataset indicates that the model effectively reduces false positives on this dataset. The comprehensive F1 score is 0.848 and 0.849, further demonstrating the model's stable performance in balancing recall and precision. In summary, the performance of this model is primarily reflected in its image segmentation accuracy, detection completeness, and false positive control, as well as its good generalization and robustness across various datasets.

[0129] like Figure 9 、 Figure 10 As shown in Figure 2, the proposed model systematically validates the contribution of each key algorithmic module to model performance through ablation experiments and conducts an in-depth evaluation on a self-built database without data augmentation. The core goal of the ablation experiments is to analyze the importance of the U-Net structure, multi-scale feature fusion, and attention mechanism in the SegB2-Unet semantic segmentation task, as well as the role and impact of each module in crack identification.

[0130] Experimental results show that the overall segmentation performance of the base model (Segformer-B2) is relatively limited without the use of additional improved modules. Specifically, its PA (pixel accuracy) is 98.41%, mIoU (mean intersection over union) is 86.37%, recall is only 81.20%, precision is 89.90%, and F1 score is only 85.33%. This indicates that when relying solely on Segformer-B2 for crack identification, the model's detection capabilities are limited, especially in identifying crack integrity and preserving edge information. The model performs well on relatively simple crack structures, but suffers from a high rate of missed detections and false detections when faced with complex crack morphologies.

[0131] The introduction of the U-Net structure significantly improved the model's segmentation capabilities. Experimental data showed that mIoU increased from 86.37% to 92.05%, Recall increased to 91.77%, Precision increased to 92.22%, and F1 Score increased from 85.33% to 91.99%, demonstrating that the U-Net structure effectively enhances the model's decoding capabilities. Its core advantage lies in its use of the Skip Connection mechanism, which fuses the encoder's early high-resolution features with the decoder's low-resolution features. This restores detailed information about crack regions while enhancing the model's spatial perception. The addition of the U-Net structure enhances the model's performance in crack edge identification, crack width differentiation, and connectivity preservation, significantly reducing the false detection rate.

[0132] On this basis, multi-scale feature fusion was introduced to further optimize the model's overall segmentation capabilities. Experimental data showed that mIoU increased from 92.05% to 92.18%, and Recall increased from 91.77% to 93.52%. Although Precision decreased slightly (90.82%), the F1 Score remained high (92.15%). The core function of multi-scale feature fusion is to integrate crack information at different scales, enabling the model to simultaneously identify small cracks and larger crack extension areas. The introduction of this module improves the model's detection capabilities for cracks of different sizes and enhances its adaptability to complex crack morphologies.

[0133] Ultimately, the addition of the attention mechanism further optimized the model's performance, achieving mIoU of 92.25%, recall of 93.08%, precision of 91.38%, and an F1 score of 92.22%. The attention mechanism's primary contribution lies in enhancing the model's focus on crack regions while reducing the impact of background interference. In crack identification tasks, traditional CNNs can be affected by irrelevant information due to interference factors such as noise, texture variations, and non-structural cracks on the concrete surface. The attention mechanism automatically assigns higher weights to crack regions, thereby improving the model's reliability. Furthermore, the attention mechanism plays a key role in the detection and segmentation of crack edges, making the crack morphology more complete and enhancing the model's applicability in practical engineering environments.

[0134] In summary, ablation experiments fully demonstrate the importance of the U-Net structure, multi-scale feature fusion, and attention mechanism in the semantic segmentation task of this invention. The U-Net structure significantly improves the overall crack segmentation capability, multi-scale feature fusion enhances the model's adaptability to cracks of different scales, and the attention mechanism optimizes the model's focus on crack regions, reducing false detections and improving overall segmentation accuracy.

[0135] The present invention verifies the effectiveness of each module through ablation experiments and achieves excellent performance on a self-built database without data enhancement, providing reliable technical support for high-precision crack identification, and at the same time providing a more practical technical solution for future automatic detection of concrete cracks and health monitoring of engineering structures.

[0136] Figure 11The performance changes of the present invention before and after data augmentation on a self-built dataset are demonstrated. It can be seen that after data augmentation, all evaluation indicators are improved, indicating the positive impact of the data augmentation strategy on the crack segmentation task. First, in terms of PA (pixel accuracy), it was 99.01% before data augmentation and slightly increased to 99.46% after augmentation, indicating that the overall segmentation accuracy remains high and data augmentation has little effect on PA. MPA (mean pixel accuracy) increased from 96.23% to 97.16%, indicating that the enhanced model has a more balanced classification of different categories (cracks and background). mIoU (mean intersection over union) increased from 92.25% to 94.50%, indicating that the enhanced model has a higher degree of overlap between the crack area and the true annotated area, and the segmentation is more accurate. Recall (recall rate) increased from 93.08% to 94.62%, indicating that the model's ability to detect cracks has been enhanced, reducing missed detections. Precision also increased from 91.38% to 94.37%, indicating more accurate crack detection and reduced false positives. Finally, F1Score increased from 92.22% to 94.49%, demonstrating that data augmentation improved the balance between recall and precision, resulting in more stable overall segmentation results. This is shown in Table 1.

[0137] Table 1 Optimal model results

[0138] After data enhancement, the SegB2-Unet of the present invention achieves optimal segmentation performance on a self-built ECC crack dataset, and all indicators are close to the theoretical optimal values, indicating that the model performs excellently in accuracy, robustness, and generalization ability.

[0139] First, the model achieved a pixel accuracy (PA) of 99.98%, indicating that across the entire dataset, the model accurately classified cracks and background with near-perfect accuracy, with minimal misclassification. This means the model can provide highly reliable segmentation results across a wide range of crack detection tasks. Furthermore, the mean pixel accuracy (MPA) was 99.75%, demonstrating that the model's classification accuracy across different classes (cracks and background) was very balanced, minimizing the impact of class imbalance on model performance.

[0140] In terms of crack segmentation accuracy, the mean intersection over union (mIoU) reached 99.04%, indicating that the model's predicted crack regions overlap highly with the actual crack regions, accurately capturing the crack's morphology and boundary details. Furthermore, the recall rate was 99.52%, indicating that nearly all cracks were detected, with an extremely low miss rate. This is crucial for crack monitoring tasks, ensuring that potential structural risks are not missed. Furthermore, the precision rate reached 99.23%, demonstrating that while maintaining a high recall rate, the model also effectively reduces false detections, meaning that the probability of background areas being misclassified as cracks is extremely low.

[0141] Overall, the F1 Score reached 99.37%, demonstrating that the model achieves an optimal balance between completeness and accuracy in crack identification, ensuring comprehensive crack detection while avoiding false detections. This result demonstrates that our SegB2-Unet, trained and optimized after data augmentation, achieves extremely high detection capabilities and can be stably and efficiently applied to crack monitoring tasks in practical engineering applications.

[0142] Data augmentation significantly improves model performance. During training, through data augmentation strategies such as rotation and flipping, the model learns crack characteristics in different orientations and shapes, improving its adaptability. Furthermore, data augmentation optimizes the model's crack edge detection capabilities, enabling it to better segment crack boundaries and reduce edge blur, thereby improving mIoU and F1 scores. Furthermore, data augmentation increases the diversity of crack data, giving the model stronger discrimination capabilities in complex background environments, avoiding false background detections, improving precision, and ensuring high reliability of detection results.

[0143] In summary, the data-enhanced SegB2-Unet achieved excellent performance on a self-built ECC crack dataset, with all key metrics approaching 99% or higher, demonstrating extremely high segmentation accuracy and stability. This result demonstrates the strong engineering applicability of SegB2-Unet and its widespread application in concrete structural health monitoring tasks, such as automated crack detection and maintenance in infrastructure such as bridges, tunnels, and roads. This provides reliable technical support for future civil engineering structural safety monitoring.

[0144] In this embodiment, the present invention accurately segments crack regions and analyzes crack length and width. This analysis is based on crack semantic segmentation. By extracting a binary mask from the crack segmentation, combined with pixel scaling and morphological analysis methods, the length and width distribution of the crack is calculated.

[0145] To analyze crack length, this invention employs an automated method based on image skeleton extraction for quantitative crack analysis. As shown in the figure, the model's predicted image is first subjected to a morphological thinning algorithm to extract the crack skeleton. Subsequently, bifurcation points in the skeleton image are identified and removed by detecting the number of nonzero pixels within a neighborhood, ensuring that the crack path remains a single connected path for subsequent analysis. Connected domain segmentation is then used to partition the skeleton into regions, and a modified Dijkstra algorithm is used to calculate the true crack length of each connected region, taking into account diagonal distances. To convert pixel lengths to actual physical dimensions, a pre-calibrated scale factor (pixels / mm) is used for conversion. The lengths of each crack segment are then statistically analyzed and averaged.

[0146] In terms of crack width analysis, the present invention uses digital image processing technology to quantitatively analyze the crack width of ECC specimens. Figure 12 As shown in the figure, the model's predicted image is first skeletonized to extract the crack centerline. A 3×3 convolution is then performed to detect bifurcation points within the skeleton, thereby segmenting the continuous crack into multiple independent crack segments. The crack segments are then cropped based on their minimum bounding rectangle within a local region, and a distance transform is calculated on the cropped local binary image. Distance values ​​are then extracted only for the local skeleton region, averaged within a preset half-width range, and multiplied by two to obtain the average crack width. To eliminate noise interference, only valid crack segments with an area greater than a preset threshold are retained for width calculation.

[0147] like Figure 13 、 Figure 14 As shown in Figure 2, experimental results demonstrate that the model maintains high-precision crack parameter extraction performance and good stability across a wide range of crack morphologies and background complexities. In particular, the model effectively avoids false and missed detections of fine cracks and branching cracks, thereby ensuring the reliability of the identification results. Furthermore, by optimizing morphological processing and regional analysis algorithms, the model further improves the accuracy of crack feature extraction, ensuring that automated crack measurement results meet practical engineering requirements.

[0148] Overall, the method of the present invention performs well in crack feature extraction and size calculation statistics, and can provide high-quality crack parameter data for subsequent crack development assessment and structural health monitoring, providing strong support for the safety assessment and maintenance of concrete structures.

[0149] This paper proposes a deep learning model (SegB2-Unet) based on the combination of Segformer-B2 and U-Net for high-precision semantic segmentation of uniaxial tension cracks in high-ductility concrete (ECC). This model combines the global feature extraction capabilities of the Transformer architecture with the high-resolution feature recovery capabilities of the U-Net architecture. Combined with multi-scale feature fusion and an attention mechanism, it effectively improves the accuracy and robustness of crack identification.

[0150] Experimental results demonstrate that SegB2-Unet achieves superior performance on a self-constructed ECC crack dataset without data augmentation. It demonstrates high accuracy in terms of mIoU (92.25%), Recall (93.08%), Precision (91.38%), and F1 Score (92.22%), demonstrating that the method can accurately identify cracks in complex backgrounds and exhibits strong generalization capabilities. Furthermore, SegB2-Unet demonstrates stable performance on public datasets, with experimental results for mIoU (85.8%-85.6%), Recall (84.6%-80.9%), and Precision (85.2%-89.3%) further validating the model's robustness and adaptability.

[0151] Ablation experiments further validated the crucial role of the U-Net architecture, multi-scale feature fusion, and the attention mechanism in crack identification. The experiments showed that when using Segformer-B2 alone for crack identification, the mean Intersection Over Union (MIoU) was only 72.83%. However, by gradually introducing the U-Net architecture, multi-scale feature fusion, and attention mechanism, the mIoU ultimately increased to 92.25%, demonstrating the effectiveness of each module. The U-Net architecture significantly improved the ability to recover crack edge details, multi-scale feature fusion enhanced the detection of cracks of varying sizes, and the attention mechanism optimized the model's focus on crack regions, improving overall segmentation accuracy.

[0152] Furthermore, the proposed method not only achieves high-precision crack segmentation but also further identifies crack length and width, providing more accurate parameter information for structural health monitoring. Experimental results demonstrate that the model maintains high-precision crack feature extraction capabilities across diverse crack morphologies and background complexities, demonstrating excellent stability in engineering applications.

[0153] Overall, the SegB2-Unet model proposed in this paper overcomes the limitations of traditional crack identification methods, excelling in complex crack morphology, fine crack identification, and efficient computation. This method not only improves the automation level of crack identification but also provides an efficient and accurate solution for crack segmentation and parameter identification for concrete structure health monitoring. In the future, this model can be further optimized to accommodate more engineering scenarios and practical application needs, such as crack identification and maintenance in infrastructure such as bridges, tunnels, and roads, thereby providing more reliable technical support for the safety of civil engineering structures.

[0154] Example 2

[0155] The present invention also provides a system for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method. The system is used in the above method, and the system includes: an acquisition module, a preprocessing module, a construction module, and an identification module;

[0156] An acquisition module is used to construct an image feature dataset based on features of high-toughness concrete crack image samples;

[0157] A preprocessing module, used to preprocess the image feature dataset to obtain a preprocessed dataset;

[0158] A construction module is used to construct a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and train the SegB2-Unet neural network using the pre-processed dataset to obtain a semantic segmentation model of high-toughness concrete cracks;

[0159] The recognition module is used to identify the high-toughness concrete tensile crack based on the high-toughness concrete crack semantic segmentation model, extract the length and width of the crack, and obtain crack morphology information.

[0160] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on deep learning methods, characterized in that: The method comprises: An image feature dataset is constructed based on the features of high-toughness concrete crack image samples; Preprocessing the image feature dataset to obtain a preprocessed dataset; Constructing a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and training the SegB2-Unet neural network using the pre-processed dataset to obtain a semantic segmentation model of high-toughness concrete cracks; Identifying high-toughness concrete tensile cracks based on the high-toughness concrete crack semantic segmentation model, extracting the length and width of the cracks, and obtaining crack morphology information; The method for constructing a SegB2-Unet neural network based on the pre-trained Segformer-B2 model and the improved U-net network includes: using Segformer-B2 as the encoder, combining it with the U-Net decoder, and integrating multi-scale feature fusion and attention mechanisms. The overall network structure is divided into an encoding part, a bottleneck part, and a decoding part; In the encoder part, Segformer-B2 is used as the feature extraction network. This network first processes the input image through overlapping patch embedding, which divides the image into several patches with overlapping areas. The image then enters a four-layer Transformer block. Each layer contains an efficient self-attention mechanism and a hybrid feed-forward network to capture crack features at different scales. Bottleneck layer: Located between the encoder and decoder, the bottleneck layer is used to further extract high-level features and reduce redundant information. This layer uses 3×3 convolution for feature extraction, combined with batch normalization and ReLU activation functions to ensure effective information propagation and alleviate the vanishing gradient problem. In addition, the bottleneck layer also has a channel adjustment function. The number of channels is expanded from 512 output by the Transformer block to 1024, and deep-level features are convolved to extract high-level features before returning to 512 channels. Decoder: The decoder uses a multi-stage upsampling strategy to gradually restore the spatial resolution of the feature map. In each stage, a transposed convolution module is first used for preliminary upsampling, and the feature map size is precisely adjusted through bilinear interpolation. To compensate for the spatial details that may be lost during information transmission in deep networks, each stage fully utilizes skip connections to directly transfer high-resolution features from the encoder to the corresponding decoding layer stage, enhancing the ability to capture crack edges and other subtle details. Each skip connection is followed by a convolutional fusion layer with 3×3 convolution, BatchNorm, and ReLU to further fuse the upsampled features with the features transferred from the encoder.

2. The method according to claim 1, characterized in that The characteristics of high-toughness concrete crack image samples include: input features and output features; The input features include: experimentally produced ECC crack images and manually annotated ECC crack images; The output features include: crack semantic segmentation image, and numerical information of crack length and width.

3. The method according to claim 1, characterized in that The method of preprocessing the image feature dataset to obtain the preprocessed dataset includes: Performing maximum and minimum normalization processing on the image feature dataset to obtain a normalized image dataset; The normalized image data set is divided based on a random division method to obtain the preprocessed data set.

4. A system for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on a deep learning method, the system being used to implement the method according to any one of claims 1 to 3, characterized in that: The system includes: an acquisition module, a pre-processing module, a construction module and a recognition module; The acquisition module is used to construct an image feature data set based on the features of the high-toughness concrete crack image samples; The preprocessing module is used to preprocess the image feature dataset to obtain a preprocessed dataset; The construction module is used to construct a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and train the SegB2-Unet neural network using the pre-processed data set to obtain a high-toughness concrete crack semantic segmentation model; The recognition module is used to identify high-toughness concrete tensile cracks based on the high-toughness concrete crack semantic segmentation model, extract the length and width of the cracks, and obtain crack morphology information.

Citation Information

Patent Citations

  • Road crack detection and segmentation method based on improved U-Net model

    CN119810445A