High-toughness concrete multi-seam behavior recognition and characterization method and system based on deep learning method
By combining Segformer-B2 and improved U-Net network deep learning methods, the problems of low efficiency and low accuracy in high toughness concrete crack recognition and quantitative analysis are solved, and efficient and accurate crack recognition and parameter analysis are achieved.
Patent Information
- Application Number
- CN202510535529.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The prior art has problems such as low recognition efficiency, low accuracy and high image quality in crack recognition and quantitative analysis of high toughness concrete, especially in complex background environments and diverse crack forms, which are difficult to achieve accurate recognition.
A deep learning-based method is adopted, combined with Segformer-B2 and improved U-Net network to build the SegB2-Unet neural network, and train the neural network through preprocessing data sets to realize semantic segmentation and quantitative analysis of high-strength concrete cracks.
It realizes efficient and accurate high-toughness concrete crack identification and quantitative analysis, improves identification efficiency and accuracy, can accurately identify cracks of different scales in complex backgrounds, and provides accurate parameter information of crack length and width.
Smart Images

Figure CN120070833A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision and high-performance concrete mechanical testing, and particularly relates to a method and system for identifying and characterizing the multi-crack development behavior of high-toughness concrete based on deep learning methods. Background Art
[0002] Engineered Cementitious Composite (ECC) is a high-performance fiber-reinforced cementitious composite, which has received much attention due to its unique pseudo strain-hardening behavior and excellent multiple cracking tensile properties, and has gradually occupied an important position in modern civil engineering. Compared with ordinary concrete, ECC has obvious advantages in many structural applications because it has high tensile strain capacity and multiple cracking behavior, and has a unique crack control ability, that is, under tensile stress, it will not produce a single wide crack, but multiple micro-cracks will appear, and the strain can reach 3-8% under tension, and the crack width can be controlled to <100μm before final failure. This multi-crack characteristic enhances the ductility of the material, effectively extends the service life of the structure, and reduces the impact of crack propagation on the structural integrity.
[0003] With the continuous in-depth research on ECC materials and the continuous improvement of preparation technologies, the cost of ECC has gradually decreased and its performance has been continuously improved. At present, ECC has been widely used in the repair and reinforcement of concrete beam structures and the construction of new engineering structures. Although ECC materials have excellent crack control ability, the appearance and propagation of cracks are still the key factors affecting their structural performance. The formation, distribution, width, and propagation of cracks have a direct impact on the performance of ECC material structures, especially in terms of corrosion, penetration, and durability. If these cracks cannot be identified and repaired in time, they may further propagate, resulting in structural damage. Therefore, the early detection and monitoring of ECC cracks are crucial, which can help to timely discover potential structural problems and carry out repairs, thereby avoiding more serious damage.
[0004] The traditional crack detection methods used in the early stage mainly relied on manual detection, which was not only time-consuming and laborious, but also the results were easily affected by the inspector's experience, making it difficult to achieve precise detection. Therefore, both the efficiency and accuracy were relatively low, and problems such as data loss or errors were likely to occur.
[0005] Subsequently, some non-destructive testing methods emerged, such as ultrasonic, infrared thermography, and electrical sensing. Although ultrasonic and infrared thermography can provide more detailed crack information, they require expensive equipment, complex operations, and have the drawback of a small measurement range. With the development of computer technology, image processing technology based on computer vision technology has emerged. People collect images through equipment and then use a computer to preprocess the crack images, such as edge detection, thresholding, clustering, binarization, histogram equalization, and denoising techniques, to achieve image segmentation and image pixel detection. However, the crack recognition technology using image processing technology has certain limitations. For example, this technology has high requirements for image quality. If there are problems such as image blurring, noise, or uneven illumination, it may lead to inaccurate crack recognition. In addition, the complex background environment and the diversity of crack morphologies also increase the difficulty of recognition, making it difficult for the algorithm to take into account all situations, and some images still require manual recognition and other issues.
[0006] In recent years, with the rapid development of deep learning, the application of deep learning algorithms in concrete has shown great potential, and more and more scholars have begun to apply deep learning technology to aspects such as the health monitoring of concrete structures. Among them, the detection and recognition of concrete cracks is exactly an important research direction. For the detection and recognition of ECC cracks, deep learning technology learns the rules and features from samples through a multi-layer neural network, so as to achieve the recognition and positioning of objects in the image. By training a deep neural network model, the ECC crack recognition method based on deep learning can achieve automated crack detection, with the characteristics of high efficiency, accuracy, and strong stability, while avoiding the interference of human subjective factors. This not only greatly improves the efficiency and accuracy of concrete structure detection but also helps to effectively ensure the safety of the structure and extend its service life.
[0007] In existing research, from the perspective of methods, the current semantic segmentation models for ECC and other concrete cracks still have the following deficiencies: (1) The receptive fields of encoders such as traditional convolutional neural networks (such as U-Net) are limited, making it difficult to capture the long-range dependencies and multi-scale features of ECC cracks; (2) For the machine learning models currently applied in this field, the existing models still have insufficient feature extraction and limited segmentation accuracy when dealing with fine cracks, and also consume a large amount of time, which is not conducive to the application of the model in actual engineering; from the perspective of research objects, there are few relevant studies on the recognition of ECC uniaxial tensile cracks based on deep learning methods currently. Existing research often focuses on the recognition of ordinary concrete cracks and lacks relevant research on the recognition of ECC tensile cracks. Especially for ECC cracks, the cracking morphology and cracking mode of ECC cracks are different from those of ordinary concrete, and the durability and permeability of the ECC structure depend on the crack width. Therefore, crack recognition is carried out to provide parameter information about the crack length and width in ECC tension. Summary of the Invention
[0008] To solve the problems existing in the prior art, the present invention provides a method and system for identifying and characterizing the multi-crack development behavior of high-toughness concrete based on deep learning methods, which is applicable to the accurate identification and quantitative analysis of cracks in the uniaxial tensile test of high-toughness concrete.
[0009] To achieve the above object, the present invention provides the following solutions:
[0010] A method for identifying and characterizing the multi-crack development behavior of high-toughness concrete based on deep learning methods, the method comprising:
[0011] Constructing an image feature dataset based on the features of high-toughness concrete crack image samples;
[0012] Preprocessing the image feature dataset to obtain a preprocessed dataset;
[0013] Constructing a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and training the SegB2-Unet neural network with the preprocessed dataset to obtain a high-toughness concrete crack semantic segmentation model;
[0014] Identifying the tensile cracks of high-toughness concrete based on the high-toughness concrete crack semantic segmentation model, and extracting the length and width of the cracks to obtain crack morphology information.
[0015] Preferably, the features of the high-toughness concrete crack image samples include: input features and output features;
[0016] Among them, the input features include: the ECC crack image made by experiments, the ECC crack image manually annotated;
[0017] The output features include: crack semantic segmentation images, numerical information on the length and width of the cracks.
[0018] Preferably, the method for preprocessing the image feature dataset to obtain a preprocessed dataset includes:
[0019] Performing maximum-minimum normalization processing on the image feature dataset to obtain a normalized image dataset;
[0020] Dividing the normalized dataset based on a random division method to obtain the preprocessed dataset.
[0021] Preferably, the method for constructing the SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network includes: using Segformer-B2 as the encoder, combining with the U-Net decoder, and integrating multi-scale feature fusion and attention mechanism. The overall network structure is divided into an encoding part, a bottleneck part, and a decoding part.
[0022] Preferably, in the encoder part, Segformer-B2 is used as the feature extraction network. This network first processes the input image through overlapping Patch embedding, so that the image is segmented into several Patches with overlapping regions. Subsequently, the image enters four layers of Transformer blocks, each layer containing an efficient self-attention mechanism and a hybrid feed-forward network to capture crack features at different scales.
[0023] Preferably, the bottleneck part is located between the encoder and the decoder and is used to further extract high-level features. This part uses 3×3 convolution for feature refinement, combined with batch normalization and ReLU activation function. The bottleneck part also has the function of channel adjustment, so that the deep features of the encoder are adapted to the input requirements of the decoder.
[0024] Preferably, in the decoder part, the U-Net structure is used to gradually restore the resolution of the feature map to generate the crack segmentation result. The decoder consists of multiple upsampling modules, feature fusion modules, and convolution operations, namely Conv3×3 + BatchNorm + ReLU. Skip connections are introduced in the decoder, so that the high-resolution features of different scales in the encoder can be directly transmitted to the corresponding decoding layers. An attention module based on the CBAM idea is introduced in the decoder. Specifically, this module first uses the channel attention mechanism to extract global information through adaptive average pooling and max pooling, and calculates the channel attention coefficient through a fully connected layer to weight the input features and enhance the model's attention to important channels. Subsequently, the spatial attention mechanism is used to fuse the spatial information of global average pooling and max pooling, and calculate the spatial attention coefficient through a convolutional layer and a Sigmoid activation function to improve the model's sensitivity to the crack region.
[0025] The present invention also provides a high-toughness concrete multi-crack development behavior recognition and characterization system based on a deep learning method. The system is used for the aforementioned method. The system includes: an acquisition module, a preprocessing module, a construction module, and an identification module;
[0026] The acquisition module is used to construct an image feature dataset based on the features of high-toughness concrete crack image samples;
[0027] The preprocessing module is used to preprocess the image feature dataset to obtain a preprocessed dataset;
[0028] The building block is used to construct the SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and train the SegB2-Unet neural network through the preprocessed data set to obtain a high-toughness concrete crack semantic segmentation model;
[0029] The recognition module is used to recognize the tensile cracks of high-toughness concrete based on the high-toughness concrete crack semantic segmentation model, extract the length and width of the cracks, and obtain crack morphology information.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] The present invention proposes an end-to-end ECC tensile crack feature recognition model based on deep learning and a multi-scale feature fusion module. First, a deep learning model combining the Segformer-b2 model and U-net is used to replace the traditional single-network deep learning model, improving the network's feature mining ability and semantic segmentation ability. Second, the multi-scale feature fusion and attention mechanism technologies are used to further enhance the network's feature segmentation ability, that is, reducing the time consumption compared with the original single-network model and further improving the segmentation accuracy. The present invention proves that the proposed SegB2-Unet method has good recognition ability for cracks generated during the ECC tensile test through 466 image samples collected from experiments. The corresponding evaluation indicators PA, MPA, MIoU, Recall, Precision, and F1 Score are 99.46, 97.16, 94.50, 94.62, 94.37, and 94.41 respectively. This method is superior to neural network deep learning methods such as single U-net and Segformer-B2. In summary, the SegB2-Unet method proposed by the present invention realizes high-precision semantic segmentation of ECC cracks and automatic measurement of crack length and width. This method can be widely applied to crack detection and maintenance of infrastructure such as bridges, tunnels, and roads, providing an efficient and reliable technical solution for structural health monitoring in the field of civil engineering. Description of the Drawings
[0032] In order to more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic diagram of the specimen size of the embodiment of the present invention;
[0034] Figure 2Schematic diagram of the ECC crack image and the manually labeled ECC crack image made for the experiments of the embodiments of the present invention;
[0035] Figure 3 Schematic diagram of the model network architecture of the embodiments of the present invention;
[0036] Figure 4 Schematic diagram of the actual image and the prediction result of the present invention in the embodiments of the present invention;
[0037] Figure 5 Schematic diagram of the model comparison in the embodiments of the present invention;
[0038] Figure 6 Schematic diagram of the comparison of the model generalization performance effect in the embodiments of the present invention;
[0039] Figure 7 Schematic diagram of the performance of the model of the embodiments of the present invention on the Hao et al. dataset;
[0040] Figure 8 Schematic diagram of the performance of the model of the embodiments of the present invention on the Deepcrack dataset;
[0041] Figure 9 Schematic diagram of the performance index of the ablation experiment in the embodiments of the present invention;
[0042] Figure 10 Schematic diagram of the ablation experiment result in the embodiments of the present invention;
[0043] Figure 11 Schematic diagram of the comparison before and after data augmentation in the embodiments of the present invention;
[0044] Figure 12 Analysis process of the crack length in the embodiments of the present invention, where (a) is the original crack; (b) is the crack segment skeleton; (c) is the length recognition result;
[0045] Figure 13 Analysis process of the crack width in the embodiments of the present invention, where (a) is the original crack; (b) is the crack segment skeleton; (c) is the width calculation;
[0046] Figure 14 Analysis of the local crack width in the embodiments of the present invention, where (a) is the local crack; (b) is the local crack segment skeleton; (c) is the local crack width;
[0047] Figure 15 Schematic diagram of the flow of a method for identifying and characterizing the multi-crack development behavior of high-toughness concrete based on a deep learning method in the embodiments of the present invention. Detailed implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0050] Embodiment 1
[0051] As Figure 15 shown, the present invention provides a method for identifying and characterizing the multi-crack development behavior of high-toughness concrete based on a deep learning method. The method includes:
[0052] Constructing an image feature dataset based on the features of high-toughness concrete crack image samples;
[0053] Preprocessing the image feature dataset to obtain a preprocessed dataset;
[0054] Constructing a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and training the SegB2-Unet neural network through the preprocessed dataset to obtain a high-toughness concrete crack semantic segmentation model;
[0055] Identifying the tensile cracks of high-toughness concrete based on the high-toughness concrete crack semantic segmentation model, and extracting the length and width of the cracks to obtain crack morphology information.
[0056] In this embodiment, the present invention proposes a deep learning model (SegB2-Unet) based on the combination of Segformer-B2 and U-Net for high-precision semantic segmentation of uniaxial tensile cracks in Engineered Cementitious Composites (ECC). This method combines the global feature extraction ability of Transformer and the high-resolution feature restoration ability of the U-Net structure, and integrates multi-scale feature fusion and attention mechanism to effectively improve the ability to extract crack features, achieving accurate segmentation of cracks and precise identification of their length and width. In addition, while ensuring the recognition accuracy, this method optimizes the computational efficiency and reduces the calculation time, enabling it to be applicable to large-scale crack recognition tasks. The data processing, training, and analysis of the present invention are based on the integrated development environment of PyCharm 2021, and programming is carried out using Python 3.11. PyTorch 2.4.1 is selected as the deep learning framework. The GPU and CPU of the workstation used are NVIDIA GeForce RTX 3060 Ti and 13th Gen Intel(R) Core(TM) i5-13400F@2.50 GHz respectively. CUDA 12.4 is used as the parallel computing platform to accelerate training.
[0057] The specific implementation steps are as follows:
[0058] In this embodiment, the present invention proposes a deep learning model (SegB2-Unet) that combines Segformer-B2 and U-Net for semantic segmentation of uniaxial tensile cracks in Engineered Cementitious Composites (ECC). This model combines the global feature extraction ability of Transformer and the fine segmentation ability of Convolutional Neural Network (CNN) to achieve high-precision recognition of complex crack patterns, and further extracts the length and width of the cracks to support crack morphology analysis and structural health monitoring.
[0059] In this embodiment, the principle of Segformer-B2: Segformer is a semantic segmentation framework that combines Transformer and a lightweight multi-layer perceptron (MLP) decoder. Its encoder part adopts a hierarchical Transformer design, which can extract multi-scale features and avoid using positional encoding, thus maintaining stable performance when the test resolution is inconsistent with the training resolution. This design ensures that Segformer has powerful feature representation ability while performing efficient calculations, enabling it to extract global and local feature information in crack recognition tasks, improving the segmentation accuracy and generalization ability.
[0060] Principle of U-Net: U-Net is a classic CNN structure widely used in semantic segmentation tasks. Its structure consists of a symmetric encoder (downsampling path) and decoder (upsampling path), forming a typical U-shaped architecture. By introducing skip connections between the encoder and decoder, U-Net can effectively fuse features at different scales and enhance the ability to recover high-resolution spatial information. This design enables U-Net to accurately extract crack edges in crack semantic segmentation tasks, ensuring the detection integrity of fine cracks and reducing false detections or missed detections caused by resolution loss.
[0061] Design of SegB2-Unet Model: In the SegB2-Unet structure, the encoder part uses Segformer-B2 to utilize its powerful global feature extraction ability to ensure effective identification of cracks at different scales. The decoder part uses an improved U-Net structure, combined with multi-scale feature fusion and attention mechanism, to further enhance the model's ability to identify fine cracks. This design gives full play to the advantages of Transformer in global modeling, while combining the local detail recovery ability of the CNN structure, enabling the model to maintain high-precision crack segmentation even in complex background environments and further extract the geometric parameters (length, width) of the cracks to support crack morphology analysis and engineering applications.
[0062] Attention Mechanism: In the SegB2-Unet model proposed in this invention, the attention mechanism is introduced into the decoder part to improve the detection ability of the crack area, reduce background interference, and improve the accuracy of crack segmentation. Cracks often exhibit characteristics such as slender, non-uniform, and low contrast, making it easy for traditional CNNs to miss important information during feature extraction. Through the attention mechanism, the model can dynamically enhance the expression of crack features, suppress irrelevant background information, and optimize the crack segmentation effect. This invention adopts the CBAM (Convolutional Block Attention Module) mechanism, which adaptively weights the input features in two consecutive steps:
[0063] First is channel attention, and then spatial attention. In the channel attention stage, the module generates 1×1 feature descriptors in the spatial dimension through adaptive average pooling and adaptive max pooling respectively, and these two pooling methods capture the global average information and the most significant activations of the input features respectively.
[0064] Next, these two descriptors will go through a shared fully-connected network, which includes a dimensionality reduction layer (using linear transformation to reduce the number of channels to 1 / reduction of the original number of channels), a ReLU activation function, and a linear layer to restore the number of channels, in order to extract the dependencies between channels. The outputs of the two fully-connected layers are added together and normalized to the interval (0, 1) through Sigmoid activation to form fine-grained channel attention weights, and then the original features are weighted channel by channel to highlight the key information.
[0065] Subsequently, in the spatial attention stage, for the feature map weighted by channels, the module calculates the average value and the maximum value along the channel dimension respectively to generate two spatial descriptor maps, which reflect the overall activation level and the most significant response of the local area respectively. After concatenating these two descriptor maps along the channel dimension, they are fused through a convolutional layer (kernel size is 7 to ensure a sufficient receptive field), and finally a spatial attention map is generated through Sigmoid activation.
[0066] Finally, the feature map is weighted pixel by pixel using the spatial attention map, effectively enhancing the features of the key regions while suppressing the background noise. This dual attention mechanism can help the model capture subtle and discontinuous features more accurately when dealing with tasks such as crack detection, significantly improving the overall segmentation accuracy.
[0067] Through the spatial attention mechanism and the channel attention mechanism, the present invention further optimizes the performance of SegB2-Unet in ECC crack recognition, enabling it to achieve efficient and accurate semantic segmentation in complex crack environments, providing reliable technical support for the health monitoring of concrete structures.
[0068] Multi-scale feature fusion: The morphology and size of cracks vary greatly, and the widths, lengths, and depths of different cracks may show obvious differences in the same image. It is difficult to obtain robust detection results by relying only on single-scale features for segmentation. Therefore, the present invention introduces a multi-scale feature fusion (Multi-Scale Feature Fusion, MSFF) strategy in SegB2-Unet to simultaneously utilize local detail features and global semantic information to improve the detection accuracy of cracks.
[0069] The present invention uses a Segformer-B2 encoder to extract multi-scale features. The model extracts feature maps with different resolutions at different Transformer layers, and these feature maps respectively contain high-resolution local detail information (crack edges) and low-resolution global structural information (overall crack morphology).
[0070] To make full use of multi-scale information, during the decoding process, the model combines the Skip Connection with the Two-stage Up-Sampling mechanism to fully integrate the low-resolution global features and high-resolution detailed information, thereby enhancing the reconstruction ability of the crack area. Specifically, in each decoding stage, initial upsampling is first performed through transposed convolution, and then the size of the feature map is precisely adjusted through Bilinear Interpolation to strictly align it with the corresponding encoder features. Subsequently, concatenation and feature fusion are carried out. In addition, the combination of the Channel and Spatial Attention (CBAM) mechanism helps to enhance the feature expression of the crack area, while suppressing background noise and improving the robustness of crack detection. The experimental results show that this decoding strategy can effectively improve the detection ability for cracks of different scales, avoid the problem that a single-scale model performs poorly on cracks of a specific size, while enhancing the integrity of the crack morphology, ensuring that the segmentation result covers the entire crack area, reducing the situation of crack fracture or omission, and thus improving the crack recognition ability of the model. The experimental results of multi-scale feature fusion show that: it improves the detection ability for cracks of different scales and avoids the problem that a single-scale model performs poorly at certain crack sizes; it strengthens the integrity of the crack morphology, ensures that the detection result covers the entire crack area, and avoids crack fracture or omission; it improves the generalization ability of the model, adapts to different types of cracks and complex background environments, and improves the applicability in real engineering environments.
[0071] The SegB2-Unet of the present invention effectively improves the accuracy and robustness of crack recognition through the multi-scale feature fusion mechanism, combining the Skip Connection and the upsampling strategy. This method can adapt to cracks of different sizes and morphologies, making the SegB2-Unet have broad application prospects in the field of ECC structural health monitoring and being able to meet the automated crack recognition needs of infrastructure such as bridges, roads, and tunnels.
[0072] In this embodiment, the model of the present invention is trained and tested using public datasets and experimental datasets respectively to comprehensively evaluate the detection performance and generalization ability of the model in different crack environments. In the public dataset, we selected a high-quality dataset containing various crack types to verify the adaptability of the model to different crack morphologies, background conditions, and size changes. The experimental dataset is obtained from laboratory tests and contains image data of uniaxial tensile cracks in Engineered Cementitious Composites (ECC) to further verify the reliability and stability of the model in real engineering application scenarios.
[0073] The cracks used in the experimental dataset were generated through uniaxial tensile tests on dog-bone specimens prepared from ECC materials. The tests were conducted in a controlled experimental environment, and the cracks gradually expanded under standard loading conditions, ensuring that the crack formation process conforms to the actual engineering situation, thereby enhancing the application adaptability and engineering value of the model.
[0074] Through systematic testing on the public dataset and the experimental dataset, the detection ability of the model of the present invention has been comprehensively analyzed under different environments, and is quantitatively evaluated by PA (Pixel Accuracy), mIoU (Mean Intersection over Union), Recall, Precision, and F1 Score to ensure the reliability, stability, and generalization ability of the model.
[0075] Public dataset: In this study, multiple public concrete crack datasets were selected to evaluate the generalization ability and robustness of the model under different crack environments. These datasets contain crack images with different backgrounds, different lighting conditions, and various crack morphologies, covering smooth surface cracks, rough surface cracks, shallow cracks, and deep cracks, which can effectively improve the adaptability of the model in different application scenarios. These datasets have been professionally annotated and are widely used in crack recognition research, providing high-quality supervised data for model training.
[0076] The model of the present invention directly uses the public dataset for training and testing without performing additional processing on the data to ensure the fairness and objectivity of the comparative experiment. All public datasets are independently used for model evaluation to verify the adaptability and generalization of the model on different data sources.
[0077] Experimental dataset: In the present invention, the experimental dataset consists of ECC concrete tensile test images actually obtained in the laboratory, mainly used to evaluate the performance of the model in real experimental crack recognition scenarios.
[0078] The experimental specimens were prepared from high-toughness concrete (ECC) materials, and their mix proportions include cement, fly ash, silica fume, sand, water, water reducer, and fibers. After 28 days of standardized curing, the specimens were formed. The specimen size refers to the JSCE standard specification, and uniaxial tensile tests were carried out on a universal testing machine at a loading rate of 0.5 mm / min. During the test, the formation and expansion process of the cracks were recorded by a high-resolution camera, and the cracks were annotated. Finally, a high-quality experimental dataset containing 466 samples was constructed.
[0079] The design of this experimental dataset ensures the authenticity of crack morphology, and its crack growth pattern is consistent with the crack propagation mechanism in actual engineering structures. Therefore, it has important value for the engineering adaptability evaluation of the model. The introduction of the experimental dataset enables the model to not only have good generalization ability on the public dataset but also maintain a high recognition accuracy in the real experimental crack recognition task. As Figure 1 、 Figure 2 shown.
[0080] In this embodiment, data preprocessing: To further improve the generalization ability and stability of the model, the experimental dataset was preprocessed before training to enhance the model's detection ability under different crack morphologies and complex backgrounds.
[0081] First, we used methods such as image rotation (e.g., 、 ), horizontal flipping, and vertical flipping to augment the data samples, enabling the model to adapt to the detection tasks of cracks in different directions and enhancing the recognition ability of non-uniform crack growth patterns.
[0082] In addition, to improve the training speed of the model and enhance its recognition ability for small-scale cracks, we divided a single sample into multiple small samples of 352×480 for training. This method not only reduces the GPU computing burden but also ensures that the model can learn finer-grained crack information, optimizing the crack edge detection and detail recovery capabilities.
[0083] In summary, the present invention adopts strict experimental design and data preprocessing methods to ensure the stability, robustness, and high computational efficiency of the model in different environments, enabling it to be applicable to complex crack morphology detection and structural health monitoring tasks.
[0084] In this embodiment, the model architecture is as Figure 3 shown: The proposed SegB2-Unet in the present invention uses Segformer-B2 as the encoder, combines with the U-Net decoder, and integrates multi-scale feature fusion and attention mechanism to enhance the ability of ECC crack recognition. The overall network structure can be divided into an encoding part (Encoder), a bottleneck part (Bottleneck), and a decoding part (Decoder).
[0085] Encoder Part: In the encoder part, Segformer-B2 is used as the feature extraction network. This network first processes the input image through Overlap Patch Embeddings, dividing the image into several patches with overlapping regions to reduce information loss. Subsequently, the image enters four layers of Transformer Blocks, each layer containing an Efficient Self-Attention mechanism and a Mix-FFN. This structure ensures that the model has strong global feature extraction capabilities and can capture crack features at different scales. During this process, the Transformer blocks at different levels output feature maps of different scales, with sizes of (64, H / 4, W / 4), (128, H / 8, W / 8), (320, H / 16, W / 16), and (512, H / 32, W / 32) respectively, to ensure the multi-scale perception ability of cracks.
[0086] Bottleneck Part: The bottleneck layer is located between the encoder and the decoder and is used to further extract high-level features and reduce redundant information. This part uses 3×3 convolution (Conv3×3) for feature refinement, combined with Batch Normalization and the ReLU activation function to ensure the effective propagation of information and alleviate the problem of gradient disappearance. In addition, the bottleneck layer also has a channel adjustment function. The number of channels is expanded from 512, the number of channels output by the Transformer block, to 1024, and high-level features are extracted through convolution of deep features, and then returned to 512 channels. This processing can not only integrate information at a deeper level, but also play a role in compressing and refining features, and provide more refined high-level semantic information for the decoder stage and adapt to the input requirements of the decoder.
[0087] Decoder Part: In the decoder part, the resolution of the feature map is restored by means of progressive upsampling to generate high-precision crack prediction results. The decoder consists of a two-stage upsampling module combining multiple transposed convolutional layers (ConvTranspose2d) with bilinear interpolation, as well as a convolutional fusion layer (Conv2d) and convolutional operations (Conv3×3 + BatchNorm + ReLU). Among them, transposed convolution is used for feature learning, while the interpolation operation ensures precise size matching to optimize the feature alignment effect of the skip connection. This structure helps to improve the smoothness and continuity of the prediction results while restoring the details of the crack boundary. The transposed convolutional layer module performs transposed convolution operations through a learnable convolutional kernel to achieve upsampling. It has a trainable convolutional kernel that can learn an upsampling method suitable for crack feature restoration. For the crack detection task, retaining texture details is crucial for edge recognition, so the learnable upsampling method is more conducive to crack boundary restoration. In the decoder part of this model, although the upsampling and skip connection strategy is still adopted, the overall design is not a simple symmetric structure, but is carefully designed to fully adapt to the features output by the MiT-B2 encoder at multiple scales. First, before entering the decoding stage, a bottleneck module processes the high-dimensional features of the last layer of the encoder, aiming to compress redundant information and extract more discriminative feature representations. This bottleneck module not only adjusts the number of channels but also effectively improves the compactness and discriminability of the features through the combination of convolution, batch normalization, and non-linear activation functions. Subsequently, the decoder adopts a multi-stage upsampling strategy to gradually restore the spatial resolution of the feature map. In each stage, first, the transposed convolution (ConvTranspose2d) module is used for preliminary upsampling, and the size of the feature map is precisely adjusted by bilinear interpolation. This not only increases the spatial size of the feature map but also realizes an effective transformation of the features through trainable parameters. Since the transposed convolution may not be able to precisely restore to the target size in some cases, bilinear interpolation is introduced to finely adjust the size of the upsampled feature map to strictly align it with the corresponding skip connection feature map. In this way, the transposed convolution and the interpolation operation are effectively combined, thus effectively alleviating the problem of size mismatch in the output of the MiT-B2 encoder at multiple scales and ensuring seamless integration of information. In addition, to make up for the spatial details that may be lost during information transmission in the deep network, each stage makes full use of the skip connection to directly transfer the high-resolution features in the encoder to the corresponding decoding layer stage, enhancing the ability to capture crack edges and other tiny details.After each skip connection, a convolutional fusion layer (3×3 convolution, BatchNorm, and ReLU) follows to further fuse the upsampled features with the features passed from the encoder. This convolutional fusion layer not only alleviates the matching problem between different semantic levels but also refines the reconstructed feature information, ensuring that the final output has good smoothness and continuity, thus achieving high-precision crack semantic segmentation. In addition, to further enhance the feature expression and the ability to capture edge details, an attention module based on the CBAM idea is introduced at the final stage of the decoder. After this module is applied to the output features of the decoder, it first weights the importance of different channels through the channel attention branch, automatically adjusting the response degree of each channel's features to highlight the key semantic information. Subsequently, through the spatial attention branch, it carefully analyzes the spatial distribution of the feature map, strengthens the regions with important local information (such as crack edges), and suppresses redundant or noisy information. This combined strategy can further optimize the feature fusion effect, improve the smoothness and continuity of the segmentation result, and enhance the recognition sensitivity to subtle crack features.
[0088] The CBAM (Convolutional Block Attention Module) attention mechanism consists of two parts: the channel attention mechanism and the spatial attention mechanism. The purpose of the channel attention is to learn the importance of different channels and enhance the feature response of important channels; the purpose of the spatial attention is to learn the importance of different spatial positions of the feature map and highlight key regions, such as cracks. The formula for the CBAM attention mechanism is as follows:
[0089] Among them, represents the Sigmoid function; is the final feature generated by the model and are the average value and the maximum value of channel pooling respectively; is the convolution operation with a convolution kernel size of ; represents the output of the channel attention mechanism; represents the output of the spatial attention mechanism.
[0090] Final Output: Finally, in the last stage of the decoder, the generation of the crack segmentation map is achieved through a 1×1 convolutional layer combined with the Sigmoid activation function. In this segmentation map, the pixel values corresponding to the crack regions are 1, while the pixel values of the background regions are 0, ensuring that the model can maintain a high-precision crack recognition ability at different scales. Generally speaking, the SegB2-Unet model combines the global feature extraction advantages of Segformer-B2 and the excellent performance of the U-Net structure in fine-grained segmentation. Through multi-scale feature fusion and the introduction of the CBAM attention mechanism, this model significantly improves the accuracy and robustness of crack recognition. Experimental results show that the model exhibits excellent performance in the ECC uniaxial tensile crack recognition task, demonstrating its application potential in the fields of crack automatic detection and structural health monitoring.
[0091] Among them, Skip Connections: In terms of Skip Connections, due to the symmetry of the U-Net structure, the high-resolution features extracted in the early stage of the encoder can be directly transmitted to the corresponding levels of the decoder through Skip Connections, effectively compensating for the problem of feature loss and improving the segmentation accuracy. In the crack recognition task, cracks often present slender or discontinuous forms, and ordinary deep networks may cause small cracks to be ignored or missegmented. Skip Connections can retain and transmit the local detail information of the shallow layer to the decoding stage, enabling the spatial information of the cracks to be retained. At the same time, this connection method can significantly improve the detection effect of crack edges, avoid problems such as edge blurring or incompleteness, improve the overall segmentation accuracy, and reduce information loss, making the final segmentation result more accurate.
[0092] Attention Module (Convolutional Block Attention Module): In terms of the attention module, the CBAM (Convolutional Block Attention Module) mechanism is adopted, which adaptively weights the input features in two consecutive steps: first, channel attention, and then spatial attention. In the channel attention stage, the module generates 1×1 feature descriptors in the spatial dimension through adaptive average pooling and adaptive max pooling respectively. These two pooling methods capture the global average information and the most significant activations of the input features respectively. Next, these two descriptors will pass through a shared fully connected network, which includes a dimensionality reduction layer (using a linear transformation to reduce the number of channels to 1 / reduction of the original number of channels), a ReLU activation function, and a linear layer to restore the number of channels, in order to extract the dependencies between channels. The outputs of the two fully connected layers are added together and normalized to the interval (0,1) through Sigmoid activation to form fine-grained channel attention weights, and then the original features are weighted channel by channel to highlight the key information. Subsequently, in the spatial attention stage, for the feature map weighted by channels, the module calculates the average value and the maximum value along the channel dimension respectively to generate two spatial descriptive maps, which reflect the overall activation level and the most significant response of the local region respectively. After concatenating these two descriptive maps in the channel dimension, they are fused through a convolutional layer (kernel size is 7 to ensure a sufficient receptive field), and finally, a spatial attention map is generated through Sigmoid activation. Finally, the feature map is weighted pixel by pixel using the spatial attention map, effectively enhancing the features in the key regions while suppressing the background noise. This dual attention mechanism can help the model capture subtle and discontinuous features more accurately when dealing with tasks such as crack detection, significantly improving the overall segmentation accuracy.
[0093] In summary, the SegB2-Unet proposed in this invention fully combines the advantages of Segformer-B2 and U-Net, enabling the model to have a powerful crack recognition ability. As the encoder, Segformer-B2 provides the ability to extract global features, can effectively capture long-range feature correlations, and enhance the recognition ability for complex crack morphologies. The decoding part of the U-Net structure retains the high-resolution spatial features through skip connections, ensuring that the detailed information of the crack area will not be lost, thereby improving the recognition accuracy of cracks. In addition, the multi-scale feature fusion and attention mechanism introduced in the model further optimize the crack recognition effect, enabling the model to focus on the crack area, reduce background interference, and improve the segmentation accuracy. The experimental results show that this structure demonstrates excellent performance in the ECC uniaxial tensile crack recognition task, can accurately segment the crack area, and provide reliable technical support for subsequent crack analysis, quantitative evaluation, and structural health monitoring.
[0094] In this embodiment, the present invention adopts an image skeleton extraction method and a distance transformation method respectively for extracting the crack length and width.
[0095] In terms of length extraction, first, the process of converting the binary segmentation map into a single-pixel-wide skeleton is carried out. The topological structure centerline of the image is retained through iterative erosion operations, thereby generating a single-pixel-wide skeleton. By observing the semantic segmentation results, it is found that most cracks have multiple bifurcation points, which increases the extraction difficulty. Therefore, a bifurcation point detection method is introduced to traverse the skeleton pixels and count the number of non-zero pixels in the 8-neighborhood of a certain pixel point . If a certain pixel itself is a skeleton point and the neighborhood statistical quantity ≥3, then this pixel point will be determined as a bifurcation point. The formula is as follows:
[0096]
[0097] where count represents the total number of statistics; represents the binary image matrix, which only has two cases of 0 and 1; represents the coordinates of the pixel point to be judged currently in the binary image matrix; represents the pixel value of the coordinate in the matrix; the pixel value of the pixel point to be judged currently of.
[0098] After the bifurcation point detection is completed, the pixel value of the bifurcation point is set to 0 to disconnect the connection, so that the complex crack is decomposed into multiple independent segments to avoid misjudging the cross cracks as a single crack during the subsequent connected component analysis. After skeleton extraction and bifurcation point processing, the cracks in the image are segmented into multiple independent parts, and the images of the obtained independent crack segments are subjected to connected component analysis and labeled. Subsequently, area filtering is performed on all the identified crack segments. If the length of a certain crack segment is 0 or less than 1 mm after conversion, the processing of this segment is skipped. This can exclude invalid crack segments, reduce the influence of noise and the subsequent calculation amount.
[0099] After completing the labeling of the independent crack segments, the length of the crack in the skeleton map is calculated. First, the skeleton map is modeled as a graph structure, the skeleton pixels are regarded as the nodes of the graph, and the distance between the nodes is calculated according to the adjacency relationship. The distance between horizontally or vertically adjacent nodes is 1 pixel, and the distance between diagonally adjacent nodes is pixels (about 1.414). The Dijkstra algorithm is used to find the longest path between the two end points in the skeleton map, and the length of this path is the length of the crack segment. Since this length is composed of pixels, we need to convert it into a physical length. It is roughly obtained through the correspondence between the image pixels and the physical dimensions in the present invention Pixels. Therefore, the physical length can be obtained by dividing the pixel length by 57.0, i.e.:
[0100] where, represents the physical length; represents the pixel length.
[0101] In terms of width extraction, similar to the calculation of length, the image needs to be skeletonized and the bifurcation points need to be processed. After the processing is completed, the segmented independent skeleton crack segments are corresponded to the cracks in the corresponding positions of the original binary image to calculate the width of the cracks. After the correspondence is completed, each crack segment is extracted, and the local distance transformation is calculated to obtain the half-width of the local area. Then, the distance transformation values of the skeleton pixels of the local crack segment are averaged and multiplied by 2 to obtain the average crack width. The calculation formula is as follows:
[0102] where, represents the Euclidean distance (i.e., half-width) from the skeleton pixel of the crack segment to the nearest background pixel; represents the coordinates of the current skeleton pixel; represents the coordinates of the background pixel.
[0103] where, represents the sum of the half-widths of all skeleton pixels of the crack segment; represents the total number of skeleton pixels.
[0104] In this embodiment, in terms of the details of model operation, the present invention conducts refined model training and optimization for the ECC uniaxial tensile crack semantic segmentation task to ensure that SegB2-Unet can efficiently and stably complete the crack semantic segmentation task. The entire training process includes key steps such as data preparation, training configuration, optimization strategy, and inference process. In addition, the present invention can not only segment the crack area, but also provide the length and width parameter information of the cracks, providing comprehensive data support for subsequent structural health monitoring.
[0105] In terms of data preparation, first, the dataset is divided into a training set (80%), a validation set (10%), and a test set (10%) according to the ratio of 8:1:1 to ensure that the model can learn the characteristics of cracks on sufficient samples and at the same time evaluate the generalization ability on the validation set. Before the data is input, it is standardized to normalize the pixel values to the range of [0,1] to improve the stability of training. In addition, in order to further enhance the generalization ability of the model, data augmentation techniques are applied during the training process, including rotation (±45° and ±90°), flipping, etc. At the same time, the image is segmented into multiple small blocks of 352×480 for training to improve the training speed of the model and enhance the model's ability to identify cracks in different directions.
[0106] In terms of training configuration, this study uses the Adam optimizer for weight update, sets the initial learning rate to 0.0001, and uses ReduceLROnPlateau as the learning rate scheduler, with a decay factor of 0.1 and a patience value of 10. During the training process, the learning rate is dynamically adjusted according to the change of the validation loss to improve the stability of model convergence and prevent overfitting. The loss functions used are Dice Loss and OHEM Focal Loss. Among them, DiceLoss alleviates the problem of data class imbalance by calculating the overlap degree between the prediction result and the true label, and effectively improves the detection ability of small target (crack) areas. Focal Loss further enhances the attention to difficult-to-separate samples and introduces the online hard example mining (OHEM) strategy. Only part of the samples with higher loss values are selected for gradient backpropagation, so as to optimize the learning effect of the model on complex crack areas. The final loss is the weighted sum of Dice Loss (weight 0.6) and OHEM Focal Loss (weight 0.4). The smoothing parameter of Dice Loss is 0.000001, the positive class weight factor (alpha) of OHEM Focal Loss is 0.25, the focusing factor (gamma) is 2.0, the OHEM ratio (ohem_ratio) is 0.7, and the total number of training epochs (Epochs) is 100. After each round of training, pixel accuracy (PA), mean intersection over union (mIoU), recall, precision, and F1Score are calculated to measure the performance of the model.
[0107] The formula is as follows:
[0108]
[0109]
[0110]
[0111] Among them, is the smoothing coefficient; is the value of the th pixel predicted by the model; is the value of the th pixel of the true label; is the intersection of the predicted crack area and the true crack area (i.e., the sum of the correctly predicted crack pixels); is the sum of the predicted crack area and the true crack area (measuring the overall size of the crack area); is the class weight factor, which gives higher weight to the crack area (positive class) to alleviate the class imbalance problem; is the focusing factor, which improves the attention to difficult-to-classify samples; dynamically adjusts the weights of easy and difficult samples, the larger it is, the higher the attention of the model to misclassified samples; and are the Dice Loss weight and the FocalLoss weight respectively.
[0112] The combined loss function can optimize the overall crack morphology while enhancing the recognition ability of crack boundaries and fine cracks, making the model perform better in the ECC uniaxial tensile crack prediction task.
[0113] During the inference process, the crack image to be recognized first undergoes the same preprocessing as in the training stage (mainly tensor conversion), and then is fed into the trained SegB2-Unet model. For large-size images, the present invention adopts a sliding window strategy, divides the image into small blocks of a fixed size, and makes predictions block by block. The model uses a multi-scale feature extraction module on each window, efficiently fuses features at all levels through skip connections and attention mechanisms, thereby capturing the fine structure of the crack, and gradually restores the spatial information in the decoding stage, and finally generates a binary segmentation mask (the crack area is marked as 1 or 255, and the background is 0).
[0114] In addition, to improve the prediction effect, the output result also undergoes post-processing operations such as non-local mean denoising, morphological dilation, and median filtering to ensure that the segmentation result is smoother and more coherent.
[0115] Through the above refined training and optimization strategies, the SegB2-Unet model of the present invention can achieve efficient, stable, and accurate crack segmentation in the ECC uniaxial tensile crack recognition task, and accurately identify the length and width of the crack, providing reliable data support and technical guarantee for subsequent structural health monitoring, crack development analysis, and engineering maintenance.
[0116] In this embodiment, the present subject uses seven evaluation indicators to evaluate the semantic segmentation model and the prediction model for ECC crack images, namely pixel accuracy (Pixel Accuracy, PA), mean pixel accuracy (Mean Pixel Accuracy, MPA), mean intersection over union (Mean Intersection over Union, MIoU), recall, precision, and F1 score (F1 Score).
[0117] The calculation of the evaluation indexes of this model is all based on the confusion matrix. Indicates originally When the class is the same and predicted as class, that is, true positive (TP) and true negative (TN), Indicates originally class is predicted as class, that is, false positive (FP) and false negative (FN). If the class is the positive class, at this time, then indicates TP, indicates TN, indicates FP, indicates FN.
[0118] The specific calculation formulas of the model evaluation indexes are as follows:
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125] Figure 4 The prediction results of the present invention on large-size high-resolution images (1760*3840) are shown. For large-size high-resolution images, the present invention adopts a sliding window strategy for step-by-step reasoning of local regions. Subsequently, non-local mean denoising, morphological dilation operation and median filtering are successively introduced to post-process the prediction results, removing noise points, connecting broken regions and reducing the false detection rate. The efficient processing of large-size images is realized, and while ensuring the robustness of the model, the precise segmentation and the retention of edge details are also taken into account, demonstrating its advantages in crack detection applications.
[0126] The experimental results show that the present invention has carried out efficient processing on large-size images and effectively obtained precise prediction results.
[0127] Figure 5The performance evaluation of the proposed SegB2-Unet model on the self-built dataset of the present invention is shown. From the data in the figure, it can be seen that the performance of a single Segformer-B2 in terms of pixel accuracy (PA), mean pixel accuracy (MPA), mean intersection over union (MIoU), recall, precision, and F1 score is approximately 98.41%, 90.32%, 86.37%, 81.20%, 89.90%, and 85.31% respectively. In contrast, U-Net has a significant improvement in the above metrics, reaching 99.06%, 96.13%, 92.19%, 92.82%, 91.41%, and 92.10% respectively; while Segformer-B0 and Segformer-B1 have lower effects in terms of PA, MPA, MIoU, recall, precision, and F1 score. It is worth noting that although the level of SegB2-Unet in terms of PA is 99.01% slightly lower than that of U-Net's 99.06%, SegB2-Unet is slightly higher than U-Net in terms of recall at approximately 93.00%, and at the same time, precision and F1 score reach 91.73% and 92.36% respectively. Therefore, SegB2-Unet is slightly superior to U-Net in terms of comprehensive performance. The above results indicate that while maintaining a relatively high overall performance, SegB2-Unet achieves a higher recall rate for the target area, providing potential value for subsequent applications in scenarios sensitive to missed detections to achieve a more comprehensive improvement in segmentation performance.
[0128] Figure 6 , Figure 7 , Figure 8 The performance evaluation of the proposed SegB2-Unet model and its prediction results on two groups of publicly available datasets of the present invention are shown. The experimental results show that the mean intersection over union (mIoU) of the model on the two datasets is 0.858 and 0.856 respectively, indicating its high segmentation accuracy and robustness in the image segmentation task. In addition, the recall rate of the first group of datasets is 0.846, while that of the second group is 0.809, suggesting that the model has a higher crack detection integrity in the former, but there may be a certain degree of missed detection problem in the latter; the corresponding precision rates are 0.852 and 0.893 respectively, and the higher precision rate of the second group of datasets indicates that the model can effectively reduce the false positive rate on this dataset. The comprehensive index F1 score is 0.848 and 0.849 respectively, further reflecting the stable performance of the model in balancing recall rate and precision. In summary, the performance of the model in this study is mainly reflected in the accuracy of image segmentation, detection integrity and false positive control, as well as the good generalization ability and robustness demonstrated on different datasets.
[0129] As Figure 9 , Figure 10 shown, the model of the present invention systematically verifies the contributions of each key algorithm module to the model performance through ablation experiments, and conducts in-depth evaluations on a self-built database without data augmentation. The core objective of the ablation experiment is to analyze the importance of the U-Net structure, multi-scale feature fusion, and attention mechanism in the SegB2-Unet semantic segmentation task, as well as the roles and impacts of each module in crack recognition.
[0130] The experimental results show that the overall segmentation performance of the basic model (Segformer-B2) is relatively limited without using additional improvement modules. Specifically, its PA (Pixel Accuracy) is 98.41%, mIoU (Mean Intersection over Union) is 86.37%, Recall is only 81.20%, Precision is 89.90%, and F1 Score is only 85.33%. This indicates that when relying solely on Segformer-B2 for crack recognition, the detection ability of the model is limited, especially in terms of the integrity recognition of cracks and the retention of edge information, where there are still obvious deficiencies. The model performs relatively well on relatively simple crack structures, but there are high rates of missed detections and false detections when facing complex crack morphologies.
[0131] When the U-Net structure is introduced, the segmentation ability of the model is significantly improved. The experimental data shows that the mIoU increases from 86.37% to 92.05%, Recall increases to 91.77%, Precision increases to 92.22%, and F1 Score increases from 85.33% to 91.99%. This indicates that the U-Net structure can effectively enhance the decoding ability of the model. Its core advantage lies in using the Skip Connections mechanism to fuse the high-resolution features in the early stage of the encoder with the low-resolution features of the decoder, thereby restoring the detailed information of the crack area while enhancing the spatial perception ability of the model. The addition of the U-Net structure makes the model perform more excellently in crack edge recognition, crack width differentiation, and connectivity maintenance, and the false detection rate decreases significantly.
[0132] On this basis, multi-scale feature fusion is further introduced, and the overall segmentation ability of the model is further optimized. Experimental data shows that the mIoU increases from 92.05% to 92.18%, Recall increases from 91.77% to 93.52%. Although Precision decreases (90.82%), the F1 Score still remains at a high level (92.15%). The core role of multi-scale feature fusion is to integrate crack information at different scales, enabling the model to simultaneously identify fine cracks and larger-scale crack propagation regions. The introduction of this module improves the model's detection ability for cracks of different sizes and enhances its adaptability to complex crack morphologies.
[0133] Finally, after adding the attention mechanism, the model performance is further optimized. The mIoU reaches 92.25%, Recall increases to 93.08%, Precision increases to 91.38%, and the F1 Score reaches 92.22%. The main contribution of the attention mechanism is to improve the model's attention to the crack region and reduce the influence of background interference. In the crack recognition task, due to possible interference factors such as noise, texture changes, and non-structural cracks on the concrete surface, traditional CNNs may be affected by irrelevant information, while the attention mechanism can automatically assign higher weights to the crack region, thereby improving the reliability of the model. In addition, the attention mechanism also plays a key role in the detection and segmentation of crack edges, making the crack morphology more complete and enhancing the applicability of the model in the actual engineering environment.
[0134] In summary, the ablation experiment fully verifies the importance of the U-Net structure, multi-scale feature fusion, and attention mechanism in the semantic segmentation task of the present invention. Among them, the U-Net structure significantly improves the overall segmentation ability of cracks, multi-scale feature fusion enhances the model's adaptability to cracks of different scales, and the attention mechanism optimizes the model's attention ability to the crack region, reduces false detections, and improves the overall segmentation accuracy.
[0135] The present invention verifies the effectiveness of each module through ablation experiments and achieves excellent performance on the self-built database without data augmentation, providing reliable technical support for high-precision crack recognition and a more practical technical solution for future automatic detection of concrete cracks and health monitoring of engineering structures.
[0136] Figure 11Shows the performance changes of the present invention before and after data augmentation on the self-built dataset. It can be seen that after data augmentation, all evaluation indicators have improved, indicating the positive impact of the data augmentation strategy on the crack segmentation task. First, in terms of PA (Pixel Accuracy), it was 99.01% before data augmentation and slightly increased to 99.46% after augmentation, indicating that the accuracy of the overall segmentation remains at a high level and the data augmentation has little impact on PA. MPA (Mean Pixel Accuracy) increased from 96.23% to 97.16%, indicating that the enhanced model classifies different categories (cracks and background) more evenly. mIoU (Mean Intersection over Union) increased from 92.25% to 94.50%, indicating that the enhanced model has a higher overlap degree between the crack area and the true annotation area, and the segmentation is more accurate. Recall increased from 93.08% to 94.62%, indicating that the model's ability to detect cracks has been enhanced, reducing the missed detection situation. Precision also increased from 91.38% to 94.37%, indicating that the detected crack area is more accurate and the false detection rate has decreased. Finally, F1Score increased from 92.22% to 94.49%, indicating that data augmentation improves the balance between recall and precision, making the overall segmentation effect more stable. As shown in Table 1.
[0137] Table 1 Optimal Results of the Model
[0138] After data augmentation, the SegB2-Unet of the present invention achieved the optimal segmentation performance on the self-built ECC crack dataset, and all indicators are close to the theoretical optimal values, indicating that the model performs excellently in terms of accuracy, robustness, and generalization ability.
[0139] First of all, the Pixel Accuracy (PA) of the model reached 99.98%, indicating that in the entire dataset, the model classifies cracks and background almost completely correctly, with a very low misclassification rate. This means that the model can provide highly reliable segmentation results in a wide range of crack recognition tasks. At the same time, the Mean Pixel Accuracy (MPA) is 99.75%, indicating that the classification accuracy of the model on different categories (cracks and background) is very balanced, avoiding the impact of class imbalance problems on the model performance.
[0140] In terms of the segmentation accuracy of the crack area, the mean intersection over union (mIoU) reaches 99.04%, indicating that the crack area predicted by the model has a very high overlap with the true crack area, and can accurately capture the shape and boundary details of the cracks. In addition, the recall rate is 99.52%, indicating that almost all cracks can be detected, and the missed detection rate is extremely low, which is crucial for the crack monitoring task to ensure that potential structural risks are not missed. At the same time, the precision rate reaches 99.23%, indicating that while ensuring a high recall rate, the model also effectively reduces the false detection situation, that is, the probability of the background area being misclassified as a crack is extremely low.
[0141] Overall, the F1 Score reaches 99.37%, which indicates that the model has achieved an optimal balance between the integrity and precision of crack recognition, that is, it can ensure the comprehensive detection of cracks and avoid false detection. This result shows that the SegB2-Unet of the present invention has achieved extremely high detection capabilities under the training optimization after data augmentation, and can be stably and efficiently applied to the crack monitoring task in actual engineering.
[0142] The improvement effect of data augmentation on the model performance is particularly significant. During the training process, through data augmentation strategies such as rotation and flipping, the model is able to learn crack characteristics in different directions and forms, improving its adaptability. In addition, data augmentation optimizes the crack edge detection ability of the model, enabling it to better segment the boundaries of cracks and reducing the problem of edge blurring, thereby improving the mIoU and F1 Score. At the same time, data augmentation expands the diversity of crack data, enabling the model to have stronger discrimination ability in complex background environments, avoiding background false detection, improving the precision, and ensuring the high reliability of the detection results.
[0143] In summary, the SegB2-Unet after data augmentation has achieved excellent performance on the self-built ECC crack dataset, and all key indicators are close to over 99%, showing extremely high segmentation accuracy and stability. This result proves that the SegB2-Unet has strong engineering applicability and can be widely applied to the health monitoring tasks of concrete structures, such as automatic crack detection and maintenance of infrastructure such as bridges, tunnels, and roads, providing reliable technical support for the future safety monitoring of civil engineering structures.
[0144] In this embodiment, the present invention accurately segments the crack area and analyzes the crack length and width. This analysis process is carried out on the basis of crack semantic segmentation. By extracting the binary mask after crack segmentation and combining pixel scale conversion and morphological analysis methods, the length and width distribution of the cracks are calculated.
[0145] In terms of analyzing the crack length, the present invention adopts an automated method based on image skeleton extraction to quantitatively analyze cracks. As shown in the figure, first, the predicted result image of the model is used to extract the crack skeleton by means of a morphological thinning algorithm; subsequently, the bifurcation points in the skeleton image are identified and removed by detecting the number of non-zero pixels in the neighborhood to ensure that the crack path is a single connected path in subsequent analysis. Then, the connected domain segmentation technology is used to divide the skeleton into regions, and an improved Dijkstra algorithm is used to calculate the true crack length of each connected region considering the diagonal distance. In order to convert the pixel length into the actual physical size, it is converted according to the pre-calibrated scale factor (pixels / mm), and then the lengths of each crack segment are statistically analyzed and the average value is obtained.
[0146] In terms of analyzing the crack width, the present invention adopts digital image processing technology to quantitatively analyze the crack width of ECC specimens. As Figure 12 shown, first, the predicted result image of the model is used to extract the crack centerline by means of skeletonization technology, and the bifurcation points in the skeleton are detected by 3×3 convolution, and then the continuous crack is cut into multiple independent crack segments. In the local area of the crack segment, it is cropped according to the minimum circumscribed rectangle of each crack segment, and the distance transform is calculated on the cropped local binary image; subsequently, only the distance values are extracted for the local skeleton area, and their average is calculated within the preset half-width range, and then multiplied by 2 to obtain the average width of the crack segment. To eliminate noise interference, only the effective crack segments with an area greater than the preset threshold are retained for width calculation.
[0147] As Figure 13 、 Figure 14 shown, the experimental results show that the model can maintain high-precision crack parameter extraction performance under different crack morphologies and background complexities, and has good stability. Especially for thinner cracks and branched cracks, the model can effectively avoid false detection and missed detection, thus ensuring the reliability of the recognition results. In addition, by optimizing the morphological processing and region analysis algorithms, the model can further improve the extraction accuracy of crack features and ensure that the automatic measurement results of cracks meet the actual engineering requirements.
[0148] Overall, the method of the present invention performs excellently in crack feature extraction and size calculation and statistics, and can provide high-quality crack parameter data for subsequent crack development evaluation and structural health monitoring, providing strong support for the safety assessment and maintenance of concrete structures.
[0149] The present invention proposes a deep learning model (SegB2-Unet) based on the combination of Segformer-B2 and U-Net for high-precision semantic segmentation of uniaxial tensile cracks in Engineered Cementitious Composites (ECC). This model integrates the global feature extraction ability of Transformer and the high-resolution feature restoration ability of U-Net structure, and combines multi-scale feature fusion and attention mechanism to effectively improve the accuracy and robustness of crack recognition.
[0150] Experimental results show that SegB2-Unet achieves superior performance on the self-built ECC crack dataset without data augmentation. It shows high accuracy in terms of mIoU (92.25%), Recall (93.08%), Precision (91.38%) and F1 Score (92.22%), indicating that this method can accurately identify cracks in complex background environments and has strong generalization ability. In addition, on the public dataset, SegB2-Unet also demonstrates stable performance, and the experimental results of its mIoU (85.8% - 85.6%), Recall (84.6% - 80.9%), Precision (85.2% - 89.3%) further verify the robustness and adaptability of the model.
[0151] Ablation experiments further verify the important roles of U-Net structure, multi-scale feature fusion and attention mechanism in the crack recognition task. The experiments show that when using Segformer-B2 alone for crack recognition, the mIoU is only 72.83%, but by gradually introducing U-Net structure, multi-scale feature fusion and attention mechanism, the mIoU is finally improved to 92.25%, indicating the effectiveness of each module. Among them, the U-Net structure significantly improves the ability to restore the edge details of cracks, multi-scale feature fusion enhances the detection ability of cracks of different sizes, and the attention mechanism optimizes the attention of the model to the crack area and improves the overall segmentation accuracy.
[0152] In addition, the present invention not only has high-precision crack segmentation ability, but also can further identify the length and width of cracks, providing more accurate parameter information for structural health monitoring. Experimental results show that the model can maintain high-precision crack feature extraction ability under different crack morphologies and background complexities, and has good stability in engineering application scenarios.
[0153] Overall, the SegB2-Unet model of the present invention breaks through the limitations of traditional methods in crack recognition and performs excellently in terms of complex crack morphology, fine crack recognition, and efficient calculation. This method not only improves the automation level of crack recognition but also provides an efficient and accurate crack segmentation and parameter recognition solution for the health monitoring of concrete structures. In the future, the model can be further optimized to adapt to more engineering scenarios and actual application requirements, such as crack recognition and maintenance of infrastructure like bridges, tunnels, and roads, thereby providing more reliable technical support for the safety guarantee of civil engineering structures.
[0154] Embodiment 2
[0155] The present invention also provides a system for recognizing and characterizing the multi-crack development behavior of high-toughness concrete based on a deep learning method. The system is used for the aforementioned method and includes an acquisition module, a preprocessing module, a construction module, and an identification module.
[0156] The acquisition module is used to construct an image feature dataset based on the features of high-toughness concrete crack image samples.
[0157] The preprocessing module is used to preprocess the image feature dataset to obtain a preprocessed dataset.
[0158] The construction module is used to construct a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and train the SegB2-Unet neural network through the preprocessed dataset to obtain a high-toughness concrete crack semantic segmentation model.
[0159] The identification module is used to identify the tensile cracks of high-toughness concrete based on the high-toughness concrete crack semantic segmentation model, extract the length and width of the cracks, and obtain crack morphology information.
[0160] The above-described embodiments are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for identifying and characterizing the development behavior of multiple cracks in high-toughness concrete based on deep learning method, characterized in that: The method comprises: An image feature dataset is constructed based on the features of high-toughness concrete crack image samples; Preprocessing the image feature data set to obtain a preprocessed data set; Building a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, training the SegB2-Unet neural network through the pre-processed data set to obtain a semantic segmentation model of high-toughness concrete cracks; Based on the high-toughness concrete crack semantic segmentation model, high-toughness concrete tensile cracks are identified, the length and width of the cracks are extracted, and the crack morphology information is obtained.
2. The method according to claim 1, characterized in that The characteristics of high-toughness concrete crack image samples include: input characteristics and output characteristics; Wherein, the input features include: experimentally produced ECC crack images, manually annotated ECC crack images; The output features include: crack semantic segmentation image, and numerical information of crack length and width.
3. The method according to claim 1, characterized in that: The method of preprocessing the image feature data set to obtain the preprocessed data set includes: Performing maximum and minimum normalization processing on the image feature data set to obtain a normalized image data set; The normalized data set is divided based on a random division method to obtain the preprocessed data set.
4. The method according to claim 1, characterized in that The method of constructing a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network includes: using Segformer-B2 as an encoder, combining with a U-Net decoder, and integrating multi-scale feature fusion and attention mechanism. The overall network structure is divided into an encoding part, a bottleneck part, and a decoding part.
5. The method according to claim 4, characterized in that In the encoder part, Segformer-B2 is used as the feature extraction network, which first processes the input image through overlapping patch embedding to segment the image into several patches with overlapping areas. Then, the image enters a four-layer Transformer block, each layer of which contains an efficient self-attention mechanism and a hybrid feedforward network to capture crack features of different scales.
6. The method according to claim 4, characterized in that The bottleneck part is located between the encoder and the decoder, and is used to further extract high-level features. This part uses 3×3 convolution for feature extraction, combined with batch normalization and ReLU activation function. The bottleneck part also has a channel adjustment function, so that the deep-level features of the encoder adapt to the input requirements of the decoder.
7. The method according to claim 4, characterized in that In the decoder part, the U-Net structure is used to gradually restore the resolution of the feature map to generate crack segmentation results. The decoder consists of multiple upsampling modules, feature fusion modules and convolution operations, namely Conv3×3 + BatchNorm + ReLU. Skip connections are introduced in the decoder so that the high-resolution features of different scales of the encoder can be directly passed to the corresponding decoding layer. An attention module based on the CBAM idea is introduced in the decoder. Specifically, this module first adopts the channel attention mechanism to extract global information through adaptive average pooling and maximum pooling, and calculates the channel attention coefficient through the fully connected layer to weight the input features and enhance the model's attention to important channels. Subsequently, the spatial attention mechanism is used to fuse the spatial information of global average pooling and maximum pooling, and the spatial attention coefficient is calculated through the convolution layer and the Sigmoid activation function to improve the model's sensitivity to crack areas.
8. A system for identifying and characterizing the multi-crack development behavior of high-toughness concrete based on a deep learning method, the system being used to implement the method described in any one of claims 1 to 7, characterized in that: The system comprises: an acquisition module, a preprocessing module, a construction module and a recognition module; The acquisition module is used to construct an image feature data set based on the features of the high-toughness concrete crack image samples; The preprocessing module is used to preprocess the image feature data set to obtain a preprocessed data set; The construction module is used to construct a SegB2-Unet neural network based on the pre-trained model Segformer-B2 and the improved U-net network, and train the SegB2-Unet neural network through the pre-processed data set to obtain a high-toughness concrete crack semantic segmentation model; The recognition module is used to identify high-toughness concrete tensile cracks based on the high-toughness concrete crack semantic segmentation model, extract the length and width of the cracks, and obtain crack morphology information.
Citation Information
Patent Citations
Electron paramagnetic resonance imaging acceleration method and system
CN119784868A
Road crack detection and segmentation method based on improved U-Net model
CN119810445A
Interactive medical image analysis with recursive vision transformer networks
EP4339888A1
Cited By
Method and device for measuring diameter of RC structural steel bar based on deep learning
CN120451250A
Pavement crack semantic segmentation method based on Transform and CNN architecture
CN120807916A
Intelligent detection and analysis method for cracks of engineering cement-based composite material
CN121544613A