Transform and CNN fused crack detection and structure evaluation system and method

By integrating Transformer and CNN into the crack detection system, accurate identification of multi-scale cracks and structural health assessment are achieved, solving the problems of insufficient detection accuracy and robustness in existing technologies, and making it suitable for real-time detection in complex environments.

CN120707523APending Publication Date: 2025-09-26GUANGXI NEW DEV TRANSPORT GRP CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510820342.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing crack detection technologies have shortcomings in accuracy, robustness, multi-scale processing and data requirements, especially in complex backgrounds and low-resolution images. The detection effect is poor, and the computational complexity is high, making it difficult to meet real-time requirements.

Method used

The crack detection system, which integrates Transformer and CNN, uses multi-scale feature extraction, self-attention mechanism, dynamic attention pooling, and self-supervised learning to achieve collaborative extraction of local and global features, refine crack boundaries, and optimize the model structure in combination with structural health assessment to adapt to various application scenarios.

Benefits of technology

It significantly improves the accuracy and robustness of crack detection, can accurately identify multi-scale cracks in complex environments, reduces computational complexity, is suitable for real-time detection, reduces dependence on large-scale labeled data, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707523A_ABST
    Figure CN120707523A_ABST
Patent Text Reader

Abstract

The invention provides a Transform and CNN fused crack detection and structure evaluation system and method. The system comprises an image preprocessing module, a crack feature extraction module, a crack positioning and classification module, a crack boundary refinement module and a structure integrity evaluation module. Through combination of the CNN and the Vision Transform, local and global features in the image can be extracted at the same time, and the precision of crack detection is enhanced. The dynamic attention mechanism is used for refining fracture boundaries and improving fracture positioning and recognition effects. And the structure health assessment module combines crack information and structure stress analysis, performs structure risk assessment by using a support vector machine or a random forest, and outputs a structure health state and a repair suggestion. The invention further provides an evaluation scheme of the system. The method improves the precision and robustness of crack detection, has higher multi-scale detection capability, noise robustness and real-time performance, and is suitable for automatic monitoring and health management of civil infrastructures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to crack detection and structural health monitoring technology for civil infrastructure in the field of civil engineering, and specifically to a crack detection and structural assessment system and method that integrates Transformer and CNN (convolutional neural network). Background Art

[0002] Civil infrastructure (such as bridges, roads, tunnels, and buildings) serves as a vital pillar of national economic development, and its safety directly impacts the safety of society and the lives and property of its citizens. Cracks are a significant issue affecting the structural safety of civil infrastructure. As these structures age, cracks gradually develop. Especially when subjected to environmental factors (such as temperature, humidity, and load fluctuations), crack propagation can lead to structural instability or even catastrophic failure. Therefore, regular and accurate crack detection is crucial for infrastructure maintenance, management, and repair.

[0003] Traditional crack detection methods rely heavily on manual visual inspection and simple image processing techniques. While intuitive, manual inspection suffers from issues such as low efficiency, significant subjective error, and limited coverage. Furthermore, it can be dangerous to perform in harsh environments. Image processing techniques, which capture and analyze crack images using cameras, rely on traditional image processing methods such as edge detection, morphological processing, and region segmentation. These methods have limitations in noise handling, complex backgrounds, and the accuracy of crack feature extraction.

[0004] In recent years, with the rapid development of artificial intelligence and deep learning technologies, crack detection methods based on deep learning have gradually become a research hotspot. In particular, the widespread application of convolutional neural networks (CNNs) in image processing has achieved considerable progress. By extracting and classifying features from crack images, CNNs can significantly improve detection accuracy and efficiency. Some research has combined deep learning technology with drones and robotic platforms, utilizing automated equipment for crack detection, further enhancing the automation and adaptability of detection.

[0005] At present, crack detection technologies based on deep learning are mainly divided into three categories:

[0006] (1) Crack image classification algorithm: This method classifies images into two categories: cracks and non-cracks, and uses a trained CNN model to determine whether the image contains cracks. Although this method is simple and computationally efficient, its disadvantage is that it cannot provide information about the location of cracks, thus limiting the accuracy of structural health assessment.

[0007] (2) Crack target detection algorithm: This method detects the specific location of cracks in the image through a region proposal network (such as Faster R-CNN, YOLO, etc.). Compared with image classification algorithms, this type of method can provide crack location information, but it has the following problems:

[0008] Low detection accuracy: Especially in images with complex backgrounds and low contrast, target detection algorithms often face problems of missed detection and false detection.

[0009] Poor detection capability of small cracks: Existing object detection algorithms have difficulty in effectively detecting tiny cracks, especially in images with high noise or low resolution.

[0010] (3) Crack semantic segmentation algorithm: Semantic segmentation algorithm can achieve pixel-level crack detection by classifying each pixel, thereby providing more detailed crack information. This type of method is usually based on architectures such as FCN (Fully Convolutional Neural Network) and U-Net, and achieves accurate crack positioning through image segmentation. However, semantic segmentation technology also has some problems in practical applications:

[0011] Large demand for training data: In order to perform effective pixel-level crack segmentation, a large amount of accurately labeled training data is required. However, in actual engineering, the workload of pixel-level labeling is extremely large, and labeling errors may also affect the training effect.

[0012] Insufficient robustness: Existing semantic segmentation algorithms have poor adaptability to noise, low resolution, and background complexity, resulting in detection accuracy in real environments failing to meet high standards.

[0013] Blurred crack boundaries: Existing semantic segmentation methods often have a certain degree of ambiguity when dealing with crack boundaries, and are unable to accurately separate cracks from background areas. This is especially true when dealing with small cracks, which poses the risk of unclear detection or missed detection.

[0014] Although existing deep learning methods have achieved certain results in crack detection, they still face the following major technical challenges:

[0015] (1) Crack size differences: Crack sizes vary greatly, with great variability from tiny cracks to larger cracks. Existing models usually find it difficult to maintain high accuracy when dealing with cracks of different sizes.

[0016] (2) Complex background and environmental noise: Crack images often contain complex background information and environmental noise. Existing algorithms are not robust enough in these situations, resulting in unstable detection results.

[0017] (3) Low-resolution image processing: Most traditional crack detection methods assume that the image has a high resolution, but in actual applications, the image quality is often limited by the equipment, resulting in a decrease in detection accuracy.

[0018] (4) Lack of data sets: The training of deep learning models usually requires a large number of well-labeled data sets, but the existing crack image data sets are small in scale, and the annotation quality of many data sets is not high, which makes it difficult to meet the needs of deep learning training.

[0019] The shortcomings or problems of the existing technology are as follows:

[0020] (1) Lack of robustness: Existing crack detection models often make false detections or missed detections when processing images with low resolution, noise interference, and complex backgrounds, and cannot work stably in a changing real environment.

[0021] (2) Strong data dependence: Deep learning models are highly dependent on datasets. Most existing datasets lack diversity and annotation accuracy. Data quality issues during training affect model performance.

[0022] (3) Poor ability to handle multi-scale cracks: Existing deep learning methods have performance bottlenecks when dealing with cracks of different scales, especially for small cracks or cracks with complex morphology, the detection effect is poor.

[0023] (4) High computational complexity: The computational complexity of some deep learning models is high, especially in real-time detection, which makes it difficult to meet the real-time requirements of engineering applications.

[0024] In summary, existing crack detection technologies have varying degrees of deficiencies in accuracy, robustness, multi-scale processing, and data requirements. There is an urgent need for a new detection method that can overcome these problems. Summary of the Invention

[0025] In response to the problems existing in the existing technologies, such as low detection accuracy, poor robustness, and difficulty in multi-scale crack detection, the present invention provides a crack detection and structural assessment system and method that integrates multi-scale Transformer and CNN. The purpose is to integrate CNN and Vision Transformer for the first time to achieve collaborative extraction of local and global features and enhance the recognition capability of complex crack morphologies. A dynamic attention pooling mechanism is introduced to refine crack boundaries and improve boundary recognition accuracy. An integrated framework for detection and risk assessment is constructed to achieve end-to-end processing from image input to structural health output. Combined with a self-supervised learning mechanism, it reduces dependence on manually labeled data and reduces deployment costs. The model structure is optimized to improve real-time performance and edge deployment capabilities to adapt to a variety of application scenarios. The accuracy, robustness, and applicability of crack detection are significantly improved, providing an efficient and accurate solution for infrastructure detection in the civil engineering field.

[0026] In order to achieve the above object, the specific scheme of the present invention is as follows:

[0027] The crack detection and structural assessment system that integrates Transformer and CNN includes:

[0028] An image preprocessing module is used to receive the original crack image, perform denoising, enhancement and scale transformation on the original crack image, and output preprocessed image data;

[0029] A crack feature extraction module, connected to the image preprocessing module, is used to receive the preprocessed image data and extract local and global features of the crack image by combining the convolutional neural network submodule with the Transformer submodule in the Vision Transformer through multi-scale feature extraction and self-attention mechanism, and output crack feature data;

[0030] A crack location and classification module, connected to the crack feature extraction module, is used to receive crack feature data, locate crack positions and classify cracks based on the extracted local and global features of the crack image using a target detection algorithm, and output crack position and type information;

[0031] a crack boundary refinement module, connected to the crack location and classification module, configured to receive crack location and type information, process the located and classified crack areas, refine the crack boundaries using dynamic attention pooling and boundary refinement algorithms, and output refined crack boundary data;

[0032] The structural integrity assessment module is connected to the crack boundary refinement module and is used to receive the refined crack boundary data. It combines the refined crack boundary feature information with the stress analysis results of the structure, uses a support vector machine or a random forest model to perform structural health assessment, calculates the risk score of the crack, and outputs the risk level of the structure and corresponding repair suggestions based on the risk score and stress conditions of the crack. In cases where data is scarce or difficult to label, a self-supervised learning process is performed.

[0033] Furthermore, the image preprocessing module includes:

[0034] A denoising autoencoder, configured to receive an original crack image and remove noise from the original crack image;

[0035] an image enhancement unit, connected to the denoising autoencoder, configured to receive the denoised crack image and perform histogram equalization and edge enhancement on the crack image;

[0036] a scale transformation module, connected to the image enhancement unit, configured to receive the enhanced crack image, perform multi-scale processing on the crack image, and obtain image copies with different resolutions;

[0037] The crack feature extraction module includes:

[0038] The convolutional neural network submodule is used to extract local features of the crack image, including edges and textures. It uses a multi-layer convolutional neural network architecture, which includes at least one convolution layer, one pooling layer, and one activation layer to capture the subtle structural features of the crack.

[0039] A Transformer submodule, connected to the convolutional neural network submodule, is used to extract global features of the crack image through a self-attention mechanism, enhance the importance of the crack region in the image, and capture long-range dependencies. The specific implementation includes parallel processing of multiple self-attention heads and an encoder layer based on the Transformer architecture. The encoder layer performs global weighted aggregation of input features through weighted summation, thereby strengthening the global feature representation of the crack region.

[0040] The crack location and classification module includes:

[0041] The target detection algorithm unit is used to receive the feature information output by the crack feature extraction module, use the Faster R-CNN algorithm or the YOLOv5 algorithm to generate candidate frames for the cracks, and determine the locations of the cracks;

[0042] A classifier, connected to the target detection algorithm unit, configured to receive crack candidate frames generated by the target detection algorithm unit and classify crack types using a softmax classifier;

[0043] The crack boundary refinement module includes:

[0044] The dynamic attention pooling unit receives the crack area information output by the crack localization and classification module, calculates the attention weight of each pixel through the self-attention mechanism in the Vision Transformer, and dynamically enhances the features of the crack area, including the crack boundary.

[0045] A boundary refinement algorithm unit, connected to the dynamic attention pooling unit, is used to receive the dynamically enhanced crack feature information and perform pixel-level optimization on the crack boundary using a convolutional neural network submodule and morphological operations;

[0046] The structural integrity assessment module includes:

[0047] Crack parameter extraction unit, used to extract the crack length, width, depth, location, and type characteristics from the crack image, and combine them with the load conditions and material properties of the structure;

[0048] a structural stress analysis module, connected to the crack parameter extraction unit, for simulating the effect of the crack on the structure through finite element analysis based on the extracted crack parameters, and calculating the stress in the crack area;

[0049] The risk assessment module, connected to the structural stress analysis unit, is used to perform a structural health assessment using a support vector machine or random forest algorithm based on crack parameters and structural stress analysis results, outputting a structural risk level and corresponding repair recommendations based on the structural risk level. Crack parameters include crack length, width, type, and depth.

[0050] The image preprocessing module, crack feature extraction module, crack location and classification module, crack boundary refinement module and structural integrity assessment module are all driven by a deep neural network (DNN) to achieve end-to-end automated processing.

[0051] A crack detection and structural assessment method integrating Transformer and CNN includes the following steps:

[0052] Step 1, image preprocessing step: receiving the original crack image, and performing denoising, enhancement and scale transformation on the original crack image to generate preprocessed image data, including multiple image copies at different scales;

[0053] Step 2, crack feature extraction step: The image data preprocessed in step 1 is input into the convolutional neural network submodule, and the local features of the cracks, including edges and textures, are extracted through the convolutional layer. The extracted local features are input into the Transformer submodule in the VisionTransformer, and the global features of the cracks are extracted through the self-attention mechanism, which enhances the importance of the crack area in the image and captures long-range dependencies. Features are extracted from the image copies of different scales described in step 1 to obtain features at different scales.

[0054] Step 3: The features of different scales described in step 2 are fused through cross-layer connection and cascade convolution method to obtain a fused feature map;

[0055] Step 4, crack location and classification: Input the fused feature map from step 3 into the target detection algorithm unit, use the Faster R-CNN algorithm or the YOLOv5 algorithm to generate candidate crack boxes and determine the crack locations. Input the crack candidate boxes into the softmax classifier to classify the crack types and output the crack location and type information.

[0056] Step 5, crack boundary refinement step: The crack location and type information described in step 4 is input into the dynamic attention pooling unit. The self-attention mechanism in the Vision Transformer is used to calculate the attention weight of each pixel, dynamically enhancing the features of the crack area, including the crack boundary. The dynamically enhanced crack feature information is input into the boundary refinement algorithm unit. The convolutional neural network submodule and morphological operations are used to optimize the crack boundary at the pixel level to generate refined crack boundary data.

[0057] Step 6, structural integrity assessment step: Generate refined crack boundary data from step 5 and extract crack parameters from the crack image. The crack parameters include the characteristics of the length, width, depth, location and type of the crack, and are combined with the load conditions and material properties of the structure; based on the extracted crack parameters, simulate the impact of the crack on the structure through finite element analysis and calculate the stress in the crack area; based on the crack parameters and the structural stress analysis results, use support vector machines or random forests to perform structural health assessment, calculate the risk score of the crack, and output the health status of the structure and corresponding repair suggestions based on the risk score and stress conditions of the crack.

[0058] Furthermore, the denoising of the original crack image in step 1 is performed using a denoising autoencoder, which is trained by minimizing the following loss function:

[0059] ,

[0060] Where: is the input image, is the denoised image, is the number of pixels in the image;

[0061] The original crack image is enhanced by using histogram equalization and edge enhancement technology to enhance image contrast and edge information. The formula for histogram equalization is as follows:

[0062] ,

[0063] Where: is the pixel value after equalization; is the pixel value of the original image; Is the mapping function, representing the original grayscale value To the gray value after equalization Mapping; Indicates an 8-bit grayscale image, and L takes the value of 256; is the probability density function of the original image; Indicates a small change in the integral, an integral operation performed to calculate the cumulative probability;

[0064] Edge enhancement uses the Sobel operator to enhance the crack edge, and uses the convolution kernel to detect the crack edge in the image to enhance the edge information. The convolution kernel is as follows:

[0065] ,

[0066] The image scale transformation of the original crack image is to use a pooling layer to reduce the image size and obtain image copies of multiple scales. The pooling operation formula is as follows:

[0067] ,

[0068] Where, is the original image, is the pooled copy of the image.

[0069] Furthermore, the convolution operation formula for extracting the local features of the crack in step 2 is as follows:

[0070] ,

[0071] Where: Represents the position in the output feature map The value of Represents the position in the convolution kernel The weight of Indicates that the input image is located at The pixel value of the position; is the input image, is the convolution kernel, is the feature map after convolution; m and n represent the offset index within the convolution kernel, which is used to extract the local area by sliding the window in the input image; 、 Represents the location coordinates of the previous output feature map.

[0072] The self-attention mechanism formula for extracting the global features of the crack is as follows:

[0073] ,

[0074] Where: is the query matrix; is the bond matrix; is a value matrix; is the dimension of the key matrix; the self-attention mechanism helps identify the relationships and patterns of cracks in complex backgrounds; T represents the matrix transpose operation.

[0075] Furthermore, the formula for feature fusion in step 3 is:

[0076] ,

[0077] Where: 、 They are features of different scales, It is the fused feature map.

[0078] Furthermore, the calculation formula for the candidate box of the crack in step 4 is as follows:

[0079] ,

[0080] Where: and are the coordinates of the upper left and lower right corners of the crack candidate box;

[0081] The position of the crack is determined by using Intersection over Union to measure the overlap between the candidate box and the actual crack. The formula is as follows:

[0082] ,

[0083] The function formula of the Softmax classifier is:

[0084] ,

[0085] Among them, e represents a natural constant, which is approximately equal to 2.71828. It is the base of the exponential function and is used to calculate the exponential value; Indicates that the crack belongs to The class score is used to normalize the probability of the entire class; Indicates the Unnormalized probability values ​​of the classes, used for summing; Is the crack belong to Class score; Indicates the Unnormalized probability values ​​of the classes; Represents input features Next, the crack belongs to class probability; is the total number of crack classification categories.

[0086] Furthermore, the formula for calculating the attention weight of each pixel in step 5 is as follows:

[0087] ,

[0088] Where: It's a pixel Pixel The attention weight of Represents pixels With pixels similarity between Represents pixels With pixels similarity between Indicates the total number of pixels or feature vectors in the image; 、 、 Indicates the image 、 and Feature representation of pixels; Indicates the sum of all pixels in the image for normalization; Represents an exponential function used to enhance differences and make items with high similarity more prominent.

[0089] The boundary refinement algorithm unit further refines the crack boundary through morphological operations. The morphological operations include erosion and dilation. The formula of the erosion operation is as follows:

[0090]

[0091] Where: is the image at position after corrosion Pixel value, output image; It means taking the minimum value of all pixel values ​​in the area covered by the structural element to achieve the "erosion" effect, that is, shrinking the boundary; represents the original input image; Represents the original image I in the structural element The pixel position under coverage The corresponding value; It is a structural element that defines the shape and extent of corrosion; 、 Represents corroded structural elements The offset of each pixel relative to the center point.

[0092] Furthermore, the formula for calculating the stress in the crack area in step 6 is as follows:

[0093] ,

[0094] Where: is the stress in the crack area; is the externally applied load; is the effective cross-sectional area of ​​the crack region;

[0095] The calculation formula of the risk score of the crack is as follows:

[0096] ,

[0097] Where: Score the risk of cracks; represents a comprehensive risk scoring function, which is used to calculate the final risk score based on multiple input fracture parameters and model output; Output represents the crack damage grade result output by the support vector machine, which is used to reflect the severity of the crack and is a key input parameter of the crack risk scoring function; Indicates the length characteristics of the crack; Represents the width characteristics of the crack.

[0098] Furthermore, the structural integrity assessment step in step 6 further includes performing a self-supervised learning process when data is scarce or difficult to label. The self-supervised learning process includes:

[0099] Step 61: using some manually annotated crack image data to train a preliminary training model combining a convolutional neural network and a VisionTransformer;

[0100] Step 62, generating pseudo labels: using the preliminarily trained model in step 1 to predict the unlabeled crack images and generate pseudo labels;

[0101] Step 63, calculating the loss function of self-supervised learning: using the pseudo labels for training, and calculating the loss function of self-supervised learning to ensure that the error between the pseudo labels and the actual labels is minimized. The loss function of self-supervised learning is defined as:

[0102] ,

[0103] Where: is the pseudo label generated by the model; is the actual label; is the number of training samples; by minimizing the loss function, the learning process of the model is optimized;

[0104] Step 64, data expansion and optimization: Use pseudo-labeled and unlabeled data for further training and optimization to expand the dataset.

[0105] Advantages of the present invention

[0106] The crack detection and structural assessment system and method of this invention, which integrates a Transformer and CNN (convolutional neural network), is suitable for intelligent civil engineering inspection and automated crack detection and real-time assessment using autonomous equipment such as drones and robots. This is particularly true for crack detection tasks in complex environments (such as high altitudes, confined spaces, and inclement weather). By incorporating a deep learning framework, this invention overcomes the shortcomings of traditional manual inspection and image processing-based crack detection technologies in terms of efficiency, accuracy, and robustness, further promoting the intelligent and automated development of civil infrastructure inspection. It also offers the following advantages:

[0107] 1. Multi-scale crack detection capability: The convolutional neural networks (CNNs) used in traditional crack detection methods often focus on crack features at a single scale, resulting in insufficient detection of small cracks and complex crack morphologies. However, this invention introduces multi-scale feature extraction, processing crack images through multiple copies at different resolutions. This ensures that the model can simultaneously detect both micro and large cracks, enhancing the robustness and comprehensiveness of crack detection. Cracks of varying scales can be accurately processed within a unified framework, eliminating the loss of important information due to low image resolution. This significantly improves the detection accuracy of cracks of various sizes (particularly microcracks and low-resolution crack images) and effectively addresses the difficulty traditional crack detection methods have in effectively handling cracks of varying sizes, particularly distinguishing between micro and large cracks.

[0108] 2. A hybrid deep learning architecture combining Transformer and CNN: This method combines the local feature extraction capabilities of CNN with the global context modeling capabilities of Vision Transformer. CNN efficiently extracts local features, while the Transformer submodule models long-range dependencies in the image through a self-attention mechanism. This improves the global recognition of cracks and improves crack detection accuracy, especially in complex backgrounds and with complex crack morphologies. This effectively addresses the problem that existing crack detection methods, which are mostly based on a single convolutional neural network (CNN) architecture, have a poor ability to capture global information and long-range dependencies in crack images, resulting in unclear crack edges and inaccurate crack identification.

[0109] 3. Adaptive Crack Boundary Refinement: This method introduces dynamic attention pooling and boundary refinement algorithms to adaptively focus on crack boundary areas, refine the boundaries, and eliminate boundary blurring caused by noise or low resolution. This significantly improves boundary accuracy, enhances the robustness of crack detection in complex backgrounds and low-resolution images, and reduces the risk of false and missed detections. This avoids the false and missed detection problems caused by boundary blurring in traditional methods.

[0110] 4. Joint Processing of Crack Location and Structural Risk Assessment: This invention combines crack detection with structural health assessment. Using machine learning models such as support vector machines (SVM) or random forests (RF), it comprehensively considers crack location, size, type, and its impact on structural stresses to perform a structural health risk assessment. It also outputs structural health status and repair recommendations. This enables a joint assessment of cracks and structural health, providing a scientific basis for engineering maintenance and improving monitoring accuracy and decision-making effectiveness. This avoids the existing problem of separating crack detection and structural assessment, improving the accuracy of structural health monitoring and the effectiveness of decision-making support.

[0111] 5. Effectiveness of self-supervised learning in data-scarce situations: This invention utilizes self-supervised learning, enabling effective training based on a small amount of labeled data. By generating pseudo-labels and training on unlabeled data, it improves training efficiency and data utilization, reduces reliance on large-scale labeled datasets, lowers labeling costs, and improves training efficiency and accuracy when data is scarce. It is suitable for real-time, on-site data collection. This solves the problem that most existing crack detection systems rely on large, labeled datasets, which are often limited by time, cost, and other factors, resulting in data scarcity or imbalance.

[0112] 6. Real-time Deployment Capability and Efficiency: This invention utilizes optimized algorithms and model compression to ensure real-time processing on edge devices. This efficient design, combined with parallel computing, enables rapid crack detection and structural health assessment, enabling efficient, real-time crack detection and structural health assessment. This technology is suitable for mobile platforms and large-scale monitoring, providing immediate decision support for maintenance. This addresses the computational complexity of existing crack detection technologies, which hinders their application in real-time monitoring, particularly in deployment on mobile platforms such as drones and robots.

[0113] 7. Enhanced noise robustness and low-resolution adaptability: The present invention uses a denoising autoencoder and image enhancement methods to remove noise, improve the detection effect of low-resolution images, ensure the stable operation of the system under harsh environmental conditions, improve the robustness of the detection system in complex backgrounds and low-quality images, reduce errors, and solve the problem that traditional crack detection methods perform poorly in noisy or low-resolution images, resulting in false detections and missed detections and causing errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1 Flowchart of the crack detection and structure assessment method integrating multi-scale Transformer and CNN of the present invention DETAILED DESCRIPTION

[0115] The present invention will be further explained and illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be noted that this specific embodiment is not intended to limit the scope of rights of the present invention.

[0116] like Figure 1 As shown, this specific embodiment provides a crack detection and structural assessment system and method that integrates a multi-scale Transformer and a CNN, which is particularly suitable for automated monitoring of structures such as roads, bridges, tunnels, and buildings. The crack detection and structural assessment system that integrates a multi-scale Transformer and a CNN includes:

[0117] An image preprocessing module is used to receive the original crack image, perform denoising, enhancement and scale transformation on the original crack image, and output preprocessed image data;

[0118] Specifically, the image preprocessing module includes:

[0119] A denoising autoencoder, configured to receive an original crack image and remove noise from the original crack image;

[0120] an image enhancement unit, connected to the denoising autoencoder, configured to receive the denoised crack image and perform histogram equalization and edge enhancement on the crack image;

[0121] a scale transformation module, connected to the image enhancement unit, configured to receive the enhanced crack image, perform multi-scale processing on the crack image, and obtain image copies with different resolutions;

[0122] A crack feature extraction module, connected to the image preprocessing module, is used to receive the preprocessed image data and extract local and global features of the crack image by combining the convolutional neural network submodule with the Transformer submodule in the Vision Transformer through multi-scale feature extraction and self-attention mechanism, and output crack feature data;

[0123] Specifically, the crack feature extraction module includes:

[0124] The convolutional neural network submodule is used to extract local features of the crack image, including edges and textures. It uses a multi-layer convolutional neural network architecture, which includes at least one convolution layer, one pooling layer, and one activation layer to capture the subtle structural features of the crack.

[0125] The Transformer submodule is connected to the convolutional neural network submodule and is used to extract global features of the crack image through a self-attention mechanism, enhance the importance of the crack area in the image, and capture long-range dependencies. The specific implementation includes parallel processing of multiple self-attention heads and an encoder layer based on the Transformer architecture. The input features are globally weighted aggregated through weighted summation, thereby strengthening the global feature representation of the crack area.

[0126] A crack location and classification module, connected to the crack feature extraction module, is used to receive crack feature data, locate crack positions and classify cracks based on the extracted local and global features of the crack image using a target detection algorithm, and output crack position and type information;

[0127] Specifically, the crack location and classification module includes:

[0128] The target detection algorithm unit is used to receive the feature information output by the crack feature extraction module, use the Faster R-CNN algorithm or the YOLOv5 algorithm to generate candidate frames for the cracks, and determine the locations of the cracks;

[0129] The classifier is connected to the target detection algorithm unit, and is used to receive the crack candidate box generated by the target detection algorithm unit and classify the crack type through a softmax classifier.

[0130] a crack boundary refinement module, connected to the crack location and classification module, configured to receive crack location and type information, process the located and classified crack areas, refine the crack boundaries using dynamic attention pooling and boundary refinement algorithms, and output refined crack boundary data;

[0131] Specifically, the crack boundary refinement module includes:

[0132] The dynamic attention pooling unit is used to receive the crack area information output by the crack localization and classification module. This unit is mainly used to calculate the attention weight of each pixel through the self-attention mechanism in the Vision Transformer during the crack boundary refinement stage. It assigns dynamic attention weights to different areas in the crack image (especially the boundary areas), dynamically enhances the features of the crack area, including the crack boundary part, and suppresses background interference.

[0133] The structure and execution process of the dynamic attention pooling unit are as follows:

[0134] 1. Composition, Input and Output of Dynamic Attention Pooling Unit

[0135] Input: Feature map from the "Crack Localization and Classification Module"

[0136] Output: Weighted crack area feature map for use by the "boundary refinement algorithm unit"

[0137] Internal composition:

[0138] Query-Key-Value Mapping Layer (Linear Projection)

[0139] Self-attention weight calculation unit

[0140] Dynamic pooling aggregation unit (weighted pooling)

[0141] Output feature fusion layer

[0142] 2. Operation process and calculation steps

[0143] 1: Feature encoding is query, key, value ( 、 、 )

[0144] Input feature map Divided into pixel vector , projected through three linear layers as follows:

[0145] , ,

[0146] Where: is the input image feature tensor, with dimension ; 、 、 are query, key, and value matrices, each with dimensions ; , , is a learnable linear transformation weight matrix with dimension ; Embedding dimension for attention; is the number of pixels in the image, equal to .

[0147] 2: Calculate the self-attention weight matrix

[0148] Each position and location The similarity is calculated by dot product:

[0149] ,

[0150] Perform softmax normalization on each row to obtain the attention allocation probability:

[0151] ,

[0152] Where: For query and key similarity between From the position To location The attention weight indicates the information contribution ratio; is a scaling factor used to avoid gradient explosion.

[0153] 3: Attention pooling output features

[0154] Use the weight matrix to calculate the value vector Weighted summation gives a new representation for each pixel:

[0155] ,

[0156] Where, Indicates the New feature representation of each position; represents the attention weight and the position and location the relationship between or intensity of attention; Indicates the The feature vector of the position; Indicates the total number of pixels or feature vectors.

[0157] Finally, all positions are summarized into the attention-weighted feature map:

[0158] ,

[0159] Where: : Dynamic enhancement features of the output; can be reshaped into , which is used by the subsequent boundary refinement algorithm. : Query vector, which represents the "query" information of the current feature in the global context. It contains the features of the current position; : key vector, representing the “key” feature of each position, which is usually the same as the query vector Compare to calculate the attention weight. ; : By query vector and key vector The attention weight matrix determined by the similarity between them; : Value vector, representing the "value" or "feature" of each position. It contains the feature information of each position and will eventually be weighted and aggregated through the attention mechanism to obtain the output feature.

[0160] 4: Data interaction mechanism with previous and next modules

[0161] Upstream interface: Receives feature maps from the "Crack Location and Classification Module";

[0162] Downstream interface: Enhanced features Input to the "boundary refinement algorithm unit" and used to dynamically select important pixels in the erosion / dilation operation;

[0163] Channel adjustment: If , the channel dimension is adjusted by 1×1 convolution to maintain consistency.

[0164] 3. Key Parameters and Default Settings

[0165] : The default value is 64, description: self-attention embedding dimension;

[0166] , : The default value is 32,32, description: feature map space size;

[0167] Number of heads: The default value is 4. Description: Multi-head parallel attention enhances boundary information representation;

[0168] Positional Encoding, the default value is yes, description: use two-dimensional sinusoidal position encoding to maintain spatial structure;

[0169] Dropout: The default value is 0.1. Description: An attention regularization measure to prevent overfitting.

[0170] 4. Advantages and functions

[0171] Ability to dynamically adjust the feature intensity of different regions so that the model automatically focuses on the crack edge;

[0172] Effectively suppress background noise and improve boundary clarity;

[0173] Significantly improves boundary recognition accuracy and robustness in low-contrast, blurred crack images.

[0174] The boundary refinement algorithm unit is connected to the dynamic attention pooling unit, and is used to receive the dynamically enhanced crack feature information and use a convolutional neural network submodule and morphological operations to perform pixel-level optimization on the crack boundary.

[0175] The structural integrity assessment module is connected to the crack boundary refinement module, and is used to receive the refined crack boundary data, combine the refined crack boundary feature information and the stress analysis results of the structure, use the support vector machine or random forest model to perform structural health assessment, predict the health status of the structure and generate repair suggestions, and output the structural health assessment results and repair suggestions.

[0176] Specifically, the structural integrity assessment module includes:

[0177] Crack parameter extraction unit, used to extract the crack length, width, depth, location, and type characteristics from the crack image, and combine them with the load conditions and material properties of the structure;

[0178] a structural stress analysis module, connected to the crack parameter extraction unit, for simulating the effect of the crack on the structure through finite element analysis based on the extracted crack parameters, and calculating the stress in the crack area;

[0179] The risk assessment module is connected to the structural stress analysis unit and is used to perform structural health assessment based on crack parameters and structural stress analysis results using a support vector machine or random forest, output the risk level of the structure, and output corresponding repair suggestions based on the risk level of the structure. In cases where data is scarce or difficult to label, a self-supervised learning process is performed.

[0180] The structures mentioned in the above system refer to infrastructure structures in the field of civil engineering, including but not limited to the following categories:

[0181] Roads: including highways, urban roads, rural roads, etc., the surface or foundation of these roads may develop cracks, affecting the structural integrity and service life of the roads.

[0182] Bridges: Bridges are an important part of transportation infrastructure and their structures may be affected by cracks, especially the bridge deck, piers and beams.

[0183] Tunnels: The lining and rock mass of a tunnel may develop cracks that may affect the stability and safety of the tunnel.

[0184] Buildings: including residential, commercial, and industrial buildings, may have cracks in their walls, foundations, and floor slabs, affecting the structural safety of the buildings.

[0185] Dams and dams: Dams and dam structures in water conservancy projects may have cracks on their surfaces and inside due to water pressure, geological changes and other factors.

[0186] Other civil infrastructure: such as ports, docks, airport runways, etc. These structures may also develop cracks due to various reasons.

[0187] This system enables rapid crack detection, precise classification, and health assessment. It boasts higher detection accuracy, enhanced environmental adaptability, and more efficient real-time processing capabilities, significantly improving the maintenance efficiency and safety of civil infrastructure. It not only effectively addresses the limitations of traditional crack detection methods, such as insufficient precision, poor noise processing, and difficulty detecting multi-scale cracks, but also enables high-precision crack detection in complex environments, providing an accurate basis for structural health assessment.

[0188] A crack detection and structural assessment method integrating Transformer and CNN includes the following steps:

[0189] Step 1, image preprocessing step: receiving the original crack image, and performing denoising, enhancement and scale transformation on the original crack image to generate preprocessed image data, including multiple image copies at different scales to ensure that the image quality is suitable for subsequent processing;

[0190] The denoising of the original crack image is performed by using a denoising autoencoder to compress and decode the image through a convolutional network to remove noise and retain the details of the crack.

[0191] Encoder: Use convolutional layers to compress the image to obtain a low-dimensional representation.

[0192] Decoder: The compressed low-dimensional representation is reconstructed into a denoised image through a deconvolution layer.

[0193] The denoising autoencoder is trained by minimizing the following loss function:

[0194] ,

[0195] Where: is the input image, is the denoised image, is the number of pixels in the image;

[0196] The enhancement of the original crack image is to enhance the image contrast and edge information by using histogram equalization and edge enhancement technology, so that the crack is easier to identify.

[0197] The formula for histogram equalization is as follows:

[0198] ,

[0199] Where: is the pixel value after equalization; is the pixel value of the original image; Is the mapping function, representing the original grayscale value To the gray value after equalization Mapping; Indicates an 8-bit grayscale image, and L takes the value of 256; is the probability density function of the original image; Indicates a small change in the integral, an integral operation performed to calculate the cumulative probability;

[0200] Edge enhancement uses the Sobel operator to enhance the crack edge, and uses the convolution kernel to detect the crack edge in the image to enhance the edge information. The convolution kernel is as follows:

[0201] ,

[0202] In order to detect cracks of different sizes, the original crack image is scaled using a pooling layer to reduce the image size, obtain image copies of multiple scales, and input them into the deep learning model for processing. The pooling operation formula is as follows:

[0203] ,

[0204] Where, is the original image, is the pooled copy of the image.

[0205] Step 2, crack feature extraction: The image data preprocessed in step 1 is input into the convolutional neural network submodule (CNN module for short), and the local features of the cracks, including edges and textures, are extracted through the convolutional layer. The extracted local features are input into the Transformer submodule of the Vision Transformer (ViT for short) to identify the overall morphology of the cracks. The global features of the cracks are extracted through the self-attention mechanism, which enhances the importance of the crack area in the image and captures long-range dependencies. Features are extracted from the image copies of different scales described in step 1 to obtain features at different scales.

[0206] The local features of the cracks are extracted by extracting features at different levels through convolution layers. The convolution operation formula is as follows:

[0207] ,

[0208] Where: Represents the position in the output feature map The value of Represents the position in the convolution kernel The weight of Indicates that the input image is located at The pixel value of the position; is the input image, is the convolution kernel, is the feature map after convolution; m and n represent the offset index within the convolution kernel, which is used to extract the local area by sliding the window in the input image; i and j represent the position coordinates of the previous output feature map.

[0209] The self-attention mechanism formula for extracting the global features of the crack is as follows:

[0210] ,

[0211] Where: is the query matrix; is the bond matrix; is a value matrix; is the dimension of the key matrix; the self-attention mechanism helps identify the relationship and pattern of cracks in complex backgrounds; T represents the matrix transpose operation.

[0212] In step 3, the features of different scales described in step 2 are fused through cross-layer connection and cascade convolution method to obtain a fused feature map; to ensure that the model can simultaneously use local and global information for crack detection.

[0213] The formula for feature fusion is:

[0214] ,

[0215] Where: 、 They are features of different scales, It is the fused feature map.

[0216] Step 4, crack location and classification: Input the fused feature map from step 3 into the target detection algorithm unit, use the Faster R-CNN algorithm or the YOLOv5 algorithm to generate candidate crack boxes and determine the crack locations. Input the crack candidate boxes into the softmax classifier to classify the crack types and output the crack location and type information.

[0217] The calculation formula of the candidate frame of the crack is as follows:

[0218] ,

[0219] Where: and are the coordinates of the upper left and lower right corners of the crack candidate box;

[0220] The position of the crack is determined by using Intersection over Union to measure the overlap between the candidate box and the actual crack. The formula is as follows:

[0221] ,

[0222] When the IoU is greater than the set threshold, the candidate box is determined to be the crack area.

[0223] Use the Softmax classifier to classify the crack types. The classifier is trained using the features extracted by CNN and outputs the crack type. The function formula of the Softmax classifier is:

[0224] ,

[0225] Among them, e represents a natural constant, which is approximately equal to 2.71828. It is the base of the exponential function and is used to calculate the exponential value; Indicates that the crack belongs to The class score is used to normalize the probability of the entire class; Indicates the Unnormalized probability values ​​of the classes, used for summing; Is the crack belong to Class score; Indicates the Unnormalized probability values ​​of the classes; Represents input features Next, the crack belongs to class probability; is the total number of crack classification categories.

[0226] Step 5, Crack Boundary Refinement: The crack location and type information from Step 4 is fed into the dynamic attention pooling unit. The self-attention mechanism in the Vision Transformer calculates the attention weight for each pixel, dynamically enhancing the features of the crack region, including the crack boundary, and improving the accuracy of the crack boundary. The dynamically enhanced crack feature information is fed into the boundary refinement algorithm unit, which uses a convolutional neural network submodule and morphological operations to optimize the crack boundary at the pixel level, generating refined crack boundary data.

[0227] The specific steps are as follows:

[0228] (1) Dynamic Attention Pooling

[0229] The dynamic attention mechanism in ViT is used to enhance the features of crack regions in the image. The weights of crack regions are calculated through the self-attention mechanism, and the model is made to focus on the crack boundaries.

[0230] During crack detection, traditional convolutional neural networks (CNNs) may not be able to accurately identify crack boundaries, especially when dealing with cracks with fuzzy or unclear boundaries. Therefore, this example introduces a dynamic attention mechanism based on the Vision Transformer (ViT) to enhance the features of crack regions in images and refine the crack boundaries.

[0231] Self-Attention: In ViT, the self-attention mechanism focuses on important areas by calculating the correlation between each position (pixel) in the image and other positions. In crack images, the crack boundaries are often the most critical areas. The self-attention mechanism can give higher weight to the boundary areas, strengthening the features of these areas and thus improving the accuracy of crack detection.

[0232] The calculation process of the self-attention mechanism is as follows:

[0233] For the input image , first converts it into a query, key, and value matrix through a linear transformation, representing the features of each pixel in the image. Calculate the self-attention weight of each pixel and then focus on the crack area. The specific process is as follows:

[0234] ①Calculate the query, key, and value

[0235] For the input image , query, key and value matrices are obtained through linear transformation:

[0236] , , ,

[0237] Where: 、 、 are the weight matrices used to generate queries, keys, and values, respectively.

[0238] ②Attention weight calculation

[0239] The query matrix Q and the key matrix K are used to calculate the relevance of each position in the image, and then the attention weight is obtained through the softmax function:

[0240] ,

[0241] Where: is the query matrix; is the bond matrix; is a value matrix; It is a scaling factor that prevents the dot product result from being too large and causing the gradient to disappear; is the dimension of the key matrix, ensuring computational stability; T represents the matrix transpose operation.

[0242] ③Dynamically adjust weights

[0243] To adapt to the crack characteristics in different images, the attention weights are dynamically adjusted according to changes in image content. For example, in low-contrast or noisy images, the model can dynamically increase the attention weight on the crack boundary area to reduce the influence of background noise.

[0244] For each pixel position of the image , the formula for dynamically calculating the attention weight of each pixel based on its correlation with other pixels is as follows:

[0245] ,

[0246] Where: It's a pixel Pixel The attention weight of Represents pixels With pixels similarity between Represents pixels With pixels similarity between Indicates the total number of pixels or feature vectors in the image; 、 、 Indicates the image 、 and Feature representation of pixels; Indicates the sum of all pixels in the image for normalization; Represents an exponential function used to enhance differences and make items with high similarity more prominent.

[0247] In this way, the model can "adaptively" assign higher weights to crack regions, especially when the crack boundaries are blurred or the background is complex.

[0248] ④ Enhance the characteristics of the crack area

[0249] Dynamic Attention Pooling: We use the calculated self-attention weights for each pixel to perform weighted pooling on the crack areas in the image. During dynamic attention pooling, the model focuses on the crack areas through weighted calculations, enhancing the features of the crack boundaries:

[0250] ,

[0251] Where: Represents the image feature map after attention weighting. The features of the boundary area will be further amplified and focused on the crack part.

[0252] The boundary refinement algorithm unit further refines the crack boundary through morphological operations, including erosion and dilation, to remove blurred edges caused by image noise or low resolution. The formula for the erosion operation is as follows:

[0253] ,

[0254] Where: is the image at position after corrosion Pixel value, output image; It means taking the minimum value of all pixel values ​​in the area covered by the structural element to achieve the "erosion" effect, that is, shrinking the boundary; represents the original input image; Represents the original image I in the structural element The pixel position under coverage The corresponding value; It is a structural element that defines the shape and extent of corrosion; 、 Represents corroded structural elements The offset of each pixel relative to the center point.

[0255] Step 6, Structural Integrity Assessment: Refined crack boundary data is generated from Step 5, and crack parameters are extracted from the crack image. These crack parameters include length, width, depth, location, and type. These parameters, combined with the structural load and material properties, provide the necessary input data for structural integrity assessment. Based on the extracted crack parameters, finite element analysis (FEA) is used to simulate the impact of cracks on the structure and calculate stresses in the cracked area. Based on the crack parameters and structural stress analysis results, a risk assessment model, using a support vector machine (SVM) or random forest (RF), is trained to perform structural health assessment and calculate a crack risk score. This score comprehensively considers multiple crack dimensions and structural response, providing a quantitative indicator of the initial risk posed by the crack to the structure. Finally, based on the crack risk score and stress conditions, the structural health status and corresponding repair recommendations are output. In cases where data is scarce or labeling is difficult, a self-supervised learning process is implemented.

[0256] The specific steps are as follows:

[0257] 1. Crack parameter extraction:

[0258] The refined crack boundary data is generated from step 5, and crack parameters are extracted from the crack image. The crack parameters include the length, width, depth, location, and type of the crack, which are used for subsequent structural health assessment. The specific extraction method is as follows:

[0259] Crack Length: Calculate the total length of the cracks through the crack detection frame.

[0260] Crack Width: Calculate the crack width based on image edge detection or deep learning models.

[0261] Crack depth: Get crack depth information based on 3D imaging or sensor data.

[0262] Crack location: Use crack detection algorithms to locate the specific location of cracks.

[0263] Crack type: Crack types are classified using a classification model.

[0264] 2. Structural stress analysis

[0265] Combined with the load conditions and material properties of the structure, based on the extracted crack parameters, finite element analysis (FEA) is used to simulate the impact of cracks on the structure.

[0266] (1) Stress analysis: Based on the location, type and size of the cracks, simulate the effect of the cracks on the structural stress. The formula for calculating the stress in the crack area is as follows:

[0267] ,

[0268] Where: is the stress in the crack area; is the externally applied load; is the effective cross-sectional area of ​​the crack region.

[0269] (2) Structural strength assessment: Combined with the stress analysis results of the crack area, assess whether the structure will fail or produce excessive deformation due to the cracks.

[0270] 3. Structural health assessment

[0271] (1) Crack risk assessment

[0272] Based on the crack parameters and structural stress analysis results, support vector machines (SVM) or random forests (RF) are used to perform crack risk assessment, predict the potential threat of cracks to the structure, and output the structural risk level and corresponding repair recommendations. The specific steps are as follows:

[0273] 1) Training of risk assessment models

[0274] Crack parameters and structural health characteristics are trained using support vector machines (SVMs), random forests (RFs), or other machine learning methods. Crack parameters include length, width, type, and depth, while structural health characteristics include stress analysis results. During the training process, crack data is combined with structural performance data to predict the potential risk of cracks to the structure.

[0275] SVM training formula:

[0276] ,

[0277] Where: is the weight of the decision boundary, is the slack variable, is the regularization parameter that controls the classification error.

[0278] The training dataset contains information such as crack type, location, size, and damage level.

[0279] The formula for outputting the risk level of cracks is as follows:

[0280] ,

[0281] Where: Score the risk of cracks; represents a comprehensive risk scoring function, which is used to calculate the final risk score based on multiple input fracture parameters and model output; Output represents the crack damage grade result output by the support vector machine, which is used to reflect the severity of the crack and is a key input parameter of the crack risk scoring function; Indicates the length characteristics of the crack; Represents the width characteristics of the crack.

[0282] 2) Health status prediction

[0283] Based on the crack characteristics and the structural stress analysis results, the health status of the structure is output. The risk assessment model can identify the long-term impact of cracks on the structure and determine areas requiring repair or reinforcement based on the assessment results.

[0284] The trained model is combined with parameters such as crack length, width, type, depth, and structural stress analysis results to perform structural health assessment.

[0285] 3) Repair suggestions

[0286] Recommendations for repair or reinforcement are given based on the results of the structural health assessment. For example, if the structural risk assessment result is "severe damage," immediate repair or reinforcement is recommended; if the assessment is "minor damage," regular inspections are recommended.

[0287] (4) Self-supervised learning

[0288] When data is scarce or difficult to label, self-supervised learning methods are used to enhance model training and reduce dependence on labeled data. The steps of self-supervised learning are as follows:

[0289] Step 61: Using some manually annotated crack image data, a preliminary training model combining a convolutional neural network (CNN) and a vision transformer (ViT) is trained to extract crack features and perform preliminary classification.

[0290] The preliminary training model in this embodiment refers to the model obtained by supervised training based on some manually annotated crack image data using a convolutional neural network (CNN) and a Vision Transformer (ViT) fusion model structure before performing self-supervised learning. It serves as the baseline model for generating pseudo labels in the self-supervised stage. The training process of the preliminary training model is as follows:

[0291] 1. Data Source and Labeling Method

[0292] (1) Image acquisition equipment

[0293] Use high-resolution industrial cameras (2000×1500 pixels) mounted on drones and bridge inspection robots;

[0294] (2) Collection objects

[0295] Cracks in infrastructure such as road and bridge decks, concrete columns, and tunnel linings;

[0296] (3) Data sample size

[0297] Number of annotated images: 3000;

[0298] Number of unlabeled images: 7,000;

[0299] (3) Marking method

[0300] Use the LabelMe tool to accurately label polygons, including crack location (bounding box) and type (horizontal, vertical, oblique cracks, and mesh);

[0301] (5) Training and validation set division: Divide into training set and validation set at an 8:2 ratio.

[0302] 2. Data Preprocessing Steps

[0303] (1) Denoising:

[0304] Use the Denoising Autoencoder to optimize the reconstruction loss function:

[0305] ,

[0306] Where: is the original crack image pixel, is the reconstructed image after denoising, is the total number of pixels.

[0307] (2) Enhanced operation:

[0308] Histogram equalization enhances contrast, mapping function:

[0309] ,

[0310] Where: is the pixel value after equalization; is the pixel value of the original image; Is the mapping function, representing the original grayscale value To the gray value after equalization Mapping; Indicates an 8-bit grayscale image, and L takes the value of 256; is the probability density function of the original image; Indicates a small change in the integral, an integral operation performed to calculate the cumulative probability;

[0311] The Sobel operator is used for edge enhancement, using the following convolution kernel:

[0312] ,

[0313] ,

[0314] (3) Multi-scale image generation: The maximum pooling operation is used to construct three scale copies. The original scale is 512×512, which is scaled to 256×256 and 128×128.

[0315] ,

[0316] 3. Model Structure (CNN + ViT Fusion)

[0317] (1) Local feature extraction (CNN module)

[0318] Architecture: 3 convolutional layers, ReLU activation function, and 2×2 pooling layer;

[0319] Convolution formula:

[0320] ,

[0321] in, is the median value in the output feature map ( )’s pixel value; is the position in the input image ( )’s pixel value; is the position in the convolution ( )’s weight; Represents the row index of the convolution kernel, which controls the vertical displacement of the convolution kernel in the input image; Represents the column index of the convolution kernel, which controls the horizontal displacement of the convolution kernel in the input image; Indicates the height of the convolution kernel (number of rows); Indicates the width of the convolution kernel (number of columns).

[0322] (2) Global feature extraction (ViT module)

[0323] The input is converted into a sequence through Patch Embedding, with a length of

[0324] Multi-head self-attention mechanism:

[0325] ,

[0326] in: , , : are query, key, and value matrices respectively; The dimension of the key.

[0327] The number of heads is 8 and the number of Transformer encoder layers is 4.

[0328] (3) Feature fusion

[0329] local features With global features cascade:

[0330] ,

[0331] (4) Detection and classification module

[0332] Candidate box extraction: adopt YOLOv5s structure, and IoU calculation is used for non-maximum suppression;

[0333] Crack type classification: Use softmax multi-class classifier:

[0334] , ,

[0335] Where: Represents the input sample Belong to category probability; For category The corresponding model output value; For category The corresponding model output value; Representation category logits The indexed value is calculated in the category The relative importance of meeting the category when the probability of meeting the category is greater than . Indicates all categories logits The exponentially scaled values ​​are summed across all categories in the denominator to ensure that the sum of the probabilities is 1.

[0336] 4. Training Process

[0337] (1) Loss function (joint loss)

[0338] ,

[0339] For the candidate box regression loss, use Smooth L1: ,

[0340] is the cross entropy classification loss;

[0341] Parameter settings: ,

[0342] (2) Optimizer settings

[0343] Optimizer: Adam;

[0344] Initial learning rate: 0.0001;

[0345] Weight decay: 0.0005;

[0346] Momentum coefficient ( , ):(0.9, 0.999);

[0347] Batch Size: 16

[0348] Number of training epochs: 100

[0349] (3) Learning rate scheduling:

[0350] Use the cosine annealing strategy to dynamically adjust the learning rate:

[0351] ,

[0352] Where: is the learning rate corresponding to the current training round (epoch); The initial maximum learning rate, usually the initial learning rate set at the beginning of model training; The minimum learning rate is usually set to a very small value and is used for stable fine-tuning at the end of training; The index number of the current epoch in the total training rounds; is the total number of training rounds, i.e. the maximum number of epochs; is the circumference constant of pi, approximately equal to 3.14159.

[0353] 5. Evaluation Indicators and Training Effects

[0354] The initial model after training has the following performance on the validation set:

[0355] Average detection precision (mAP@0.5): 0.882;

[0356] Classification accuracy: 89.3%;

[0357] Recall rate: 86.7%

[0358] F1-score: 0.874;

[0359] Boundary positioning error (IoU mean): 0.728.

[0360] These results serve as a basis for quality assurance of pseudo-label generation in self-supervised learning.

[0361] Step 62, pseudo label generation: Generate pseudo labels using the preliminarily trained model from step 1. The pseudo labels are predicted based on the similarity of image content. The unlabeled crack images are output with pseudo labels through the prediction of the model.

[0362] Step 63, self-supervised learning loss function: Use the generated pseudo labels for training and calculate the loss function of self-supervised learning to ensure that the error between the pseudo labels and the actual labels is minimized. The loss function of self-supervised learning is defined as:

[0363] ,

[0364] Where: is the pseudo label generated by the model; is the actual label; is the number of training samples. By minimizing the loss function, the learning process of the model is optimized.

[0365] 3) Data augmentation and optimization: Pseudo-labeled and unlabeled data are used for further training and optimization to expand the dataset and improve the model’s generalization capabilities. Even with limited labeled data, the model can still be effectively trained and improve the accuracy of crack detection and structural health assessment.

[0366] The above crack detection and structural assessment method that integrates multi-scale Transformer and CNN is used for crack detection and structural integrity assessment of civil engineering structures, including roads, bridges, tunnels, buildings, dams and hydroelectric dams.

[0367] The crack detection and structural assessment system and method of this embodiment integrates a convolutional neural network (CNN) with a Vision Transformer (ViT) to achieve collaborative extraction of local and global features, significantly improving the ability to identify complex cracks. The system and method adapt to diverse scenarios and scales, making it suitable for real-world engineering applications. This addresses the problem that most existing CNN-based technologies are only good at extracting local features and struggle to capture long-range dependencies.

[0368] By introducing a dynamic attention pooling mechanism, the weights of pixel-level features are dynamically adjusted based on self-attention, highlighting boundary areas and achieving high-precision boundary restoration. This significantly outperforms traditional morphological and U-Net edge detection methods. This solves the problem of fuzzy boundary detection in existing CNN / FCN methods, especially the serious false detection problem in low-resolution images.

[0369] By constructing multi-scale image replicas and a multi-scale feature fusion mechanism, and using cascaded convolutions to handle scale variations, a single model can accommodate both narrow and wide cracks, providing more robust detection capabilities. This addresses the existing problem of single-scale modeling, which prevents unified identification of microcracks and wide cracks and results in a high rate of missed detection.

[0370] By integrating a unified framework for crack identification and structural health risk assessment, using support vector machine (SVM) / radix function (RF) models combined with finite element analysis (FEA), the system delivers more accurate assessment results. It automatically outputs crack levels and structural repair recommendations, making them suitable for decision-making and deployment. This addresses the disconnected nature of existing systems, where most systems focus solely on crack detection or structural assessment.

[0371] By introducing a self-supervised learning mechanism, the system automatically utilizes unlabeled data through pseudo-labeling, improving model training efficiency. This effectively addresses the real-world challenges of data scarcity and labeling in civil engineering, while also reducing deployment costs. This addresses the existing issues of relying heavily on manually labeled data, resulting in high training costs and difficulty in acquiring data.

Claims

1. A crack detection and structural assessment system integrating Transformer and CNN, characterized by: include: An image preprocessing module is used to receive the original crack image, perform denoising, enhancement and scale transformation on the original crack image, and output preprocessed image data; A crack feature extraction module, connected to the image preprocessing module, is used to receive the preprocessed image data and extract local and global features of the crack image by combining the convolutional neural network submodule with the Transformer submodule in the Vision Transformer through multi-scale feature extraction and self-attention mechanism, and output crack feature data; A crack location and classification module, connected to the crack feature extraction module, is used to receive crack feature data, locate crack positions and classify cracks based on the extracted local and global features of the crack image using a target detection algorithm, and output crack position and type information; a crack boundary refinement module, connected to the crack location and classification module, configured to receive crack location and type information, process the located and classified crack areas, refine the crack boundaries using dynamic attention pooling and boundary refinement algorithms, and output refined crack boundary data; The structural integrity assessment module is connected to the crack boundary refinement module and is used to receive the refined crack boundary data. It combines the refined crack boundary feature information with the stress analysis results of the structure, uses a support vector machine or a random forest model to perform structural health assessment, calculates the risk score of the crack, and outputs the health status of the structure and corresponding repair suggestions based on the risk score and stress conditions of the crack. In cases where data is scarce or difficult to label, a self-supervised learning process is performed.

2. The system according to claim 1, wherein: The image preprocessing module includes: A denoising autoencoder, configured to receive an original crack image and remove noise from the original crack image; an image enhancement unit, connected to the denoising autoencoder, configured to receive the denoised crack image and perform histogram equalization and edge enhancement on the crack image; a scale transformation module, connected to the image enhancement unit, configured to receive the enhanced crack image, perform multi-scale processing on the crack image, and obtain image copies with different resolutions; The crack feature extraction module includes: The convolutional neural network submodule is used to extract local features of the crack image, including edges and textures. It uses a multi-layer convolutional neural network architecture, which includes at least one convolution layer, one pooling layer, and one activation layer to capture the subtle structural features of the crack. A Transformer submodule, connected to the convolutional neural network submodule, is used to extract global features of the crack image through a self-attention mechanism, enhance the importance of the crack region in the image, and capture long-range dependencies. The specific implementation includes parallel processing of multiple self-attention heads and an encoder layer based on the Transformer architecture. The encoder layer performs global weighted aggregation of input features through weighted summation, thereby strengthening the global feature representation of the crack region. The crack location and classification module includes: The target detection algorithm unit is used to receive the feature information output by the crack feature extraction module, use the Faster R-CNN algorithm or the YOLOv5 algorithm to generate candidate frames for the cracks, and determine the locations of the cracks; A classifier, connected to the target detection algorithm unit, configured to receive crack candidate frames generated by the target detection algorithm unit and classify crack types using a softmax classifier; The crack boundary refinement module includes: The dynamic attention pooling unit receives the crack area information output by the crack localization and classification module, calculates the attention weight of each pixel through the self-attention mechanism in the Vision Transformer, and dynamically enhances the features of the crack area, including the crack boundary. A boundary refinement algorithm unit, connected to the dynamic attention pooling unit, is used to receive the dynamically enhanced crack feature information and perform pixel-level optimization on the crack boundary using a convolutional neural network submodule and morphological operations; The structural integrity assessment module includes: Crack parameter extraction unit, used to extract the crack length, width, depth, location, and type characteristics from the crack image, and combine them with the load conditions and material properties of the structure; a structural stress analysis module, connected to the crack parameter extraction unit, for simulating the effect of the crack on the structure through finite element analysis based on the extracted crack parameters, and calculating the stress in the crack area; The risk assessment module is connected to the structural stress analysis unit and is used to perform structural health assessment using a support vector machine or a random forest based on the crack parameters and the structural stress analysis results, output the risk level of the structure, and output corresponding repair suggestions based on the risk level of the structure.

3. A crack detection and structural assessment method integrating Transformer and CNN, characterized in that: The following steps are involved: Step 1, image preprocessing step: receiving the original crack image, and performing denoising, enhancement and scale transformation on the original crack image to generate preprocessed image data, including multiple image copies at different scales; Step 2, crack feature extraction step: The image data preprocessed in step 1 is input into the convolutional neural network submodule, and the local features of the cracks, including edges and textures, are extracted through the convolutional layer. The extracted local features are input into the Transformer submodule in the VisionTransformer, and the global features of the cracks are extracted through the self-attention mechanism, which enhances the importance of the crack area in the image and captures long-range dependencies. Features are extracted from the image copies of different scales described in step 1 to obtain features at different scales. Step 3: The features of different scales described in step 2 are fused through cross-layer connection and cascade convolution method to obtain a fused feature map; Step 4, crack location and classification: Input the fused feature map from step 3 into the target detection algorithm unit, use the Faster R-CNN algorithm or the YOLOv5 algorithm to generate candidate crack boxes and determine the crack locations. Input the crack candidate boxes into the softmax classifier to classify the crack types and output the crack location and type information. Step 5, crack boundary refinement step: The crack location and type information described in step 4 is input into the dynamic attention pooling unit. The self-attention mechanism in the Vision Transformer is used to calculate the attention weight of each pixel, dynamically enhancing the features of the crack area, including the crack boundary. The dynamically enhanced crack feature information is input into the boundary refinement algorithm unit. The convolutional neural network submodule and morphological operations are used to optimize the crack boundary at the pixel level to generate refined crack boundary data. Step 6, structural integrity assessment step: Generate refined crack boundary data from step 5 and extract crack parameters from the crack image. The crack parameters include the characteristics of the crack length, width, depth, location and type, and are combined with the load conditions and material properties of the structure; based on the extracted crack parameters, simulate the impact of the crack on the structure through finite element analysis and calculate the stress in the crack area; based on the crack parameters and the structural stress analysis results, use support vector machines or random forests to perform structural health assessment, calculate the crack risk score, and output the structural health status and corresponding repair suggestions based on the crack risk score and stress conditions.

4. The method according to claim 3, characterized in that The denoising of the original crack image described in step 1 is performed using a denoising autoencoder, which is trained by minimizing the following loss function: , Where: is the input image, is the denoised image, is the number of pixels in the image; The original crack image is enhanced by using histogram equalization and edge enhancement technology to enhance image contrast and edge information. The formula for histogram equalization is as follows: , Where: is the pixel value after equalization; is the pixel value of the original image; Is the mapping function, representing the original grayscale value To the gray value after equalization Mapping; Indicates an 8-bit grayscale image, and L takes the value of 256; is the probability density function of the original image; Indicates a small change in the integral, an integral operation performed to calculate the cumulative probability; Edge enhancement uses the Sobel operator to enhance the crack edge, and uses the convolution kernel to detect the crack edge in the image to enhance the edge information. The convolution kernel is as follows: , The image scale transformation of the original crack image is to use a pooling layer to reduce the image size and obtain image copies of multiple scales. The pooling operation formula is as follows: , Where, is the original image, is the pooled copy of the image.

5. The method according to claim 3, characterized in that The convolution operation formula for extracting the local features of the crack in step 2 is as follows: , Where: Represents the position in the output feature map The value of Represents the position in the convolution kernel The weight of Indicates that the input image is located at The pixel value of the position; is the input image, is the convolution kernel, It is the feature map after convolution; m and n represent the offset index within the convolution kernel, which is used to extract the local area by sliding the window in the input image; i and j represent the position coordinates of the previous output feature map; The self-attention mechanism formula for extracting the global features of the crack is as follows: , Where: is the query matrix; is the bond matrix; is a value matrix; is the dimension of the key matrix; the self-attention mechanism helps identify the relationships and patterns of cracks in complex backgrounds; T represents the matrix transpose operation.

6. The method according to claim 3, characterized in that The formula for feature fusion in step 3 is: , Where: 、 They are features of different scales, It is the fused feature map.

7. The method according to claim 3, characterized in that The calculation formula for the candidate box of the crack in step 4 is as follows: , Where: and are the coordinates of the upper left and lower right corners of the crack candidate box; The position of the crack is determined by using Intersection over Union to measure the overlap between the candidate box and the actual crack. The formula is as follows: , The function formula of the Softmax classifier is: , Among them, e represents a natural constant, which is approximately equal to 2.71828. It is the base of the exponential function and is used to calculate the exponential value; Indicates that the crack belongs to The class score is used to normalize the probability of the entire class; Indicates the Unnormalized probability values ​​of the classes, used for summing; Is the crack belong to Class score; Indicates the Unnormalized probability values ​​of the classes; Represents input features Next, the crack belongs to class probability; is the total number of crack classification categories.

8. The method according to claim 3, characterized in that The formula for calculating the attention weight of each pixel described in step 5 is as follows: , Where: It's a pixel Pixel The attention weight of Represents pixels With pixels similarity between Represents pixels With pixels similarity between Indicates the total number of pixels or feature vectors in the image; 、 、 Indicates the image 、 and Feature representation of pixels; Indicates the sum of all pixels in the image for normalization; represents an exponential function, which is used to enhance differences and make items with high similarity more prominent; The boundary refinement algorithm unit further refines the crack boundary through morphological operations. The morphological operations include erosion and dilation. The formula of the erosion operation is as follows: , Where: is the image after corrosion at position Pixel value, output image; It means taking the minimum value of all pixel values ​​in the area covered by the structural element to achieve the "erosion" effect, that is, shrinking the boundary; represents the original input image; Represents the original image I in the structural element The pixel position under coverage The corresponding value; It is a structural element that defines the shape and extent of corrosion; 、 Represents corroded structural elements The offset of each pixel relative to the center point.

9. The method according to claim 3, characterized in that The formula for calculating the stress in the crack area described in step 6 is as follows: , Where: is the stress in the crack area; is the externally applied load; is the effective cross-sectional area of ​​the crack region; The calculation formula of the risk score of the crack is as follows: , Where: Score the risk of cracks; represents a comprehensive risk scoring function, which is used to calculate the final risk score based on multiple input fracture parameters and model output; Output represents the crack damage grade result output by the support vector machine, which is used to reflect the severity of the crack and is a key input parameter of the crack risk scoring function; Indicates the length characteristics of the crack; Represents the width characteristics of the crack.

10. The method according to claim 3, characterized in that The structural integrity assessment step in step 6 also includes performing a self-supervised learning process when data is scarce or difficult to label. The self-supervised learning process includes: Step 61: using some manually annotated crack image data to train a preliminary training model combining a convolutional neural network and a VisionTransformer; Step 62, generating pseudo labels: using the preliminarily trained model in step 1 to predict the unlabeled crack images and generate pseudo labels; Step 63, calculating the loss function of self-supervised learning: using the pseudo labels for training, and calculating the loss function of self-supervised learning to ensure that the error between the pseudo labels and the actual labels is minimized. The loss function of self-supervised learning is defined as: , Where: is the pseudo label generated by the model; is the actual label; is the number of training samples; by minimizing the loss function, the learning process of the model is optimized; Step 64, data expansion and optimization: Use pseudo-labeled and unlabeled data for further training and optimization to expand the dataset.

Citation Information

Cited By

  • Pavement crack detection equipment based on double spectrums

    CN121095246A

  • Method and system for detecting defects of annular concrete pole body

    CN121236034A

  • Tunnel environment safety risk assessment method and system based on deep learning

    CN121365875A

  • Tunnel crack detection method based on image recognition

    CN121415265A

  • River bank collapse identification method and system based on multi-feature progressive detection and fusion verification

    CN121708476A