Bridge crack identification and automatic evaluation method based on image identification and AI modeling

By combining the CPSeg segmentation model and the multi-layer perceptron evaluation model, high-precision identification of bridge cracks and structural risk assessment are achieved, solving the problems of low efficiency and insufficient accuracy in existing technologies and providing an efficient and reliable bridge health monitoring solution.

CN120689308AInactive Publication Date: 2025-09-23TAIZHOU UNIV
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510784878.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing bridge crack detection methods rely on manual inspections, which are inefficient and prone to misjudgment. In addition, existing image processing algorithms lack recognition accuracy in complex environments, making it difficult to achieve high-precision crack identification and structural risk assessment.

Method used

By adopting the context-aware CPSeg segmentation model and the multi-layer perceptron evaluation model, an integrated process for bridge crack identification and risk assessment is constructed. Through image preprocessing, pixel-level segmentation, geometric feature extraction and supervised learning, automated assessment from crack images to structural risk levels is achieved.

Benefits of technology

It achieves high-precision positioning of bridge cracks and objective assessment of structural risk levels, improves recognition accuracy and the reliability of assessment results, and is suitable for bridge health monitoring and intelligent operation and maintenance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689308A_ABST
    Figure CN120689308A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge crack identification and automatic evaluation method based on image identification and AI modeling, and the method comprises the following steps: S1, obtaining an original image of a bridge structure surface, and carrying out the image preprocessing; s2, inputting the standardized image into an image recognition model, performing pixel-level segmentation on a crack region in the image, and outputting a crack mask graph; s3, performing feature extraction processing on the crack mask graph, extracting geometric feature parameters of the crack, and constructing a crack feature vector; s4, constructing an evaluation model based on a supervised learning method, and training the evaluation model; and S5, inputting the crack feature vector into an evaluation model, evaluating the structural risk level of the crack, and outputting a structural risk label. According to the method, image recognition and AI modeling are fused, automatic crack recognition and evaluation are achieved, and the method has the advantages of being high in precision, clear in boundary and intelligent in evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of civil engineering structure health monitoring, and in particular to a bridge crack identification and automatic assessment method based on image recognition and AI modeling. Background Art

[0002] With the continuous development of urban infrastructure construction, a large number of bridge structures have entered the middle and late stages of service. Their safety status is directly related to the safety of people's lives and property and transportation efficiency. Cracks are one of the common forms of bridge damage, reflecting the risks of fatigue, aging or local instability in the structure. Therefore, timely identification of cracks and structural risk assessment are of great significance.

[0003] Traditional bridge crack detection methods mainly rely on manual inspections or taking photos with portable devices and then conducting manual analysis. Such methods are not only limited by the experience and subjective judgment of inspectors, prone to missed detections or misjudgments, but also inefficient and with limited coverage, making it difficult to meet the engineering needs of large-scale and frequent monitoring. With the development of computer vision technology, researchers have begun to try to use image processing methods for crack identification, such as edge detection and threshold segmentation. However, these traditional algorithms have unstable performance in environments with uneven lighting and severe background interference, and are particularly prone to false detections in the presence of fine cracks or irregular texture interference.

[0004] In recent years, deep learning-based image segmentation technology has provided a new solution for structural defect identification. Convolutional neural networks have shown high accuracy in crack identification. However, existing models focus more on overall accuracy and ignore the refined expression of crack boundaries, resulting in poor accuracy of segmentation results in fuzzy boundary areas. At the same time, existing methods often only focus on crack identification and fail to further convert the identification results into geometric features or risk levels that can be used for structural safety assessment, lacking the evaluation capabilities for practical engineering decision-making.

[0005] Therefore, there is an urgent need for a closed-loop automated method from crack image recognition, fine boundary segmentation to structural risk level assessment, which can not only improve the recognition accuracy but also meet the requirements of engineering applications for result credibility and alignment with specifications. Summary of the Invention

[0006] One purpose of the present invention is to propose a method for bridge crack identification and automatic assessment based on image recognition and AI modeling. The present invention integrates image recognition and AI modeling technologies, adopts the context-aware CPSeg segmentation model and the multi-layer perceptron assessment model, and constructs an integrated process for crack identification and risk assessment. It can achieve high-precision positioning of bridge cracks, geometric feature extraction and structural risk level judgment, and has the advantages of high recognition accuracy, clear boundaries, and objective and reliable assessment results. It is suitable for bridge health monitoring and intelligent operation and maintenance scenarios.

[0007] According to an embodiment of the present invention, a method for identifying and automatically assessing bridge cracks based on image recognition and AI modeling includes the following steps:

[0008] S1. Obtaining an original image of the bridge structure surface and performing image preprocessing on the original image to obtain a standardized image;

[0009] S2. Input the standardized image into an image recognition model. The image recognition model is built based on a context-aware pyramid structure, adopts the CPSeg architecture, and introduces a boundary optimization module to perform pixel-level segmentation on the crack areas in the image and output a crack mask map containing the spatial distribution and boundary information of the cracks.

[0010] S3, performing feature extraction processing on the crack mask image, extracting geometric feature parameters of the crack, and constructing them into a crack feature vector;

[0011] S4. Build an evaluation model based on supervised learning methods and train the evaluation model;

[0012] S5. Input the crack feature vector into the evaluation model, evaluate the structural risk level of the crack, and output a structural risk label.

[0013] Optionally, the image preprocessing includes image distortion correction, grayscale conversion, denoising, brightness equalization and size normalization.

[0014] Optionally, the S2 specifically includes:

[0015] S21. Input the standardized image into an image recognition model. The image recognition model adopts a CPSeg model based on context perception and pyramid structure. The CPSeg model includes a backbone feature extraction network, a context perception module, a pyramid decoding module, an edge detection module, and a boundary optimization module to achieve pixel-level segmentation of the crack area.

[0016] S22. Performing multi-layer convolution operations on the standardized image through a backbone feature extraction network to extract semantic feature maps at different levels, wherein the semantic feature maps have different spatial resolutions and semantic depths;

[0017] S23. Input the semantic feature map into the context perception module, perform multi-scale spatial pyramid pooling to obtain global context information, and output the context feature map:

[0018] P l (x,y)=Concat(Pool s1 (F l ),Pool s2 (F l ),...,Pool sN (Fl ));

[0019] Among them, P l (x,y) layer l, Pool si Indicates that the pooling kernel size is s i The spatial pooling operation, s i represents the pooling scale of the i-th layer, N represents the number of pooling layers, Concat represents the concatenation function on the channel dimension, P l (x,y) represents the fused context feature map, F l Represents the feature map extracted by the lth layer, (x, y) represents the pixel coordinates;

[0020] S24, input all context feature maps into the pyramid decoding module, perform upsampling and layer-by-layer feature fusion operations, combine the semantic feature maps extracted from the backbone feature extraction network, restore spatial details through skip connections, and output a semantic prediction map;

[0021] S25. Input the semantic prediction map into the edge detection module, and calculate the boundary prediction map by gradient amplitude calculation and convolution operation based on the semantic prediction map:

[0022]

[0023] in, Represents the boundary prediction graph, describing the boundary probability predicted by the model, Represents the predicted probability map, GradMag represents the execution of the Sobel operator to obtain the gradient magnitude, Conv 1×1 represents a 1×1 convolution operation, and σ represents a sigmoid activation function;

[0024] S26. Input the semantic prediction map into the boundary optimization module. The boundary optimization module includes an edge-guided supervision mechanism and an edge-sensitive loss function. The edge-guided supervision mechanism constructs a boundary-guided loss function based on the boundary supervision label map and the boundary prediction map. The boundary supervision label map is calculated by the Sobel operator through the real mask map to obtain the gradient amplitude, and compared with the set threshold to obtain a binary boundary map:

[0025]

[0026] in, represents the boundary guidance loss function, E(x,y) represents the boundary supervision label map, which describes whether the pixel point (x,y) is the real boundary;

[0027] S27. Compare the semantic prediction map with the real mask map of the crack at the same time and construct an edge-sensitive loss function:

[0028]

[0029] in, Represents the edge-sensitive loss function, describing the edge-weighted mean square error loss value, w(x,y) represents the edge weight coefficient of the pixel point (x,y), S * (x, y) represents the true mask image, which describes whether the pixel belongs to the crack area. The value is 1, which means it belongs to the crack area.

[0030] S28, forming a total loss function by weighted combination of the main segmentation loss function, the boundary guidance loss function, and the edge sensitive loss function, wherein the main segmentation loss function uses a cross entropy loss function to evaluate the overall classification error between the semantic prediction map and the true mask map;

[0031] S29. The image recognition model trained using the total loss function receives any standardized image in the inference phase, outputs a semantic prediction map, and compares it with the set threshold to generate a crack mask map:

[0032]

[0033] Among them, M(x,y) represents the crack mask map, which describes the classification result of the pixel point (x,y), 1 represents the crack area, 0 represents the non-crack area, and T represents the fixed threshold.

[0034] Optionally, the backbone feature extraction network uses a ResNet50 network.

[0035] Optionally, the semantic prediction map describes the probability value of a pixel being predicted as a crack.

[0036] Optionally, the geometric characteristic parameters include crack length, crack width, crack direction, crack density and crack curvature.

[0037] Optionally, the S3 specifically includes:

[0038] S31, obtaining a crack mask map, and performing connected domain analysis on the crack mask map using the 8-neighborhood method, realizing crack area identification through pixel clustering, scanning all pixels point by point, and marking the pixel with the label of the neighborhood if the current pixel value is 1 and there is a labeled pixel in the 8-neighborhood. Otherwise, a new neighborhood label is created, and a connected domain label map is output, which indicates the label number of the crack to which each pixel belongs;

[0039] S32: For each crack region, perform crack feature extraction operations, including:

[0040] Based on the crack mask image, use the Zhang-Suen algorithm to output a single-pixel-width skeleton image and traverse all pixels with a value of 1 in the skeleton image to generate a skeleton point set for each crack mask image. The skeleton image is a binary image of the same size as the crack mask image.

[0041] Define the crack length as the maximum Euclidean distance between skeleton points and calculate the crack length:

[0042]

[0043] Among them, L i represents the length of the i-th crack, S i represents the skeleton point set, p and q represent any two points in the skeleton point set, x p and x q Indicates the horizontal coordinates of the skeleton points p and q, y p and y q Indicates the ordinate of the skeleton points p and q, and max indicates the maximum value;

[0044] The crack width is obtained by calculating the average minimum distance from the skeleton point to the boundary point:

[0045]

[0046] Among them, W i represents the width of the i-th crack, |S i | represents the total number of skeleton points, x b and y b represents the coordinates of the boundary point b;

[0047] The coordinates of the skeleton point set are combined into a sample matrix, and the covariance matrix of the sample matrix is ​​decomposed by eigenvalue, and the angle corresponding to the first principal component direction vector is selected as the crack direction:

[0048] θ i =arctan2(v y ,v x );

[0049] Among them, θ i represents the main direction angle of the i-th crack, v x and v y They represent the horizontal and vertical components of the direction vector of the first principal component, respectively, and arctan represents the inverse tangent function;

[0050] The crack density is calculated based on the ratio of the total number of crack pixels to the area of ​​the minimum bounding rectangle:

[0051]

[0052] Among them, D irepresents the i-th crack density, |M i | represents the number of pixels in the crack area, A i The area of ​​the minimum circumscribed rectangle representing the crack area;

[0053] Calculate the crack curvature, which is defined as the mean of the angle changes between adjacent points on the skeleton path:

[0054]

[0055] Among them, C i represents the crack curvature, arccos represents the inverse cosine function, and p j+1 represents the j+1th skeleton point, p j represents the j-th skeleton pixel, p j-1 represents the j-1th skeleton pixel;

[0056] S33. Normalize the crack length, crack width, crack direction, crack density, and crack curvature and then concatenate them into a crack feature vector.

[0057] Optionally, the S4 specifically includes:

[0058] S41, constructing an evaluation model, wherein the evaluation model is based on a multi-layer perceptron network;

[0059] S42. Using historical bridge crack data with crack geometric parameters and corresponding risk labels as training samples, the assessment model is trained through supervised learning, wherein the risk labels include normal structure, needing monitoring, and high risk;

[0060] S43. During the model training process, the validation set is used to adjust the hyperparameters to obtain a trained evaluation model.

[0061] The beneficial effects of the present invention are:

[0062] First, the present invention provides a bridge crack identification and automatic assessment method based on image recognition and AI modeling, which can realize an end-to-end automated process from crack image acquisition, pixel-level identification, geometric parameter extraction to structural risk level assessment, solving the problems of low crack identification accuracy, blurred boundaries, insufficient utilization of feature information, and reliance on human subjective judgment in the existing technology.

[0063] Secondly, the present invention introduces the CPSeg model based on the context-aware pyramid structure and designs modules such as backbone feature extraction, context awareness, pyramid decoding, edge detection and boundary optimization. This method significantly improves the ability to recognize crack areas in the image segmentation stage, especially under conditions of complex background and weak texture interference, while still having a strong boundary positioning capability, effectively ensuring the accuracy and integrity of the crack mask map.

[0064] Furthermore, based on crack identification, this method extracts the crack skeleton through connected domain analysis and the Zhang-Suen algorithm, and further calculates geometric features such as crack length, width, direction, density, and curvature to construct a high-dimensional crack feature vector, providing quantitative and objective basic information for structural risk analysis. Compared with traditional methods that only perform identification or manual assessment, this method effectively improves the scientific nature and consistency of risk assessment through standardized feature representation and AI modeling, avoiding problems such as manual misjudgment and missed judgment. In particular, the multi-layer perceptron assessment model, after supervised learning training, can accurately map crack features into structural risk level labels, providing intelligent output results that can be used as a reference for maintenance decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0066] Figure 1 This is a flow chart of a bridge crack identification and automatic assessment method based on image recognition and AI modeling proposed by the present invention;

[0067] Figure 2 This is a structural block diagram of the CPSeg image recognition model of the bridge crack identification and automatic assessment method based on image recognition and AI modeling proposed in the present invention;

[0068] Figure 3 This is a diagram of the connected domain identification and skeleton extraction process of a single crack area in a crack mask image of a bridge crack identification and automatic assessment method based on image recognition and AI modeling proposed by the present invention. DETAILED DESCRIPTION

[0069] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0070] refer to Figure 1-3 A bridge crack identification and automatic assessment method based on image recognition and AI modeling includes the following steps:

[0071] S1. Obtaining an original image of the bridge structure surface and performing image preprocessing on the original image to obtain a standardized image;

[0072] S2. Input the standardized image into an image recognition model. The image recognition model is built based on a context-aware pyramid structure, adopts the CPSeg architecture, and introduces a boundary optimization module to perform pixel-level segmentation on the crack areas in the image and output a crack mask map containing the spatial distribution and boundary information of the cracks.

[0073] S3, performing feature extraction processing on the crack mask image, extracting geometric feature parameters of the crack, and constructing them into a crack feature vector;

[0074] S4. Build an evaluation model based on supervised learning methods and train the evaluation model;

[0075] S5. Input the crack feature vector into the evaluation model, evaluate the structural risk level of the crack, and output a structural risk label.

[0076] This invention provides a method for bridge crack identification and automatic assessment based on image recognition and AI modeling. It implements a complete process from crack image acquisition, processing, segmentation, geometric feature extraction, to structural risk level determination, establishing a technical approach that integrates identification and assessment. This method achieves fully automated detection and intelligent assessment, significantly improving the accuracy of crack identification and the objectivity of assessment. It overcomes the low efficiency and subjective judgment inherent in traditional manual inspections, and boasts a comprehensive process flow, strong adaptability to complex scenarios, and a clear model structure.

[0077] In this embodiment, the image preprocessing includes image distortion correction, grayscale conversion, denoising, brightness balance and size normalization.

[0078] The present invention effectively improves the quality and standardization of image data by introducing operations such as image distortion correction, grayscale conversion, denoising, brightness equalization and size normalization in the image preprocessing stage, ensuring the input consistency and robustness of the subsequent image recognition model. This preprocessing process can effectively alleviate the problem of image quality fluctuations caused by factors such as acquisition equipment, lighting conditions, and background interference, thereby improving the accuracy and stability of model recognition from the source, enabling the image recognition system to still have good adaptability and robustness in complex environments, and improving the reliability of the overall method.

[0079] In this embodiment, S2 specifically includes:

[0080] S21. Input the standardized image into an image recognition model. The image recognition model adopts a CPSeg model based on context perception and pyramid structure. The CPSeg model includes a backbone feature extraction network, a context perception module, a pyramid decoding module, an edge detection module, and a boundary optimization module to achieve pixel-level segmentation of the crack area.

[0081] S22. Performing multi-layer convolution operations on the standardized image through a backbone feature extraction network to extract semantic feature maps at different levels, wherein the semantic feature maps have different spatial resolutions and semantic depths;

[0082] S23. Input the semantic feature map into the context perception module, perform multi-scale spatial pyramid pooling to obtain global context information, and output the context feature map:

[0083] P l (x,y)=Concat(Pool s1 (F l ),Pool s2 (F l ),...,Pool sN (F l ));

[0084] Among them, P l (x,y) layer l, Pool si Indicates that the pooling kernel size is s i The spatial pooling operation, s i represents the pooling scale of the i-th layer, N represents the number of pooling layers, Concat represents the concatenation function on the channel dimension, P l (x,y) represents the fused context feature map, F l Represents the feature map extracted by the lth layer, (x, y) represents the pixel coordinates;

[0085] S24, input all context feature maps into the pyramid decoding module, perform upsampling and layer-by-layer feature fusion operations, combine the semantic feature maps extracted from the backbone feature extraction network, restore spatial details through skip connections, and output a semantic prediction map;

[0086] S25. Input the semantic prediction map into the edge detection module, and calculate the boundary prediction map by gradient amplitude calculation and convolution operation based on the semantic prediction map:

[0087]

[0088] in, Represents the boundary prediction graph, describing the boundary probability predicted by the model, Represents the predicted probability map, GradMag represents the execution of the Sobel operator to obtain the gradient magnitude, Conv 1×1 represents a 1×1 convolution operation, and σ represents a sigmoid activation function;

[0089] S26. Input the semantic prediction map into the boundary optimization module. The boundary optimization module includes an edge-guided supervision mechanism and an edge-sensitive loss function. The edge-guided supervision mechanism constructs a boundary-guided loss function based on the boundary supervision label map and the boundary prediction map. The boundary supervision label map is calculated by the Sobel operator through the real mask map to obtain the gradient amplitude, and compared with the set threshold to obtain a binary boundary map:

[0090]

[0091] in, represents the boundary guidance loss function, E(x,y) represents the boundary supervision label map, which describes whether the pixel point (x,y) is the real boundary;

[0092] S27. Compare the semantic prediction map with the real mask map of the crack at the same time and construct an edge-sensitive loss function:

[0093]

[0094] in, Represents the edge-sensitive loss function, describing the edge-weighted mean square error loss value, w(x,y) represents the edge weight coefficient of the pixel point (x,y), S * (x, y) represents the true mask image, which describes whether the pixel belongs to the crack area. The value is 1, which means it belongs to the crack area.

[0095] S28, forming a total loss function by weighted combination of the main segmentation loss function, the boundary guidance loss function, and the edge sensitive loss function, wherein the main segmentation loss function uses a cross entropy loss function to evaluate the overall classification error between the semantic prediction map and the true mask map;

[0096] S29. The image recognition model trained using the total loss function receives any standardized image in the inference phase, outputs a semantic prediction map, and compares it with the set threshold to generate a crack mask map:

[0097]

[0098] Among them, M(x,y) represents the crack mask map, which describes the classification result of the pixel point (x,y), 1 represents the crack area, 0 represents the non-crack area, and T represents the fixed threshold.

[0099] The present invention adopts the CPSeg model in the image recognition stage and constructs a structure consisting of a backbone feature extraction network, a context perception module, a pyramid decoding module, an edge detection module and a boundary optimization module. It can achieve pixel-level precise segmentation of crack areas. By introducing context perception and multi-scale fusion mechanisms, the model's recognition ability in complex backgrounds and low-contrast crack scenes is effectively improved. This structure is particularly suitable for detecting small cracks and irregular morphological diseases. The recognition results have the advantages of clear boundaries and complete structure, laying a high-quality data foundation for subsequent feature extraction and evaluation.

[0100] In this embodiment, the backbone feature extraction network uses the ResNet50 network.

[0101] This paper uses ResNet50 as the backbone feature extraction network in the image recognition model. Leveraging its deep residual structure and multi-scale convolution capabilities, it can effectively extract high-level semantic information and underlying spatial structural features of crack regions in images. Compared to shallow network structures, ResNet50 has deeper network layers and stronger feature expression capabilities, which can improve the recognition accuracy and stability of the model while maintaining computational efficiency. The multi-layer feature maps extracted by this network can more comprehensively reflect the edge contours and morphological changes of cracks, providing a solid feature foundation for subsequent context-aware fusion and boundary optimization, and enhancing the model's adaptability in complex scenarios.

[0102] In this embodiment, the semantic prediction map describes the probability value of a pixel being predicted as a crack.

[0103] This method uses a semantic prediction map to output the probability of each pixel belonging to a crack area, rather than a simple binary classification label. This effectively improves the flexibility and interpretability of the model output. During the actual inference phase, users can flexibly adjust the sensitivity of the crack mask map based on a set threshold to meet the requirements of different engineering safety levels. This prediction method retains the classification confidence information of the pixel points, facilitating subsequent more refined post-processing operations such as boundary purification and weak-edge crack identification. This method enhances the quantitative expression of the model results and significantly improves the engineering adaptability and intelligent decision-making support capabilities of the bridge crack identification system.

[0104] In this embodiment, the geometric characteristic parameters include crack length, crack width, crack direction, crack density and crack curvature.

[0105] This invention achieves a comprehensive characterization of the spatial structural characteristics of cracks by defining five types of geometric characteristic parameters: crack length, width, direction, density, and curvature. This feature set covers multiple dimensions of crack morphology, including the main axis direction, area percentage, complexity, and direction. It can accurately reflect the development trend and potential risks of cracks. The extracted geometric parameters are all derived from the actual pixel positions of the mask image and skeleton image, with a high degree of physical interpretability and stability, and can serve as high-quality input for structural risk assessment models. In this way, a quantitative bridge is established between crack image data and structural safety levels, providing a solid foundation for subsequent AI modeling and evaluation.

[0106] In this embodiment, S3 specifically includes:

[0107] S31, obtaining a crack mask map, and performing connected domain analysis on the crack mask map using the 8-neighborhood method, realizing crack area identification through pixel clustering, scanning all pixels point by point, and marking the pixel with the label of the neighborhood if the current pixel value is 1 and there is a labeled pixel in the 8-neighborhood. Otherwise, a new neighborhood label is created, and a connected domain label map is output, which indicates the label number of the crack to which each pixel belongs;

[0108] S32: For each crack region, perform crack feature extraction operations, including:

[0109] Based on the crack mask image, use the Zhang-Suen algorithm to output a single-pixel-width skeleton image and traverse all pixels with a value of 1 in the skeleton image to generate a skeleton point set for each crack mask image. The skeleton image is a binary image of the same size as the crack mask image.

[0110] Define the crack length as the maximum Euclidean distance between skeleton points and calculate the crack length:

[0111]

[0112] Among them, L i represents the length of the i-th crack, S i represents the skeleton point set, p and q represent any two points in the skeleton point set, x p and x q Indicates the horizontal coordinates of the skeleton points p and q, y p and y q Indicates the ordinate of the skeleton points p and q, and max indicates the maximum value;

[0113] The crack width is obtained by calculating the average minimum distance from the skeleton point to the boundary point:

[0114]

[0115] Among them, W irepresents the width of the i-th crack, |S i | represents the total number of skeleton points, x b and y b represents the coordinates of the boundary point b;

[0116] The coordinates of the skeleton point set are combined into a sample matrix, and the covariance matrix of the sample matrix is ​​decomposed by eigenvalue, and the angle corresponding to the first principal component direction vector is selected as the crack direction:

[0117] θ i =arctan2(v y ,v x );

[0118] Among them, θ i represents the main direction angle of the i-th crack, v x and v y They represent the horizontal and vertical components of the direction vector of the first principal component, respectively, and arctan represents the inverse tangent function;

[0119] The crack density is calculated based on the ratio of the total number of crack pixels to the area of ​​the minimum bounding rectangle:

[0120]

[0121] Among them, D i represents the i-th crack density, |M i | represents the number of pixels in the crack area, A i The area of ​​the minimum circumscribed rectangle representing the crack area;

[0122] Calculate the crack curvature, which is defined as the mean of the angle changes between adjacent points on the skeleton path:

[0123]

[0124] Among them, C i represents the crack curvature, arccos represents the inverse cosine function, and p j+1 represents the j+1th skeleton point, p j represents the j-th skeleton pixel, p j-1 represents the j-1th skeleton pixel;

[0125] S33. Normalize the crack length, crack width, crack direction, crack density, and crack curvature and then concatenate them into a crack feature vector.

[0126] This method extracts geometric feature parameters from crack mask images, including crack length, width, orientation, density, and curvature, to establish a feature vector for structural risk analysis. Using connected domain analysis and the Zhang-Suen skeleton extraction method, the structural characteristics of each crack region are precisely captured, transforming pixel-level recognition results into high-level semantic information. This feature extraction method offers the advantages of high accuracy and adaptability, and the extracted results are quantifiable and interpretable, providing high-quality input for subsequent AI modeling and enhancing the scientificity and consistency of risk assessments.

[0127] In this embodiment, the S4 specifically includes:

[0128] S41, constructing an evaluation model, wherein the evaluation model is based on a multi-layer perceptron network;

[0129] S42. Using historical bridge crack data with crack geometric parameters and corresponding risk labels as training samples, the assessment model is trained through supervised learning, wherein the risk labels include normal structure, needing monitoring, and high risk;

[0130] S43. During the model training process, the validation set is used to adjust the hyperparameters to obtain a trained evaluation model.

[0131] This method introduces risk labels corresponding to historical crack data during model training and uses supervised learning methods to train a multi-layer perceptron, effectively improving the assessment model's predictive accuracy and label discrimination capabilities. By establishing a reasonable risk classification system and mapping it with crack geometry parameters, the model possesses strong structural risk identification capabilities. Furthermore, the inclusion of a validation set during training for hyperparameter optimization further enhances the model's generalization performance. This method not only simplifies the manual labeling burden but also improves the consistency between assessment results and actual risk levels.

[0132] The present invention clarifies the training process of the evaluation model, including building a network structure, preparing historical samples, supervised learning modeling and validation set tuning, so that the model has a systematic training process and a clear input-output correspondence. By constructing a multi-layer perceptron model suitable for structural crack risk classification and using real samples for training, the interpretability and application stability of the model are improved. The evaluation model has the advantages of clear structure, simple implementation and rapid deployment. It can be seamlessly integrated into the intelligent bridge detection system, providing a highly automated solution for crack assessment.

[0133] Example 1:

[0134] In order to verify the feasibility of the present invention in implementation, the present invention was applied to a typical bridge structure. This bridge spans a highway, a provincial road and an urban expressway, carries high-frequency traffic loads all year round, has a service life of more than 10 years, and has cracks of varying degrees on the surface of the structure, which is representative and has engineering application value.

[0135] In actual deployment, the inspection team used a high-resolution industrial camera mounted on a drone platform to take close-range aerial photos of the bridge surface. A total of 1,152 original images were collected, all with a resolution of 3,264 × 2,448 pixels. The images were batch processed using the image preprocessing module proposed in this invention to correct angular distortion, balance light brightness, and perform normalization. The images were then input into an image recognition model built on the CPSeg architecture. The model has been pre-trained on a training set containing over 20,000 crack images and has good generalization capabilities.

[0136] After the system completes pixel-level crack segmentation for all images, it extracts crack geometric parameters, including crack length, maximum width, main direction angle, density, and curvature, and constructs feature vectors. These vectors are then fed into a multi-layer perceptron assessment model trained through supervised learning to determine the structural risk level. The assessment levels are divided into three categories: normal structure, monitoring required, and high risk, based on the "Highway Bridge Technical Condition Assessment Standards" of the transportation industry.

[0137] In order to compare the recognition efficiency and accuracy of the method of the present invention, a traditional manual visual inspection group was set up, and experienced bridge inspection engineers manually marked the crack areas on the same batch of images and assessed the risk levels.

[0138] In addition, the system can still stably output mask images with clear boundaries and structures when facing multiple types of bridges and complex crack backgrounds (such as lighting changes, dirt occlusion, and oblique cracks). The average boundary positioning error is controlled within the range of ±2 pixels, which is significantly better than traditional image processing algorithms.

[0139] Table 1 Performance comparison of bridge crack identification and risk assessment systems

[0140]

[0141]

[0142] The system takes an average of 3.8 seconds to identify a bridge image, while traditional manual inspection takes an average of about 28 seconds. The system's processing efficiency has increased by more than 7 times. This high efficiency enables the system to be deployed on a large scale for drone inspections, edge computing terminals or bridge monitoring nodes, improving the real-time and coverage capabilities of bridge health monitoring, which is particularly suitable for the growing needs of bridge asset operation and maintenance management.

[0143] In terms of the boundary positioning capability of crack detection, the system positioning error averaged ±2.1 pixels, which shows higher accuracy than the manual average error of ±5.8 pixels. This is mainly due to the adoption of the context-aware CPSeg model architecture and the introduction of boundary optimization strategies (including edge-guided supervision and edge-sensitive loss functions) in the present invention. This enables the system to accurately identify crack edges even in complex backgrounds (such as shadows, asphalt coverings, and the junctions of structural components), avoiding problems such as blurred boundaries and poor crack connectivity that occur in traditional methods.

[0144] In terms of recognition accuracy, the average detection rate of this system is as high as 92.7%, while that of manual detection is only 89.2%. Especially when there are interference factors such as fine cracks (width less than 0.3mm) or pollution and reflection on the surface of the structure, the present invention shows stronger anti-interference ability and recognition robustness. This technical capability stems from the advantages of deep learning models in understanding image details and context modeling. It is better than traditional methods that rely on human observation, and avoids missed detections and misjudgments caused by human fatigue and subjective judgment.

[0145] The present invention also significantly improves the ability to extract crack structural parameters. The system can not only automatically calculate the length and width of the crack, but also extract the main direction angle, crack density and curvature information, and construct a feature vector to input into the evaluation model. This comprehensiveness of parameters is difficult to achieve in traditional manual operations, but relies on approximate estimation or manual measurement, which easily leads to low accuracy of evaluation results and rough risk level division. In the present invention, these high-dimensional features are standardized and input into the multi-layer perceptron model for structural risk level judgment, making the evaluation logic scientific and the process closed-loop, with stronger practicality and data traceability.

[0146] In terms of assessment consistency, the structural risk level output by the system is consistent with the manual judgment of the expert group at a rate of up to 96.2%. This shows that in practice, the method of the present invention is sufficient to replace the traditional model that relies on expert experience and judgment, reducing dependence on manual labor, while enhancing the standardization, normalization and replicability of the assessment. This has a high application and promotion value for local traffic management units, construction supervision units and third-party bridge inspection agencies.

[0147] In addition, in terms of adaptability across bridge types, the method of the present invention has achieved stable operation and efficient output under different structural forms such as simply supported beam bridges, continuous beam bridges, and steel truss bridges, reflecting the universality and generalization capabilities of the model in a variety of actual structural environments. This adaptability means that this system can be deployed as a general platform for multi-type bridge inspection tasks across the country, saving customized development costs and improving maintenance efficiency.

[0148] In summary, the present invention not only significantly surpasses traditional manual inspection methods in terms of recognition accuracy and efficiency, but also demonstrates systematic advantages in technical logic, feature dimensions, risk assessment process, and engineering deployability. Verified by comparative data with manual inspection, this method has high precision, high reliability, high versatility, and strong engineering practicality in real scenarios, and can effectively serve the intelligent and standardized operation and maintenance system of modern bridges.

[0149] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A bridge crack identification and automatic assessment method based on image recognition and AI modeling, characterized in that: The steps include: S1. Obtaining an original image of the bridge structure surface and performing image preprocessing on the original image to obtain a standardized image; S2. Input the standardized image into an image recognition model. The image recognition model is built based on a context-aware pyramid structure, adopts the CPSeg architecture, and introduces a boundary optimization module to perform pixel-level segmentation on the crack areas in the image and output a crack mask map containing the spatial distribution and boundary information of the cracks. S3, performing feature extraction processing on the crack mask image, extracting geometric feature parameters of the crack, and constructing them into a crack feature vector; S4. Build an evaluation model based on supervised learning methods and train the evaluation model; S5. Input the crack feature vector into the evaluation model, evaluate the structural risk level of the crack, and output a structural risk label.

2. The bridge crack identification and automatic assessment method based on image recognition and AI modeling according to claim 1 is characterized in that: The image preprocessing includes image distortion correction, grayscale conversion, denoising, brightness equalization and size normalization.

3. The bridge crack identification and automatic assessment method based on image recognition and AI modeling according to claim 1 is characterized in that: The S2 specifically includes: S21, inputting the standardized image into an image recognition model, wherein the image recognition model adopts a CPSeg model based on context perception and pyramid structure, wherein the CPSeg model includes a backbone feature extraction network, a context perception module, a pyramid decoding module, an edge detection module, and a boundary optimization module; S22. Performing multi-layer convolution operations on the standardized image through a backbone feature extraction network to extract semantic feature maps at different levels, wherein the semantic feature maps have different spatial resolutions and semantic depths; S23, inputting the semantic feature map into the context perception module, performing multi-scale spatial pyramid pooling to obtain global context information, and outputting the context feature map; S24, input all context feature maps into the pyramid decoding module, perform upsampling and layer-by-layer feature fusion operations, combine the semantic feature maps extracted from the backbone feature extraction network, restore spatial details through skip connections, and output a semantic prediction map; S25. Input the semantic prediction map into the edge detection module, and calculate the boundary prediction map by gradient amplitude calculation and convolution operation based on the semantic prediction map: in, Represents the boundary prediction graph, describing the boundary probability predicted by the model, Represents the predicted probability map, GradMag represents the execution of the Sobel operator to obtain the gradient magnitude, Conv 1×1 represents a 1×1 convolution operation, and σ represents a sigmoid activation function; S26, inputting the semantic prediction map into a boundary optimization module, wherein the boundary optimization module includes an edge-guided supervision mechanism and an edge-sensitive loss function, wherein the edge-guided supervision mechanism constructs a boundary-guided loss function based on the boundary supervision label map and the boundary prediction map, and the boundary supervision label map is subjected to a Sobel operator calculation to obtain a gradient amplitude through the real mask map, and is compared with a set threshold to obtain a binary boundary map; S27. Compare the semantic prediction map with the real mask map of the crack at the same time and construct an edge-sensitive loss function: in, Represents the edge-sensitive loss function, describing the edge-weighted mean square error loss value, w(x,y) represents the edge weight coefficient of the pixel point (x,y), S * (x, y) represents the true mask image, which describes whether the pixel belongs to the crack area. The value is 1, which means it belongs to the crack area. S28, forming a total loss function by weighted combination of the main segmentation loss function, the boundary guidance loss function, and the edge sensitive loss function, wherein the main segmentation loss function uses a cross entropy loss function to evaluate the overall classification error between the semantic prediction map and the true mask map; S29. The image recognition model trained using the total loss function receives any standardized image in the inference stage, outputs a semantic prediction map, and compares it with the set threshold to generate a crack mask map.

4. The method for identifying and automatically assessing bridge cracks based on image recognition and AI modeling according to claim 3 is characterized in that: The backbone feature extraction network uses the ResNet50 network.

5. The bridge crack identification and automatic assessment method based on image recognition and AI modeling according to claim 3 is characterized in that: The semantic prediction map describes the probability value of a pixel being predicted as a crack.

6. The bridge crack identification and automatic assessment method based on image recognition and AI modeling according to claim 1 is characterized in that: The geometric characteristic parameters include crack length, crack width, crack direction, crack density and crack curvature.

7. The bridge crack identification and automatic assessment method based on image recognition and AI modeling according to claim 1 is characterized in that: The S3 specifically includes: S31, obtaining a crack mask map, and performing connected domain analysis on the crack mask map using the 8-neighborhood method, realizing crack area identification through pixel clustering, scanning all pixels point by point, and marking the pixel with the label of the neighborhood if the current pixel value is 1 and there is a labeled pixel in the 8-neighborhood. Otherwise, a new neighborhood label is created, and a connected domain label map is output, which indicates the label number of the crack to which each pixel belongs; S32: For each crack region, perform crack feature extraction operations, including: Based on the crack mask image, use the Zhang-Suen algorithm to output a single-pixel-width skeleton image and traverse all pixels with a value of 1 in the skeleton image to generate a skeleton point set for each crack mask image. The skeleton image is a binary image of the same size as the crack mask image. Define the crack length as the maximum Euclidean distance between skeleton points and calculate the crack length; The crack width is obtained by calculating the average minimum distance from the skeleton point to the boundary point; The coordinates of the skeleton point set are combined into a sample matrix, and the covariance matrix of the sample matrix is ​​decomposed by eigenvalue, and the angle corresponding to the direction vector of the first principal component is selected as the crack direction; The crack density is calculated based on the ratio of the total number of crack pixels to the area of ​​the minimum bounding rectangle; Calculating the crack curvature, where the crack curvature is defined as the average of the angle changes between adjacent points on the skeleton path; S33. Normalize the crack length, crack width, crack direction, crack density, and crack curvature and then concatenate them into a crack feature vector.

8. The bridge crack identification and automatic assessment method based on image recognition and AI modeling according to claim 1 is characterized in that: The S4 specifically includes: S41, constructing an evaluation model, wherein the evaluation model is based on a multi-layer perceptron network; S42. Using historical bridge crack data with crack geometric parameters and corresponding risk labels as training samples, the assessment model is trained through supervised learning, wherein the risk labels include normal structure, needing monitoring, and high risk; S43. During the model training process, the validation set is used to adjust the hyperparameters to obtain a trained evaluation model.

Citation Information

Cited By

  • Cable sheath microcrack image identification method based on deep learning

    CN121392422A

  • Multi-scale image segmentation and damage assessment method for surface cracks of bridge structure

    CN121392622A

  • Multi-scale image segmentation and damage assessment method for surface cracks of bridge structure

    CN121392622B

  • Engineering geological exploration evaluation method based on remote sensing technology

    CN121392635A

  • Engineering geological exploration evaluation method based on remote sensing technology

    CN121392635B