Bridge crack intelligent identification method based on efficient sampling and multi-scale fusion
By constructing an intelligent bridge crack identification method with efficient sampling and multi-scale fusion, the bridge crack segmentation model of the efficient downsampling module, upsampling module and correction coordinate attention module is used to solve the high-precision detection problem of apparent cracks in bridge structures under complex backgrounds, and efficient and accurate bridge crack identification is achieved.
Patent Information
- Application Number
- CN202510491539.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing intelligent identification algorithm for apparent cracks of bridge structures based on deep learning has a high error detection rate in complex backgrounds, making it difficult to achieve high-precision detection. The traditional manual detection methods are inefficient and costly, and cannot meet the bridge detection needs.
The intelligent bridge crack identification method based on efficient sampling and multi-scale fusion is adopted, and the U-Net downsampling and upsampling operation is replaced by an efficient downsampling module and an efficient upsampling module, and a correction coordinate attention module is integrated between the encoder and the decoder to build a bridge crack segmentation model, enhance the extraction of long-distance relationships of the feature map, and take into account the learning of global channels and position information.
It improves the accuracy and robustness of the apparent crack detection of bridge structures, reduces noise interference, reduces the rate of misjudgment, and can accurately identify bridge cracks in complex backgrounds, improving detection efficiency and accuracy.
Smart Images

Figure CN120236198A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to an intelligent recognition method for bridge cracks based on efficient sampling and multi-scale fusion. Background Art
[0002] As one of the important components of road traffic infrastructure, the integrity and stability of the bridge structure have an important impact on the travel safety of vehicles and pedestrians. However, with the increase of service time, under the influence of various adverse factors such as extreme natural environment, aging of construction materials and heavy traffic loads, the bridge structure will gradually show damage and cracking (i.e. cracks). Among them, apparent cracks are the most common form of structural damage in bridge damage and cracking, and are also an important object of bridge appearance detection. If the apparent cracks of the bridge structure cannot be discovered and repaired in time, over time, the aging of concrete materials and the corrosion of steel bars inside the structure will be accelerated, further developing into more serious structural damage, which will have an adverse effect on the strength and stability of the bridge structure. This will not only greatly reduce the structural bearing capacity and service life of the bridge, but also cause unnecessary maintenance costs. At the same time, the repair of bridge cracks will also affect the normal passage of the bridge deck, and even cause traffic accidents in serious cases. Therefore, before the apparent cracks of the bridge structure develop malignantly, accurately obtaining crack information and repairing them in time can ensure the long-term service performance of the bridge and minimize the maintenance costs of the bridge structure.
[0003] Based on the shortcomings of traditional manual detection methods, such as strong subjectivity, low accuracy, low work efficiency, high missed detection rate, low safety, long time consumption and high cost, this method can no longer meet the huge detection needs of bridges in my country at this stage. In order to more objectively reflect the real situation of apparent cracks in bridge structures and provide accurate and reliable crack data support for bridge maintenance and management departments, scientific and efficient detection methods must be adopted. With the explosive development of deep learning technology, it has made major breakthroughs in engineering fields such as computer vision, which has greatly promoted the development of industrial intelligence. Based on this, relevant researchers at home and abroad have carried out a series of studies on intelligent recognition algorithms for apparent cracks in bridge structures based on deep learning, which has promoted the development of automatic detection of bridge cracks. However, based on the complex and changeable service environment of real bridges, and the presence of many background interferences such as oil stains, artificial lines, construction joints, water marks, graffiti, etc. on the surface of the structure, the existing intelligent recognition algorithms for apparent cracks in bridge structures based on deep learning have a high false detection rate, that is, they are prone to misidentify interference features similar to real crack features (such as construction joints and water marks) as cracks, resulting in a large room for improvement in overall robustness and generalization. Therefore, how to achieve high-precision detection of apparent cracks in bridge structures under complex backgrounds is a key problem that needs to be urgently solved in bridge surface inspection. Summary of the invention
[0004] In view of the key problem that urgently needs to be solved in bridge appearance detection, namely how to achieve high-precision detection of the apparent cracks in bridge structures under complex backgrounds, the present invention provides an intelligent bridge crack recognition method based on efficient sampling and multi-scale fusion.
[0005] In a first aspect, an embodiment of the present application provides an intelligent bridge crack recognition method based on efficient sampling and multi-scale fusion, and the method includes:
[0006] S100. Scanning the appearance of the bridge structure to be detected to obtain the appearance image data of the bridge structure to be detected;
[0007] S101. Reading the appearance image data of the bridge structure to be detected, preprocessing it, and inputting it into the established bridge crack segmentation model for recognition to obtain the detection result of the appearance image of the bridge structure to be detected; wherein, the bridge crack segmentation model is based on U-Net, and uses an efficient downsampling module and an efficient upsampling module to replace the original downsampling and upsampling operations of U-Net. At the same time, a corrected coordinate attention module is incorporated at the skip connection between the U-Net encoder and decoder; the efficient downsampling module is used to extract rich crack feature information and reduce the height and width of the feature map; the efficient upsampling module is used to retrieve dense crack feature details and amplify the height and width of the feature map; the corrected coordinate attention module is used to enhance the model's extraction of the long-distance relationship of the feature map, enabling it to take into account the learning of global channel and position information.
[0008] In an embodiment, the method for building the bridge crack segmentation model includes:
[0009] S1. Constructing the appearance crack image data of the bridge structure and the true value image data of the appearance crack of the bridge structure; the appearance crack image data of the bridge structure is obtained by manually selecting the appearance crack image data containing crack features from the collected appearance image data of the bridge structure and expanding it, and the true value image data of the appearance crack of the bridge structure is obtained by manually annotating the appearance crack features of the bridge structure according to the constructed appearance crack image data of the bridge structure;
[0010] S2. Performing binarization processing on the true value image data of the appearance crack of the bridge structure: in the true value image data of the appearance crack of the bridge structure, setting the pixel value of the appearance background of the bridge structure to 0 and the pixel value of the appearance crack of the bridge structure to 1 to digitally represent the true value image data of the appearance crack of the bridge structure;
[0011] S3. Training a bridge crack segmentation model with the appearance crack image data of the bridge structure and the binarized true value image data of the appearance crack of the bridge structure.
[0012] In one embodiment, step S1 includes the following sub-steps:
[0013] S11, manually selecting bridge structure appearance image data containing crack features from historical bridge structure appearance image data obtained by using bridge intelligent detection image sensors, that is, original bridge structure appearance crack image data;
[0014] S12, performing left-right flipping, up-down flipping, diagonal flipping and mosaic enhancement processing on the selected bridge structure apparent crack image data to obtain expanded bridge structure apparent crack image data;
[0015] S13. Summarize the original bridge structure apparent crack image data and the expanded bridge structure apparent crack image data to construct bridge structure apparent crack true value data.
[0016] In one embodiment, step S12 includes the following sub-steps:
[0017] S121, left-right flipping: taking the half width value in the width direction of the original bridge structure apparent crack image data as the reference axis, the data values are axially symmetrically exchanged to obtain the expanded bridge structure apparent crack image data of the original size;
[0018] S122, flipping up and down: taking the half height value in the height direction of the original bridge structure apparent crack image data as the reference axis, the data values are exchanged axisymmetrically to obtain the expanded bridge structure apparent crack image data of the original size;
[0019] S123, diagonal flipping: taking the half height value in the height direction of the bridge structure apparent crack image data after left-right flipping as the reference axis, the data values are axially symmetrically exchanged to obtain the expanded bridge structure apparent crack image data of the original size;
[0020] S124, Mosaic enhancement: Randomly select four original bridge structure apparent crack image data, center-crop each original bridge structure apparent crack image data according to the ratio of original height to width, to obtain bridge structure apparent crack image data of one quarter size, and then splice the center-cropped one quarter size bridge structure apparent crack image data according to the ratio of original height to width to obtain expanded bridge structure apparent crack image data of the original size.
[0021] In one embodiment, step S3 includes the following sub-steps:
[0022] S31, performing convolution, batch normalization and activation processing on the bridge structure apparent crack image data in sequence to obtain a feature map with a size of D, where D represents the original size of the bridge structure apparent crack image data;
[0023] S32, perform EDSM, convolution, batch normalization and activation processing on the feature map of size D, and obtain a feature map of size D / 2; similarly, obtain the feature maps of size D / 2 respectively. 2 ,...,D / 2 n feature map; where EDSM represents the efficient downsampling module; where n is the number of downsampling times;
[0024] S33, size D / 2 for S32 n The feature map of is processed by MCAM, convolution, batch normalization and activation to obtain a size of D / 2 n feature map; where MCAM represents the modified coordinate attention module;
[0025] S34, size D / 2 for S32 n-1 The feature map is processed by MCAM to obtain a feature map of size D / 2 n-1 The feature map of S33 is D / 2 n The feature map is processed by EUSM to obtain a size of D / 2 n-1 Finally, these two feature maps of size D / 2 are n-1 The feature map of the channel dimension is concatenated to obtain a size of D / 2 n-1 Similarly, we get the feature map of size D / 2 n-2 ,...,D feature map; where EUSM represents the efficient upsampling module;
[0026] S35. Perform convolution, batch normalization and activation processing on the feature map of size D obtained by splicing in S34 to obtain the final predicted output feature map, and then calculate the loss of the final predicted output feature map and the input binarized bridge structure apparent true image data, and perform back propagation with the calculated loss to update the model parameters; similarly, perform multiple batches of back propagation training on the network until the optimal weight matrix is obtained, and obtain the bridge crack segmentation model based on the optimal weight matrix.
[0027] In one embodiment, the EDSM processing in step S32 includes the following sub-steps:
[0028] S321, perform maximum pooling, convolution, batch normalization and activation processing on the input feature map in sequence, and obtain the height, width and number of channels of the input feature map. Figure 1 Half of the feature map;
[0029] S322, perform convolution, batch normalization and activation processing on the input feature map in sequence, and obtain the height, width and number of channels of the input feature map. Figure 1 Half of the feature map;
[0030] S323. Perform convolution, batch normalization, and activation processing on the input feature map successively to obtain a feature map with the same height and width as the input feature Figure 1 and the number of channels equal to that of the input feature map;
[0031] S324. Concatenate the feature maps obtained in S321, S322, and S323 in the channel dimension to obtain a feature map with the same height and width as the input feature Figure 1 and the number of channels twice that of the input feature map;
[0032] S325. Perform convolution, batch normalization, and activation processing on the feature map obtained in S324 successively to obtain a feature map with the same height and width as the input feature Figure 1 and the number of channels equal to that of the input feature map.
[0033] In one embodiment, the MCAM processing in step S33 includes the following sub-steps:
[0034] S331. Perform maximum pooling in the width dimension and transpose the data in the height and width dimensions on the input feature map to obtain a feature map with a height of 1, a width equal to the height of the input feature map, and the number of channels equal to that of the input feature map;
[0035] S332. Perform maximum pooling in the height dimension on the input feature map to obtain a feature map with a height of 1, a width and the number of channels both equal to that of the input feature map;
[0036] S333. Concatenate the feature maps obtained in S331 and S332 in the width dimension, and then perform convolution, batch normalization, and activation processing successively to obtain a feature map with a height of 1, a width equal to the sum of the height and width of the input feature map, and the number of channels one-fourth of that of the input feature Figure 4 ;
[0037] S334. Perform cutting in the width dimension on the feature map obtained in S333 to respectively obtain a feature map with a height of 1, a width equal to the height of the input feature map, and the number of channels one-fourth of that of the input feature Figure 4 and a feature map with a height of 1, a width equal to the input feature map, and the number of channels one-fourth of that of the input feature Figure 4 ;
[0038] S335. Perform transpose of the height and width dimensions, convolution, batch normalization, and activation processing successively on the feature map with a height of 1, a width equal to the height of the input feature map, and the number of channels one-fourth of that of the input feature obtained in S334 to obtain a feature map with a height equal to the input feature map, a width of 1, and the number of channels equal to that of the input feature map; Figure 4 ;
[0039] S336. For the feature map with a height of 1, a width equal to the input feature map, and the number of channels one-fourth of that of the input feature obtained in S334Figure 4 The feature map of one over [specific value] is successively subjected to convolution, batch normalization, and activation processing to obtain a feature map with a height of 1, and a width and number of channels both equal to those of the input feature map;
[0040] S337. Perform matrix multiplication processing on the feature maps obtained in S335 and S336 to obtain a feature map with a height, width, and number of channels all equal to those of the input feature map;
[0041] S338. Perform matrix multiplication processing on the feature map obtained in S337 and the original input feature map to obtain a feature map with a height, width, and number of channels all equal to those of the input feature map.
[0042] In one embodiment, the EUSM processing in step S34 includes the following sub-steps:
[0043] S341. Successively perform bilinear upsampling with a sampling factor of 2, convolution, batch normalization, and activation processing on the input feature map to obtain a feature map with a height and width both twice that of the input feature map and a number of channels half that of the input feature Figure 1 map;
[0044] S342. First perform transposed convolution, batch normalization, and activation processing on the input feature map, and then perform convolution, batch normalization, and activation processing to obtain a feature map with a height and width both twice that of the input feature map and a number of channels half that of the input feature Figure 1 map;
[0045] S343. First perform transposed convolution, batch normalization, and activation processing on the input feature map, and then perform convolution, batch normalization, and activation processing to obtain a feature map with a height and width both twice that of the input feature map and a number of channels equal to that of the input feature map;
[0046] S344. Perform feature splicing on the channel dimension of the feature maps obtained in S341, S342, and S343 to obtain a feature map with a height, width, and number of channels all twice that of the input feature map;
[0047] S345. Successively perform convolution, batch normalization, and activation processing on the feature map obtained in S344 to obtain a feature map with a height and width both twice that of the input feature map and a number of channels equal to that of the input feature map.
[0048] In a second aspect, an embodiment of the present application further provides an intelligent bridge crack recognition system based on efficient sampling and multi-scale fusion. The system is used to execute the method described in any of the above embodiments, and includes:
[0049] An acquisition module, configured to scan the apparent image data of the bridge structure to be detected to obtain the apparent image data of the bridge structure to be detected;
[0050] The recognition module is used to read the apparent image data of the bridge structure to be detected, preprocess it, and then input it into the built bridge crack segmentation model for recognition to obtain the detection result of the apparent image of the bridge structure to be detected. Among them, the bridge crack segmentation model is based on U-Net, replaces the original downsampling and upsampling operations of U-Net with an efficient downsampling module and an efficient upsampling module, and at the same time incorporates a corrected coordinate attention module at the skip connection between the U-Net encoder and decoder. The efficient downsampling module is used to extract rich crack feature information and reduce the height and width of the feature map. The efficient upsampling module is used to retrieve dense crack feature details and increase the height and width of the feature map. The corrected coordinate attention module is used to enhance the model's extraction of long-range relationships in the feature map, enabling it to take into account the learning of global channel and position information.
[0051] When constructing the convolutional neural network model of the present invention, based on the image segmentation convolutional neural network U-Net, in the downsampling process of the encoding stage, the sparse structure of the designed efficient downsampling module (EDSM) is used to extract rich crack feature information and reduce the height and width of the feature map. In the upsampling process of the decoding stage, the sparse structure of the designed efficient upsampling module (EUSM) is used to retrieve dense crack feature details and increase the height and width of the feature map. MCAM is incorporated at the skip connection between the encoding structure and the decoding structure to enhance the model's extraction of long-range relationships in the feature map, enabling it to take into account the learning of global channel and position information, thereby realizing high-precision detection of the apparent cracks of the bridge structure. Further, the segmentation effect of the bridge crack segmentation model of the present invention has good robustness, and for the apparent image data of the bridge structure with many background interferences such as oil stains, artificial drawing lines, construction joints, water stains, graffiti, etc., it can well capture the local details of the cracks and reconstruct the global information, reduce the interference of noise, reduce the misjudgment rate of the predicted crack features, and effectively solve the problem of high-precision detection of the apparent cracks of the bridge structure under complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the implementation manners will be briefly introduced below. Obviously, the drawings in the following description are only some implementation manners of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is the step flow chart of the intelligent bridge crack recognition method based on efficient sampling and multi-scale fusion of the present invention;
[0054] Figure 2Another step flowchart of the intelligent bridge crack identification method based on efficient sampling and multi-scale fusion according to the present invention;
[0055] Figure 3 The overall structural schematic diagram of the bridge crack segmentation model according to the present invention;
[0056] Figure 4 The overall structural schematic diagram of the efficient downsampling module according to the present invention;
[0057] Figure 5 The overall structural schematic diagram of the modified coordinate attention module according to the present invention;
[0058] Figure 6 The overall structural schematic diagram of the efficient upsampling module according to the present invention. Detailed implementation manners
[0059] To make the technical solutions and advantages of the present invention clearer, the following combines specific embodiments and drawings to further elaborate on the present invention. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention.
[0060] Furthermore, to better describe the present invention and its specific embodiments, the following provides an advance explanation of professional terms that may appear subsequently:
[0061] Conv k×k(stride=s,padding=p): Represents a two-dimensional convolution operation with a convolution kernel size of k, a stride of s, and a padding of p;
[0062] Conv-T k×k(stride=s,padding=p): Represents a two-dimensional transposed convolution operation with a convolution kernel size of k, a stride of s, and a padding of p;
[0063] BN: Full name "Batch Normalization", a way of standard regularization processing, making the picture feature maps after convolution satisfy the distribution law with a mean of 0 and a variance of 1;
[0064] ReLU: A type of activation function, using its characteristics of both linearity and non-linearity to provide more sensitive activation and input, avoiding saturation;
[0065] Sigmoid: A type of activation function, using the characteristics of its S-shaped curve to make the probability value finally obtained by the model after encoding-decoding be between (0,1), and the prediction is more accurate.
[0066] As Figure 1 shown, the embodiments of the present invention provide an intelligent bridge crack identification method based on efficient sampling and multi-scale fusion, and the method includes:
[0067] S100. Scan the appearance of the bridge structure to be detected to obtain the appearance image data of the bridge structure to be detected.
[0068] In step S100, specifically, use the intelligent detection imaging sensor on the truss bridge inspection vehicle to scan the appearance of the bridge structure to be detected to obtain the appearance image data of the bridge structure to be detected. The detection can be real-time detection.
[0069] S101. Read the appearance image data of the bridge structure to be detected, preprocess it, and then input it into the established bridge crack segmentation model for recognition to obtain the detection result of the appearance image of the bridge structure to be detected; among them, the bridge crack segmentation model is based on U-Net, and uses an efficient downsampling module and an efficient upsampling module to replace the original downsampling and upsampling operations of U-Net. At the same time, a modified coordinate attention module is incorporated at the skip connection between the U-Net encoder and decoder; the efficient downsampling module is used to extract rich crack feature information and reduce the height and width of the feature map; the efficient upsampling module is used to retrieve dense crack feature details and amplify the height and width of the feature map; the modified coordinate attention module is used to enhance the model's extraction of long-distance relationships in the feature map, enabling it to take into account the learning of global channel and position information.
[0070] In step S101, the efficient downsampling module (Efficient Down Sampling Module, EDSM), the efficient upsampling module (Efficient Up Sampling Module, EUSM), and the modified coordinate attention module (Modified Coordinate Attention Module, MCAM) are all modules designed by the present invention.
[0071] When constructing a convolutional neural network model, the present invention uses the image segmentation convolutional neural network U-Net as a baseline. During the downsampling process in the encoding stage, the sparse structure of the designed efficient downsampling module (EDSM) is used to extract rich crack feature information, and the height and width of the feature map are reduced, specifically by a factor of two. During the upsampling process in the decoding stage, the sparse structure of the designed efficient upsampling module (EUSM) is used to retrieve dense crack feature details, and the height and width of the feature map are amplified, specifically by a factor of two. MCAM is incorporated at the skip connection between the encoding structure and the decoding structure to enhance the model's extraction of long-range relationships in the feature map, enabling it to take into account the learning of global channel and position information, thereby achieving high-precision detection of apparent cracks in bridge structures. Further, the segmentation effect of the bridge crack segmentation model of the present invention has good robustness. For bridge structure apparent image data with many background interferences such as oil stains, artificial markings, construction joints, water stains, graffiti, etc., it can well capture the local details of cracks and reconstruct the global information, reduce the interference of noise, reduce the misjudgment rate of predicted crack features, and effectively solve the problem of high-precision detection of apparent cracks in bridge structures under complex backgrounds.
[0072] In step S101, the preprocessing includes the process:
[0073] The apparent image of the bridge structure to be detected is normalized according to the following formula:
[0074] Out(x,y) = In(x,y) / 255
[0075] where x represents the row number of the apparent image of the bridge structure to be detected, y represents the column number of the apparent image of the bridge structure to be detected, In(x,y) represents the pixel value corresponding to the original apparent image of the bridge structure to be detected, Out(x,y) represents the pixel value corresponding to the processed apparent image of the bridge structure, and the value range of Out(x,y) is 0 to 1.
[0076] As Figure 2 shown, in one embodiment, the method for building the bridge crack segmentation model includes:
[0077] S1. Construct bridge structure apparent crack image data and bridge structure apparent crack ground truth image data; the bridge structure apparent crack image data is obtained by manually selecting and expanding the bridge structure apparent crack image data containing crack features from the collected bridge structure apparent image data, and the bridge structure apparent crack ground truth image data is obtained by manually annotating the bridge structure apparent crack features according to the constructed bridge structure apparent crack image data;
[0078] S2. Binarization processing is performed on the bridge structure apparent crack true value image data: in the bridge structure apparent crack true value image data, the pixel value of the bridge structure apparent background is set to 0, and the pixel value of the bridge structure apparent crack is set to 1, so as to represent the bridge structure apparent crack true value image data in a digital manner;
[0079] S3. Train a bridge crack segmentation model using the bridge structure apparent crack image data and the binarized bridge structure apparent crack true value image data.
[0080] In the embodiment of the present application, the final segmentation effect of the bridge crack segmentation model obtained by S1 to S3 has good robustness. For the apparent image data of the bridge structure with many background interferences such as oil stains, artificial lines, construction joints, water marks, graffiti, etc., it can well capture the local details of the cracks and reconstruct the global information, reduce the interference of noise, reduce the misjudgment rate of predicting crack characteristics, and effectively solve the problem of high-precision detection of apparent cracks in bridge structures under complex backgrounds.
[0081] In one embodiment, step S1 includes the following sub-steps:
[0082] S11, manually selecting bridge structure appearance image data containing crack features from historical bridge structure appearance image data obtained by using bridge intelligent detection image sensors, that is, original bridge structure appearance crack image data;
[0083] S12, performing left-right flipping, up-down flipping, diagonal flipping and mosaic enhancement processing on the selected bridge structure apparent crack image data to obtain expanded bridge structure apparent crack image data;
[0084] S13. Summarize the original bridge structure apparent crack image data and the expanded bridge structure apparent crack image data to construct bridge structure apparent crack true value data.
[0085] Generally, the quantity and quality of the sample data set involved in model training determine the performance of the trained model. When the quality of the training sample data set is poor, no matter whether the quantity and type of the training sample data set are more or less, there is a problem of high misrecognition rate in the performance of the trained model; when the quality of the training sample data set is good, if the quantity and type of the training sample data set are too small, there is a problem of underfitting in the trained model, and if the type of the training sample data set is sufficient but the quantity is too large, there is a problem of overfitting in the trained model. For the apparent cracks on the bridge structure, due to the complex and changeable real environment where the bridge itself is located, and there are many background interferences such as oil stains, manual drawing lines, construction joints, water marks, graffiti, etc. on its structural surface, it is difficult to ensure the quantity and quality of the training sample data set of the bridge structure apparent crack images and the training sample data set of the bridge structure apparent crack ground truth images. Based on this, in this embodiment, first, a small amount of bridge structure apparent image data containing crack features is manually selected from the historical bridge structure apparent image data obtained by the bridge intelligent detection image sensor. Secondly, the selected bridge structure apparent crack image data is processed by left - right flipping, up - down flipping and diagonal flipping to increase the quantity of the bridge structure apparent crack image data. Thirdly, the selected bridge structure apparent crack image data is processed by mosaic enhancement to increase the type of the bridge structure apparent crack image data. Finally, the original bridge structure apparent crack image data and the expanded bridge structure apparent crack image data are summarized for constructing the bridge structure apparent crack ground truth data. In addition, to ensure the quality of the constructed training sample data set of the bridge structure apparent crack ground truth images, this embodiment conducts three rounds of inspection and revision on the data label annotation work of the bridge structure apparent crack image data: in the first round, several well - trained mapping personnel use the annotation software to manually annotate the cracks on the provided bridge structure apparent crack images; in the second round, the mapping personnel correct and revise each other; in the third round, the experts with many years of research on the intelligent detection technology of the bridge structure apparent crack image data conduct the final inspection and revision.
[0086] In one embodiment, step S12 includes the following sub - steps:
[0087] S121. Left - right flipping: Taking the half - width value in the width direction of the original bridge structure apparent crack image data as the reference axis, the data values are exchanged axially symmetrically to obtain the expanded bridge structure apparent crack image data with the original size.
[0088] S122. Up - down flipping: Taking the half - height value in the height direction of the original bridge structure apparent crack image data as the reference axis, the data values are exchanged axially symmetrically to obtain the expanded bridge structure apparent crack image data with the original size.
[0089] S123. Diagonal Flip: Using the mid-height value in the height direction of the apparent crack image data of the bridge structure after left-right flipping as the reference axis, perform axisymmetric exchange of its data values to obtain the expanded apparent crack image data of the bridge structure with the original size.
[0090] S124. Mosaic Enhancement: Randomly select four original apparent crack image data of the bridge structure. According to the ratio of the original height to the width, perform central cropping on each original apparent crack image data of the bridge structure to obtain the apparent crack image data of the bridge structure with a quarter size. Then, according to the ratio of the original height to the width, splice the central-cropped apparent crack image data of the bridge structure with a quarter size to obtain the expanded apparent crack image data of the bridge structure with the original size.
[0091] In one embodiment, step S3 includes the following sub-steps:
[0092] S31. Perform convolution, batch normalization, and activation processing on the apparent crack image data of the bridge structure successively to obtain a feature map with size D, where D represents the original size of the apparent crack image data of the bridge structure.
[0093] S32. Perform EDSM, convolution, batch normalization, and activation processing on the feature map with size D successively to obtain a feature map with size D / 2; similarly, successively obtain feature maps with sizes D / 2 2 ,..., D / 2 n ; where EDSM represents the efficient downsampling module; where n is the number of downsampling times.
[0094] S33. Perform MCAM, convolution, batch normalization, and activation processing on the feature map with size D / 2 in S32 n to obtain a feature map with size D / 2 n ; where MCAM represents the modified coordinate attention module.
[0095] S34. Perform MCAM processing on the feature map with size D / 2 in S32 n-1 to obtain a feature map with size D / 2 n-1 ; perform EUSM processing on the feature map with size D / 2 in S33 n to obtain a feature map with size D / 2 n-1 , and finally perform splicing processing on these two feature maps with size D / 2 n-1 in the channel dimension to obtain a feature map with size D / 2 n-1 ; similarly, successively obtain feature maps with sizes D / 2 n-2 ,..., D; where EUSM represents the efficient upsampling module.
[0096] S35. Perform convolution, batch normalization, and activation processing on the feature map of size D obtained by splicing in S34 to obtain the final predicted output feature map. Then, calculate the loss between the final predicted output feature map and the input binarized bridge structure apparent ground truth image data, and perform backpropagation with the calculated loss to update the model parameters. Similarly, perform backpropagation training of the network in multiple batches until the optimal weight matrix is obtained, and obtain the bridge crack segmentation model based on the optimal weight matrix.
[0097] In one embodiment, n is 6, and the case where n is 6 is described below. Specifically, for step S32, the feature map with size D is successively subjected to EDSM, convolution, batch normalization, and activation processing to obtain a feature map with size D / 2; the feature map with size D / 2 is successively subjected to EDSM, convolution, batch normalization, and activation processing to obtain a feature map with size D / 4; the feature map with size D / 4 is successively subjected to EDSM, convolution, batch normalization, and activation processing to obtain a feature map with size D / 8; the feature map with size D / 8 is successively subjected to EDSM, convolution, batch normalization, and activation processing to obtain a feature map with size D / 16; the feature map with size D / 16 is successively subjected to EDSM, convolution, batch normalization, and activation processing to obtain a feature map with size D / 32; the feature map with size D / 32 is successively subjected to EDSM, convolution, batch normalization, and activation processing to obtain a feature map with size D / 64; where EDSM represents an efficient downsampling module. For step S33, the feature map with size D / 64 in S32 is subjected to MCAM, convolution, batch normalization, and activation processing to obtain a feature map with size D / 64; where MCAM represents a modified coordinate attention module.For step S34, perform MCAM processing on the feature map with a size of D / 32 in S32 to obtain a feature map with a size of D / 32; perform EUSM processing on the feature map with a size of D / 64 in S33 to obtain a feature map with a size of D / 32, and finally perform concatenation processing on these two feature maps with a size of D / 32 in the channel dimension to obtain a feature map with a size of D / 32; perform MCAM, convolution, batch normalization, and activation processing on the feature map with a size of D / 16 in S32 to obtain a feature map with a size of D / 16; perform EUSM processing on the feature map with a size of D / 32 obtained by concatenation in S34 to obtain a feature map with a size of D / 16, and finally perform concatenation processing on these two feature maps with a size of D / 16 in the channel dimension to obtain a feature map with a size of D / 16; perform MCAM, convolution, batch normalization, and activation processing on the feature map with a size of D / 8 in S32 to obtain a feature map with a size of D / 8; perform EUSM processing on the feature map with a size of D / 16 obtained by concatenation in S34 to obtain a feature map with a size of D / 8, and finally perform concatenation processing on these two feature maps with a size of D / 8 in the channel dimension to obtain a feature map with a size of D / 8; perform MCAM, convolution, batch normalization, and activation processing on the feature map with a size of D / 4 in S32 to obtain a feature map with a size of D / 4; perform EUSM processing on the feature map with a size of D / 8 obtained by concatenation in S34 to obtain a feature map with a size of D / 4, and finally perform concatenation processing on these two feature maps with a size of D / 4 in the channel dimension to obtain a feature map with a size of D / 4; perform MCAM, convolution, batch normalization, and activation processing on the feature map with a size of D / 2 in S32 to obtain a feature map with a size of D / 2; perform EUSM processing on the feature map with a size of D / 4 obtained by concatenation in S34 to obtain a feature map with a size of D / 2, and finally perform concatenation processing on these two feature maps with a size of D / 2 in the channel dimension to obtain a feature map with a size of D / 2; perform MCAM, convolution, batch normalization, and activation processing on the feature map with a size of D in S32 to obtain a feature map with a size of D; perform EUSM processing on the feature map with a size of D / 2 obtained by concatenation in S34 to obtain a feature map with a size of D, and finally perform concatenation processing on these two feature maps with a size of D in the channel dimension to obtain a feature map with a size of D; where EUSM represents an efficient upsampling module.
[0098] More specifically, in the above step S31, specifically, read the bridge structure crack image data and the corresponding bridge structure apparent crack ground truth image data, and convert them both into feature matrix vectors to obtain the bridge structure apparent crack image feature matrix vector F1 and the bridge structure apparent crack ground truth image feature matrix vector T; input the bridge structure apparent crack image feature matrix vector F1 into Figure 3In the shown bridge crack segmentation model, the input image is subjected to preliminary feature extraction by performing the "Conv 3×3 (stride=1)+BN+ReLU" process twice, obtaining the feature map F2. The size of the input F1 is (512, 1024, 1), and the size of the output F2 is (512, 1024, 16). In the above step S32, the F2 in S31 is input into the Efficient Down Sampling Module (EDSM), and its sparse structure is used to extract rich crack feature information, and the height and width of F2 are reduced. Then, the "Conv 3×3 (stride=1, padding=1)+BN+ReLU" process is performed twice for feature extraction, obtaining the feature map F3 with a size of (256, 512, 32); similarly, the feature maps F4 (F4 is obtained based on F3), F5 (F5 is obtained based on F4), F6 (F6 is obtained based on F5), F7 (F7 is obtained based on F6), and F8 (F8 is obtained based on F7) are obtained in sequence, and the corresponding sizes are (128, 256, 64), (64, 128, 128), (32, 64, 256), (16, 32, 512), and (8, 16, 1024) respectively. In the above step S33, the feature map F8 with a size of (8, 16, 1024) in S32 is input into the Modified Coordinate Attention Module (MCAM) to enhance the model's extraction of long-distance relationships in the feature map, enabling it to take into account the learning of global channel and position information. Then, the "Conv 3×3 (stride=1, padding=1)+BN+ReLU" process is performed twice for feature extraction, obtaining the feature map F9 with a size of (8, 16, 1024).In the above step S34, the feature map F7 with a size of (16, 32, 512) in S32 is input into the MCAM to obtain a feature map (16, 32, 512). The feature map F9 with a size of (8, 16, 1024) in S33 is input into the Efficient Up Sampling Module (EUSM). Its sparse structure is used to retrieve dense crack feature details, and the height and width of the feature map F9 are amplified to obtain a feature map (16, 32, 1024). Then, the feature map (16, 32, 512) and the feature map (16, 32, 1024) are concatenated once in the channel dimension to obtain a feature map (16, 32, 1536). Then, two "Conv 3×3 (stride = 1, padding = 1) + BN + ReLU" processes are performed for feature extraction to obtain a feature map F10 with a size of (16, 32, 512). Similarly, feature maps F11 (F11 is obtained based on F10), F12 (F12 is obtained based on F11), F13 (F13 is obtained based on F12), F14 (F14 is obtained based on F13), and F15 (F15 is obtained based on F14) are obtained in sequence, and the corresponding sizes are (32, 64, 256), (64, 128, 128), (128, 256, 64), (256, 512, 32), and (512, 1024, 16) respectively. In the above step S35, a "Conv 1×1 (stride = 1, padding = 0) + Sigmoid" process is performed on the feature map F15 with a size of (512, 1024, 16) obtained in S34 to obtain the final predicted output feature matrix vector O. Then, the feature map and the input bridge structure apparent crack ground truth image feature matrix vector T are calculated for Dice Loss, and the calculated loss is used for backpropagation to update the model parameters. Similarly, backpropagation training of multiple batches is performed on the network until the optimal weight matrix is obtained, and the bridge crack segmentation model is obtained based on the optimal weight matrix.
[0099] In the embodiment of the present application, during the continuous backpropagation training of the network, the weights of non-target pixels will gradually decrease, while the weights of target object pixels will gradually increase until an optimal weight matrix is reached. Dice Loss is a loss function used to monitor the coincidence degree between the result recognized by the network and the bridge structure apparent ground truth image. The smaller the value of the loss function, the closer the result recognized by the network is to the bridge structure apparent ground truth image.
[0100] In one embodiment, the EDSM process in step S32 includes the following sub-steps:
[0101] S321. Perform max pooling, convolution, batch normalization, and activation processing on the input feature map in sequence to obtain a feature map with a height, width, and number of channels all being half of those of the input feature Figure 1 map;
[0102] S322. Perform convolution, batch normalization, and activation processing on the input feature map in sequence to obtain a feature map with a height, width, and number of channels all being half of those of the input feature Figure 1 map;
[0103] S323. Perform convolution, batch normalization, and activation processing on the input feature map in sequence to obtain a feature map with a height and width both being half of those of the input feature Figure 1 map and the number of channels being equal to that of the input feature map;
[0104] S324. Concatenate the feature maps obtained in S321, S322, and S323 along the channel dimension to obtain a feature map with a height and width both being half of those of the input feature Figure 1 map and the number of channels being that of the input feature map;
[0105] S325. Perform convolution, batch normalization, and activation processing on the feature map obtained in S324 in sequence to obtain a feature map with a height and width both being half of those of the input feature Figure 1 map and the number of channels being equal to that of the input feature map.
[0106] Such as Figure 4As shown, in the above step S321, specifically, first perform a max-pooling operation with a pooling kernel of 2 and a stride of 2 on the input feature map (H, W, C), and then perform a "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" operation to obtain a feature map (H / 2, W / 2, C / 2); in the above step S322, specifically, sequentially perform a "Conv 3×3 (stride = 2, padding = 1) + BN + ReLU", a "Conv 3×3 (stride = 1, padding = 1) + BN + ReLU", and a "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" operation on the input feature map (H, W, C) to obtain a feature map (H / 2, W / 2, C / 2); in the above step S323, specifically, sequentially perform a "Conv3×3 (stride = 2, padding = 1) + BN + ReLU" and a "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" operation on the input feature map (H, W, C) to obtain a feature map (H / 2, W / 2, C); in the above step S24, specifically, concatenate the feature maps (H / 2, W / 2, C / 2), the feature map (H / 2, W / 2, C / 2), and the feature map (H / 2, W / 2, C) obtained in S321, S322, and S323 in the channel dimension to obtain a feature map (H / 2, W / 2, 2C); in the above step S325, specifically, perform a "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" operation on the (H / 2, W / 2, 2C) obtained in S334 to obtain an output feature map (H / 2, W / 2, C); thus, the EDSM processing is completed.
[0107] In one embodiment, the MCAM processing in step S33 includes the following sub-steps:
[0108] S331. Perform a max-pooling operation on the input feature map in the width dimension and a transpose operation on the height and width dimension data to obtain a feature map with a height of 1, a width equal to the height of the input feature map, and a channel number equal to the input feature map;
[0109] S332. Perform a max-pooling operation on the input feature map in the height dimension to obtain a feature map with a height of 1, a width and a channel number both equal to the input feature map;
[0110] S333. Concatenate the feature maps obtained in S331 and S332 in the width dimension, and then sequentially perform convolution, batch normalization, and activation processing to obtain a feature map with a height of 1, a width equal to the sum of the height and width of the input feature map, and a channel number equal to the input feature Figure 4One - tenth of the feature map;
[0111] S334. Perform a cut on the feature map obtained in S333 in the width dimension to respectively obtain a feature map with a height of 1, a width equal to the height of the input feature map, and a channel number equal to one - tenth of the input feature Figure 4 map, and a feature map with a height of 1, a width equal to the input feature map, and a channel number equal to one - tenth of the input feature Figure 4 map;
[0112] S335. Perform high - width dimension data transposition, convolution, batch normalization, and activation processing on the feature map with a height of 1, a width equal to the height of the input feature map, and a channel number equal to one - tenth of the input feature Figure 4 map obtained in S334 to obtain a feature map with a height equal to the input feature map, a width of 1, and a channel number equal to the input feature map;
[0113] S336. Perform convolution, batch normalization, and activation processing on the feature map with a height of 1, a width equal to the input feature map, and a channel number equal to one - tenth of the input feature Figure 4 map obtained in S334 to obtain a feature map with a height of 1, a width and a channel number both equal to the input feature map;
[0114] S337. Perform matrix multiplication on the feature maps obtained in S335 and S336 to obtain a feature map with a height, a width, and a channel number all equal to the input feature map;
[0115] S338. Perform matrix multiplication on the feature map obtained in S337 and the original input feature map to obtain a feature map with a height, a width, and a channel number all equal to the input feature map.
[0116] Such as Figure 5As shown, in the above step S331, specifically, perform max pooling on the width dimension and data transposition on the height and width dimensions on the input feature map (H, W, C) in sequence to obtain a feature map (1, H, C); in the above step S332, specifically, perform max pooling on the height dimension on the input feature map (H, W, C) to obtain a feature map (1, W, C); in the above step S333, specifically, concatenate the feature map (1, W, C) and the feature map (1, H, C) obtained in S341 on the width dimension to obtain a feature map (1, (H + W), C), and perform a "Conv3×3 (stride = 1, padding = 1) + BN + ReLU" process on the feature map (1, (H + W), C) to obtain a feature map (1, (H + W), C / 4); in the above step S334, perform a cutting process on the width dimension on the feature map (1, (H + W), C) obtained in step S333 to obtain a feature map (1, H, C / 4) and a feature map (1, W, C / 4) respectively; in the above step S335, first perform data transposition on the height and width dimensions on the feature map (1, H, C / 4) obtained in step S334 to obtain a feature map (H, 1, C / 4), and then perform a "Conv 3×3 (stride = 1, padding = 1) + BN + Sigmoid" process to obtain a feature map (H, 1, C); in the above step S336, perform a "Conv 3×3 (stride = 1, padding = 1) + BN + Sigmoid" process on the feature map (1, W, C / 4) obtained in S334 to obtain a feature map (1, W, C); in the above step S337, perform a matrix multiplication process on the feature maps (H, 1, C) and (1, W, C) obtained in S335 and S336 to obtain a feature map (H, W, C); in the above step S338, perform a matrix multiplication process on the feature map (H, W, C) obtained in step S337 and the input feature map (H, W, C) to obtain an output feature map (H, W, C). Thus, the MCAM processing of the feature map is completed.
[0117] In one embodiment, the EUSM processing in step S34 includes the following sub-steps:
[0118] S341. Perform bilinear upsampling with a sampling factor of 2, convolution, batch normalization, and activation processing on the input feature map in sequence to obtain a feature map with the same height and width as the input feature map and a channel number that is half of the input feature Figure 1 map.
[0119] S342. First perform transposed convolution, batch normalization, and activation processing on the input feature map, and then perform convolution, batch normalization, and activation processing to obtain a feature map with the same height and width as the input feature map but with the number of channels being half of the input feature map; Figure 1 Half of the feature map;
[0120] S343. First perform transposed convolution, batch normalization, and activation processing on the input feature map, and then perform convolution, batch normalization, and activation processing to obtain a feature map with the same height and width as the input feature map and with the number of channels equal to the input feature map;
[0121] S344. Concatenate the feature maps obtained in S341, S342, and S343 in the channel dimension to obtain a feature map with the same height, width, and number of channels as the input feature map;
[0122] S345. Perform convolution, batch normalization, and activation processing on the feature map obtained in S344 in sequence to obtain a feature map with the same height and width as the input feature map and with the number of channels equal to the input feature map.
[0123] Such as Figure 6As shown, in the above step S341, specifically, the input feature map (H, W, C) is first subjected to a bilinear upsampling process with a sampling factor of 2, and then a "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" process is performed to obtain a feature map (2H, 2W, C / 2); in the above step S342, specifically, the input feature map (H, W, C) is sequentially subjected to a "Conv_T 3×3 (stride = 2, padding = 1) + BN + ReLU", "Conv 3×3 (stride = 1, padding = 1) + BN + ReLU", and "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" process to obtain a feature map (2H, 2W, C / 2); in the above step S343, specifically, the input feature map (H, W, C) is sequentially subjected to a "Conv_T3×3 (stride = 2, padding = 1) + BN + ReLU" and "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" process to obtain a feature map (2H, 2W, C); in the above step S344, specifically, the feature maps (2H, 2W, C / 2), (2H, 2W, C / 2), and (2H, 2W, C) obtained in S341, S342, and S343 are concatenated in the channel dimension to obtain a feature map (2H, 2W, 2C); in the above step S345, specifically, the (2H, 2W, 2C) obtained in S344 is subjected to a "Conv 1×1 (stride = 1, padding = 0) + BN + ReLU" process to obtain an output feature map (2H, 2W, C); thus, the EUSM process is completed.
[0124] Based on the original "encoding - decoding" structure of U - Net, the proposed solution of the present invention designs and adds an Efficient Down Sampling Module (EDSM), a Modified Coordinate Attention Module (MCAM), and an Efficient Up Sampling Module (EUSM). At the same time, it is deeper than the original U - Net (considering more deep - feature extraction). Among them, the designed EDSM replaces the max - pooling operation with a pooling kernel of 2 and a stride of 2 in the original U - Net, extracts rich crack - feature information using its sparse structure, avoids the possible gradient - vanishing problem caused by the single max - pooling operation, and reduces the height and width of the feature map. The designed EUSM replaces the transposed - convolution operation with a convolution kernel of 3, a convolution stride of 2, and a padding of 1 in the original U - Net, retrieves and restores dense crack - feature details using its sparse structure, avoids the possible loss of local - detail information caused by the single transposed - convolution operation, and amplifies the height and width of the feature map. The designed MCAM is integrated into the skip connection between the encoding structure and the decoding structure in the original U - Net to enhance the model's extraction of long - distance relationships in the feature map, enabling it to consider both global channel and position information learning, thereby achieving high - precision detection of apparent cracks in bridge structures.
[0125] To verify the advantages of the present invention, relevant comparative experiments were conducted. Specifically, 500 measured apparent images of bridge structures were tested based on the traditional U - Net algorithm model, DeepLabv3+ algorithm model, SegFormer algorithm model, Swin - Unet algorithm model, and the bridge - crack segmentation model of the embodiment of the present invention. The overall indicators of each algorithm model on 500 apparent images of bridge structures are shown in Table 1.
[0126] Table 1
[0127]
[0128] Among them, four most representative indicators in the current semantic - segmentation algorithm field are adopted, namely Precision, Recall, F - measure, and Intersection over Union (IOU). The specific calculations of each indicator are as follows:
[0129]
[0130]
[0131] Among them, TP is the number of true positives, FP is the number of false positives, and FN is the number of false negatives. Further, in S3, the model predicts the crack condition: when the pixel value of a certain predicted picture is greater than or equal to 0.5, the prediction result is a crack, otherwise it is regarded as the background.
[0132] It is worth mentioning that the F-measure index is the harmonic mean of Precision and Recall, which can more comprehensively reflect the excellent performance of the algorithm network. As can be seen from the above table, compared with the current mainstream network models: U-Net, DeepLabv3+, SegFormer, and Swin-Unet, the bridge crack segmentation model proposed by the present invention has obvious advantages in identifying the apparent cracks of bridges.
[0133] The embodiment of the present application also provides a bridge crack intelligent recognition system based on efficient sampling and multi-scale fusion. The system is used to execute the method described in any one of the above embodiments. The system includes:
[0134] An acquisition module, configured to scan the apparent image of the bridge structure to be detected to obtain the apparent image data of the bridge structure to be detected;
[0135] An identification module, configured to read the apparent image data of the bridge structure to be detected, preprocess it, and input it into the built bridge crack segmentation model for identification to obtain the detection result of the apparent image of the bridge structure to be detected; wherein, the bridge crack segmentation model is based on U-Net, and uses an efficient downsampling module and an efficient upsampling module to replace the original downsampling and upsampling operations of U-Net. At the same time, a modified coordinate attention module is incorporated at the skip connection between the U-Net encoder and decoder; the efficient downsampling module is used to extract rich crack feature information and reduce the height and width of the feature map; the efficient upsampling module is used to retrieve dense crack feature details and amplify the height and width of the feature map; the modified coordinate attention module is used to enhance the model's extraction of the long-distance relationship of the feature map, enabling it to take into account the learning of global channel and position information.
[0136] It should be understood that the above methods are all applicable to the system of the present invention. Therefore, the embodiments of the present invention will not be described in detail here.
[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0138] It should be noted that in this article, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method that includes the element.
[0139] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A bridge crack intelligent identification method based on efficient sampling and multi-scale fusion, characterized in that: The method comprises: S100, scanning the surface of the bridge structure to be detected to obtain surface image data of the bridge structure to be detected; S101, reading the apparent image data of the bridge structure to be detected, and inputting it into the constructed bridge crack segmentation model for identification after preprocessing, so as to obtain the detection result of the apparent image of the bridge structure to be detected; wherein, the bridge crack segmentation model takes U-Net as the baseline, and uses the efficient downsampling module and the efficient upsampling module to replace the original downsampling and upsampling operations of U-Net, and integrates the corrected coordinate attention module at the jump connection between the U-Net encoder and the decoder; the efficient downsampling module is used to extract rich crack feature information and reduce the height and width of the feature map; the efficient upsampling module is used to retrieve dense crack feature details and amplify the height and width of the feature map; the corrected coordinate attention module is used to enhance the model's extraction of long-distance relationships in the feature map, so that it can take into account the learning of global channels and position information.
2. The bridge crack intelligent identification method based on efficient sampling and multi-scale fusion according to claim 1 is characterized in that: The method for building the bridge crack segmentation model includes: S1, constructing bridge structure apparent crack image data and bridge structure apparent crack true value image data; the bridge structure apparent crack image data is obtained by manually selecting bridge structure apparent crack image data containing crack features from the collected bridge structure apparent crack image data and expanding them, and the bridge structure apparent crack true value image data is obtained by manually annotating bridge structure apparent crack features based on the constructed bridge structure apparent crack image data; S2. Binarization processing is performed on the bridge structure apparent crack true value image data: in the bridge structure apparent crack true value image data, the pixel value of the bridge structure apparent background is set to 0, and the pixel value of the bridge structure apparent crack is set to 1, so as to represent the bridge structure apparent crack true value image data in a digital manner; S3. Train a bridge crack segmentation model using the bridge structure apparent crack image data and the binarized bridge structure apparent crack true value image data.
3. The intelligent identification method for bridge cracks based on efficient sampling and multi-scale fusion according to claim 2 is characterized in that: Step S1 includes the following sub-steps: S11, manually selecting bridge structure appearance image data containing crack features from historical bridge structure appearance image data obtained by using bridge intelligent detection image sensors, that is, original bridge structure appearance crack image data; S12, performing left-right flipping, up-down flipping, diagonal flipping and mosaic enhancement processing on the selected bridge structure apparent crack image data to obtain expanded bridge structure apparent crack image data; S13. Summarize the original bridge structure apparent crack image data and the expanded bridge structure apparent crack image data to construct bridge structure apparent crack true value data.
4. The bridge crack intelligent identification method based on efficient sampling and multi-scale fusion according to claim 3 is characterized in that: Step S12 includes the following sub-steps: S121, left-right flipping: taking the half width value in the width direction of the original bridge structure apparent crack image data as the reference axis, the data values are axially symmetrically exchanged to obtain the expanded bridge structure apparent crack image data of the original size; S122, flipping up and down: taking the half height value in the height direction of the original bridge structure apparent crack image data as the reference axis, the data values are exchanged axisymmetrically to obtain the expanded bridge structure apparent crack image data of the original size; S123, diagonal flipping: taking the half height value in the height direction of the bridge structure apparent crack image data after left-right flipping as the reference axis, the data values are axially symmetrically exchanged to obtain the expanded bridge structure apparent crack image data of the original size; S124, Mosaic enhancement: Randomly select four original bridge structure apparent crack image data, center-crop each original bridge structure apparent crack image data according to the ratio of original height to width, to obtain bridge structure apparent crack image data of one quarter size, and then splice the center-cropped one quarter size bridge structure apparent crack image data according to the ratio of original height to width to obtain expanded bridge structure apparent crack image data of the original size.
5. The bridge crack intelligent identification method based on efficient sampling and multi-scale fusion according to claim 2 is characterized in that: Step S3 includes the following sub-steps: S31, performing convolution, batch normalization and activation processing on the bridge structure apparent crack image data in sequence to obtain a feature map with a size of D, where D represents the original size of the bridge structure apparent crack image data; S32, performing EDSM, convolution, batch normalization and activation processing on the feature map with a size of D, to obtain a feature map with a size of D / 2; Similarly, the sizes are feature map; where EDSM represents the efficient downsampling module; where n is the number of downsampling times; S33, for S32, the size is The feature map of performs MCAM, convolution, batch normalization and activation processing to obtain a size of feature map; where MCAM represents the modified coordinate attention module; S34, for S32, the size is The feature map of is processed by MCAM to obtain a size of The feature map of S33 is The feature map of is processed by EUSM to obtain a size of Finally, these two sizes are The feature map of is concatenated in the channel dimension to obtain a size of Similarly, we get the feature map of size Feature map; where EUSM represents the efficient upsampling module; S35. Perform convolution, batch normalization and activation processing on the feature map of size D obtained by splicing in S34 to obtain the final predicted output feature map, and then calculate the loss of the final predicted output feature map and the input binarized bridge structure apparent true image data, and perform back propagation with the calculated loss to update the model parameters; similarly, perform multiple batches of back propagation training on the network until the optimal weight matrix is obtained, and obtain the bridge crack segmentation model based on the optimal weight matrix.
6. The bridge crack intelligent identification method based on efficient sampling and multi-scale fusion according to claim 5 is characterized in that: The EDSM process in step S32 includes the following sub-steps: S321, performing maximum pooling, convolution, batch normalization and activation processing on the input feature map in sequence to obtain a feature map whose height, width and number of channels are all half of the input feature map; S322, performing convolution, batch normalization and activation processing on the input feature map in sequence to obtain a feature map whose height, width and number of channels are all half of the input feature map; S323, performing convolution, batch normalization and activation processing on the input feature map in sequence to obtain a feature map whose height and width are half of the input feature map and whose number of channels is equal to that of the input feature map; S324, concatenating the feature maps obtained in S321, S322 and S323 in terms of channel dimension to obtain a feature map whose height and width are half of the input feature map and whose number of channels is twice that of the input feature map; S325. Perform convolution, batch normalization, and activation processing on the feature map obtained in S324 in sequence to obtain a feature map whose height and width are half of the input feature map and whose number of channels is equal to that of the input feature map.
7. The bridge crack intelligent identification method based on efficient sampling and multi-scale fusion according to claim 5 is characterized in that: The MCAM processing in step S33 includes the following sub-steps: S331, perform maximum pooling on the width dimension and transpose data in the height and width dimensions on the input feature map to obtain a feature map with a height of 1, a width equal to the height of the input feature map, and a number of channels equal to the input feature map; S332, performing maximum pooling processing on the height dimension on the input feature map to obtain a feature map with a height of 1 and a width and a number of channels equal to the input feature map; S333, concatenating the feature maps obtained in S331 and S332 in the width dimension, and then successively performing convolution, batch normalization and activation processing to obtain a feature map with a height of 1, a width equal to the sum of the height and width of the input feature map, and a channel number one-fourth of the input feature map; S334, cutting the feature map obtained in S333 in the width dimension to obtain a feature map with a height of 1, a width equal to the height of the input feature map and a channel number one-fourth of the input feature map, and a feature map with a height of 1, a width equal to the input feature map and a channel number one-fourth of the input feature map; S335, the feature map obtained in S334, which has a height of 1, a width of the input feature map, and a channel number one-fourth of the input feature map, is subjected to height and width dimension data transposition, convolution, batch normalization, and activation processing in sequence, to obtain a feature map with a height equal to the input feature map, a width of 1, and a channel number equal to the input feature map; S336, performing convolution, batch normalization and activation processing on the feature map obtained in S334, whose height is 1, width is equal to the input feature map and number of channels is one quarter of the input feature map, to obtain a feature map whose height is 1, width and number of channels are equal to the input feature map; S337, performing matrix multiplication processing on the feature maps obtained in S335 and S336 to obtain a feature map whose height, width and number of channels are equal to the input feature map; S338. Perform matrix multiplication on the feature map obtained in S337 and the original input feature map to obtain a feature map whose height, width and number of channels are equal to the input feature map.
8. The bridge crack intelligent identification method based on efficient sampling and multi-scale fusion according to claim 5 is characterized in that: The EUSM process in step S34 includes the following sub-steps: S341, performing bilinear upsampling with a sampling factor of 2, convolution, batch normalization, and activation processing on the input feature map in sequence, to obtain a feature map with a height and width that are twice the input feature map and a channel number that is half the input feature map; S342, first perform transposed convolution, batch normalization and activation processing on the input feature map, and then perform convolution, batch normalization and activation processing to obtain a feature map with a height and width that are twice the input feature map and a channel number that is half the input feature map; S343, first perform transposed convolution, batch normalization and activation processing on the input feature map, and then perform convolution, batch normalization and activation processing to obtain a feature map with a height and width that are twice the input feature map and a number of channels equal to the input feature map; S344, concatenating the feature maps obtained in S341, S342 and S343 in terms of channel dimension to obtain a feature map whose height, width and number of channels are twice that of the input feature map; S345. Perform convolution, batch normalization and activation processing on the feature map obtained in S344 in sequence to obtain a feature map whose height and width are twice the input feature map and whose number of channels is equal to the input feature map.
9. An intelligent bridge crack identification system based on efficient sampling and multi-scale fusion, characterized in that: The system is used to execute the method described in any one of claims 1 to 8, including: An acquisition module is used to scan the surface of the bridge structure to be detected to obtain surface image data of the bridge structure to be detected; The recognition module is used to read the apparent image data of the bridge structure to be detected, and input the data into the established bridge crack segmentation model for identification after preprocessing, so as to obtain the detection result of the apparent image of the bridge structure to be detected; wherein, the bridge crack segmentation model takes U-Net as the baseline, and uses the efficient downsampling module and the efficient upsampling module to replace the original downsampling and upsampling operations of U-Net, and integrates the corrected coordinate attention module at the jump connection between the U-Net encoder and the decoder; the efficient downsampling module is used to extract rich crack feature information and reduce the height and width of the feature map; the efficient upsampling module is used to retrieve dense crack feature details and amplify the height and width of the feature map; the corrected coordinate attention module is used to enhance the model's extraction of long-distance relationships in the feature map, so that it takes into account the learning of global channels and position information.
Citation Information
Patent Citations
Concrete bridge crack detection method based on coding-decoding structure
CN114820511A
Bridge concrete crack detection method under complex background based on deep learning
CN116823800A
Road surface crack detection method based on U-shaped full convolutional neural network
CN117152076A
Tunnel lining crack intelligent identification method based on small target identification algorithm
CN117911677A
Unmanned aerial vehicle image pavement crack segmentation method fusing multi-scale feature extraction and attention mechanism
CN118247690A
Cited By
Bridge crack detection method based on deep learning
CN122023397A