Bridge damage assessment method based on bridge crack identification

By performing feature extraction, spatial aggregation, stepwise fusion, and hybrid sampling on bridge crack images, combined with a self-attention mechanism, the problem of insufficient accuracy in bridge crack recognition was solved, and accurate assessment of bridge damage was achieved.

CN121074501APending Publication Date: 2025-12-05HEILONGJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511209337.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

In existing technologies, the feature fusion method of bridge crack image processing algorithms is simple, resulting in insufficient feature extraction and expression capabilities, inability to effectively distinguish between cracked and non-cracked areas, insufficient recognition accuracy, and inability to accurately assess the damage to bridges.

Method used

A bridge damage assessment method based on bridge crack identification is adopted. Parallel convolution and normalization are performed through two layers of windmill convolution to extract single-level features of bridge cracks. Spatial aggregation and stepwise fusion are then performed. Combined with hybrid downsampling and hybrid upsampling, a multi-head self-attention mechanism is used for processing. Finally, a classifier is used to classify bridge cracks.

Benefits of technology

It improves the accuracy of bridge crack identification, enabling accurate assessment of bridge damage. It enhances the correlation and spatial context information of bridge crack features, balances classification and location information, and improves detection capability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074501A_ABST
    Figure CN121074501A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge damage assessment method based on bridge crack identification, and relates to the technical field of deep learning and bridge crack detection. The method comprises the steps of obtaining bridge image data and inputting the bridge image data to a trained model; initial features are extracted, multiple results are obtained through first-layer windmill convolution, and bridge crack single-stage features are obtained through second-layer windmill convolution normalization after series connection and combination; performing spatial aggregation and step-by-step fusion on the bridge crack single-stage features to obtain spatial aggregation features and step-by-step fusion features, and fusing the spatial aggregation features and the step-by-step fusion features to obtain bridge crack fusion features; the features are respectively input into a mixed down-sampling module and a mixed up-sampling module to obtain global features and local detail features, and the global features and the local detail features are spliced, processed by a multi-head self-attention mechanism and fused with the fusion features to obtain target features; and the bridge damage condition is evaluated through classification of the classifier. The method can effectively extract the key features of the bridge crack, improves the crack recognition precision, and accurately evaluates the damage condition of the bridge.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep learning and bridge crack detection, and particularly relates to a bridge damage evaluation method based on bridge crack identification. BACKGROUND

[0002] The safety and stability of bridges are the key guarantee of the transportation network. This patent focuses on bridges after being attacked by floods. The impact of floods can scour the bridge foundation, causing the foundation to be hollowed out or sink, and then causing the bridge pier to tilt or break. Bridge cracks can reduce structural strength and stiffness, accelerate material corrosion, affect stability, and thus shorten the service life of the bridge and increase safety hazards. Therefore, it is crucial to discover and handle cracks in a timely manner.

[0003] In the prior art, traditional manual detection methods rely on naked eye observation or tools such as telescopes and bridge inspection vehicles, which have limitations such as low efficiency, strong subjectivity, many blind spots in detection, and insufficient data accuracy. Unmanned aerial vehicle (UAV) inspection can quickly cover large areas and collect high-resolution image data, especially in dangerous areas that are difficult for humans to reach. By combining manual detection and UAV inspection, the UAV efficiently and quickly covers large areas to preliminarily screen the bridge, and manual detection carefully inspects key parts and complex areas in complex environments and records detailed data. By analyzing the image data collected by the UAV through image processing algorithms, cracks can be automatically identified, achieving comprehensive and automated detection, improving inspection efficiency and accuracy, reducing human risk, discovering hidden dangers in a timely manner, and ensuring the stability and safety of the bridge. The UAV carries a high-definition camera to inspect the bridge, collects image data, and transmits it back to the ground. The system can automatically identify cracks through image processing algorithms, improve inspection efficiency and accuracy, reduce human risk, discover hidden dangers in a timely manner, and ensure the safety of the bridge. Compared with traditional methods, image recognition is more efficient and accurate, can quickly discover problems, and realizes automated detection.

[0004] However, in the prior art, the feature fusion method of the bridge crack image processing algorithm is relatively simple, which makes the feature extraction and expression ability insufficient, cannot highlight the key features of the bridge cracks, and is difficult to effectively distinguish between crack and non-crack areas, resulting in insufficient crack recognition accuracy, and thus cannot accurately evaluate the damage of the bridge.

[0005] Therefore, there is an urgent need for a bridge damage evaluation method based on bridge crack identification to improve the recognition accuracy of crack and non-crack areas for accurate evaluation of the damage of the bridge. SUMMARY

[0006] Therefore, it is necessary to provide a bridge damage evaluation method based on bridge crack identification to solve the above technical problems.

[0007] The present application adopts the following technical solutions: The application provides a bridge damage evaluation method based on bridge crack identification, comprising: Obtaining bridge image data; Extracting initial features of the bridge image data, and sequentially performing parallel convolution through two windmill convolution layers and then normalizing to obtain a plurality of bridge crack single-level features; Spatially aggregating the plurality of bridge crack single-level features to obtain bridge crack spatial aggregation features; gradually fusing the plurality of bridge crack single-level features to obtain bridge crack gradual fusion features; fusing the bridge crack spatial aggregation features and the bridge crack gradual fusion features to obtain bridge crack fusion features; performing mixed down-sampling and mixed up-sampling on the bridge crack fusion features to obtain bridge crack global features and bridge crack local detail features, and sequentially performing splicing and multi-head self-attention mechanism processing and then fusing with the bridge crack fusion features to obtain target bridge crack features; Classifying the target bridge crack features through a classifier to obtain bridge crack classification results; According to the bridge crack classification results, the damage of the bridge is evaluated.

[0008] Preferably, the bridge crack classification results are obtained through a bridge crack detection model; the bridge crack detection model comprises a windmill convolution layer, a lightweight feature fusion layer and a feature output layer. The windmill convolution layer comprises two windmill convolution layers, an ordinary convolution layer and a feature fusion layer connected in cascade, and is used for extracting a plurality of bridge crack single-level features. The lightweight feature fusion layer comprises a multi-scale feature fusion module, a lightweight feature fusion module and a convolution layer, and is used for extracting target bridge crack features according to the bridge crack single-level features. The feature output layer is used for classifying the target bridge crack features to obtain bridge crack classification results.

[0009] Preferably, the plurality of bridge crack single-level features are spatially aggregated to obtain bridge crack spatial aggregation features, and specifically comprising: Adding each adjacent two bridge crack single-level features to obtain rough features; Performing convolution and activation operations on the rough features in sequence to obtain spatial aggregation weights; and performing element multiplication on the spatial aggregation weights and the rough features to obtain spatial features; The spatial features are divided into n groups, and each group of spatial features is subjected to fusion interaction and convolution interaction operations in sequence, and then subjected to activation, series joint, normalization and the like to obtain bridge crack spatial aggregation features.

[0010] Preferably, the plurality of bridge crack single-level features are gradually fused to obtain bridge crack gradual fusion features, and specifically comprising: performing average pooling and activation operations on the coarse features to obtain a feature weight threshold value; performing batch normalization on the plurality of bridge crack single-level features respectively, and dividing the plurality of bridge crack single-level features into strong attention features and weak attention features according to the feature weight threshold value; mapping the strong attention features and the weak attention features corresponding to the bridge crack single-level features to the coarse features respectively and fusing them to obtain rich features and weakened features; performing 1x1 convolution operation on the rich features to obtain detailed features, and performing depth separable convolution operation on the weakened features and then performing weighted mapping to obtain weighted features of the weakened features; combining the detailed features and the weighted features of the weakened features to obtain bridge crack step-by-step fusion features.

[0011] Preferably, the bridge crack fusion features are mixed down-sampled to obtain bridge crack global features, specifically including: performing convolution operation on the bridge crack features to extract initial bridge crack feature maps; dividing the bridge crack features into a plurality of image blocks and grouping them by channels, calculating attention masks reflecting channel correlation within each group by a self-attention mechanism and adding them to the image blocks to obtain self-attention mechanism processed image blocks; stitching the self-attention mechanism processed image blocks and fusing them with the initial bridge crack feature maps through a C2F module to obtain bridge crack global features.

[0012] Preferably, the bridge crack fusion features are mixed up-sampled to obtain bridge crack local detail features, specifically including: performing multiple convolutions on the bridge crack features respectively to obtain bridge crack features of different resolutions; performing global attention up-sampling on the bridge crack features of different resolutions to obtain up-sampled bridge crack features; dividing the up-sampled bridge crack features into a plurality of image blocks by a self-attention mechanism, and converting each image block into a query vector, a key vector and a value vector; aggregating the query vector, the key vector and the value vector corresponding to each image block into a detail feature by a multi-layer perceptron, and copying each detail feature until the size of all detail features reaches the size of the up-sampled bridge crack features to obtain bridge crack local detail features.

[0013] Preferably, the crack classification of the bridge image data includes transverse cracks, longitudinal cracks, diagonal cracks and cross cracks.

[0014] Preferably, according to the crack classification result, the damage situation of the bridge crack is evaluated, specifically including: If the crack classification result is a longitudinal crack, the evaluation result of the bridge damage condition is determined as mild danger; If the crack classification result is a transverse crack or a cross crack, the evaluation result of the bridge damage condition is determined as moderate danger; If the crack classification result is an oblique crack, the evaluation result of the bridge damage condition is determined as severe danger.The above-mentioned at least one technical scheme adopted by the present application can achieve the following beneficial effects:

[0015] In the bridge damage evaluation method based on bridge crack identification provided by the present application, through spatial aggregation and step-by-step fusion of single-stage features of bridge cracks, the obtained bridge crack fusion features are effectively enhanced in relevance, focusing on spatial context information, separating strong and weak spatial information, and realizing accurate conversion with lightweight convolution, thereby reducing redundant features and retaining key crack information; through mixed down-sampling and mixed up-sampling of bridge crack features, global features and local detail features of bridge cracks are obtained, and the obtained target bridge crack features are further fused, which balances classification and positioning information by combining shallow high-resolution features and deep semantic features, and improves the detection capability and accuracy of crack targets.

[0016] In summary, the present application effectively improves the key feature extraction capability of bridge cracks by spatial aggregation and step-by-step fusion of single-stage features of bridge cracks, and through mixed down-sampling and mixed up-sampling, thereby improving the recognition accuracy of bridge crack targets, and accurately evaluating the damage condition of the bridge according to the recognition result. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and serve to explain the present application, but do not constitute improper limitations on the present application. In the drawings:

[0018] Figure 1 A flowchart of the bridge damage evaluation method based on bridge crack identification provided by the present application is shown in the figure; Figure 2 A bridge crack diagram of the bridge damage evaluation method based on bridge crack identification provided by the present application is shown in the figure; Figure 3 A bridge crack detection model network structure diagram of the bridge damage evaluation method based on bridge crack identification provided by the present application is shown in the figure; Figure 4 A windmill convolution diagram of the bridge damage evaluation method based on bridge crack identification provided by the present application is shown in the figure; Figure 5A lightweight feature fusion schematic diagram of a bridge damage evaluation method based on bridge crack identification is provided in the present application. Figure 6 A fusion multi-head attention module schematic diagram of a bridge damage evaluation method based on bridge crack identification is provided in the present application. Figure 7 A hybrid down-sampling schematic diagram of a bridge damage evaluation method based on bridge crack identification is provided in the present application. Figure 8 A hybrid up-sampling schematic diagram of a bridge damage evaluation method based on bridge crack identification is provided in the present application. DETAILED DESCRIPTION

[0019] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in combination with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0020] The technical solutions provided by the embodiments of the present application will be described in detail below in combination with the drawings.

[0021] Figure 1 A bridge damage evaluation method flowchart based on bridge crack identification is provided in the present application, which specifically includes the following steps: S101: Obtain bridge image data.

[0022] S102: Extract initial features of the bridge image data, and sequentially perform parallel convolution through two layers of windmill convolution and then normalize to obtain multiple single-stage crack features of the bridge.

[0023] Specifically, traditional convolution has problems such as insufficient feature extraction, difficulty in effectively distinguishing targets from backgrounds, limited receptive field, and parameter redundancy when processing small targets under unmanned aerial vehicle lenses. The windmill convolution (PConv) creates horizontal and vertical convolution kernels through asymmetric padding, enhances feature extraction capability, expands the receptive field, reduces the number of parameters, improves computational efficiency, and significantly improves detection accuracy and robustness.

[0024] Specifically, referring to Figure 4The PConv module creates horizontal and vertical convolution kernels for different regions of the image by asymmetric padding, and the size of the input tensor X of the PConv module is (h1, w1, c1). After each convolution, the PConv applies batch normalization (BN) and Sigmoid linear unit (SiLU) to improve training stability and speed. The first layer performs parallel convolution using a 1x3 convolution kernel W1(1, 3, c'), and the padding parameters P(1, 0, 0, 3) represent the number of padding pixels on the left, right, top, and bottom, respectively. After the first layer convolution, the size relationship of the output feature map is:

[0025] ; where h1 is the height of the input feature map, w1 is the width of the input feature map, c1 is the number of channels of the input feature map, c2 is the number of channels of the final output feature map of the PConv module, h' is the height of the output feature map after the convolution layer, ω' is the width of the output feature map after the convolution layer, and c' is the number of channels of the output feature map after the convolution layer.

[0026] The convolution result of the first layer is merged by concatenation operation, and then normalized by a 2x2 convolution kernel W(2, 2, c2) to adjust the size of the output feature map to (h2, w2). The receptive field of PConv decreases outward, similar to a Gaussian distribution, and the features of small targets are more concentrated in the center. PConv uses grouped convolution to increase the receptive field while reducing the number of parameters.

[0027] S103: Spatially aggregate the plurality of bridge crack single-level features to obtain bridge crack spatial aggregation features; gradually fuse the plurality of bridge crack single-level features to obtain bridge crack gradual fusion features; fuse the bridge crack spatial aggregation features and the bridge crack gradual fusion features to obtain bridge crack fusion features; perform hybrid down-sampling and hybrid up-sampling on the bridge crack fusion features to obtain bridge crack global features and bridge crack local detail features, and sequentially perform splicing and multi-head self-attention mechanism processing to fuse with the bridge crack fusion features to obtain target bridge crack features.

[0028] Optionally, the plurality of bridge crack single-level features are spatially aggregated to obtain bridge crack spatial aggregation features, specifically including: adding every two adjacent bridge crack single-level features to obtain rough features; performing convolution and activation operations on the rough features in sequence to obtain spatial aggregation weights; multiplying the spatial aggregation weights with the rough features to obtain spatial features; dividing the spatial features into n groups, and performing fusion interaction and convolution interaction operations on each group of spatial features, and then sequentially performing activation, concatenation and normalization to obtain bridge crack spatial aggregation features.

[0029] Optionally, strong information features and weak information features of the plurality of bridge crack single-level features are extracted, and the weak information features are fused with the strong information features after feature conversion to obtain bridge crack step-by-step fusion features, specifically including: performing average pooling and activation operations on the rough features to obtain a feature weight threshold; performing batch normalization on the plurality of bridge crack single-level features respectively, and dividing the plurality of bridge crack single-level features into strong attention features and weak attention features according to the feature weight threshold; mapping the strong attention features and the weak attention features corresponding to the bridge crack single-level features to the rough features respectively and fusing them to obtain rich features and weakened features; obtaining detailed features by performing 1x1 convolution operation on the rich features; obtaining weighted features of the weakened features by performing depth separable convolution operation on the weakened features; and combining the detailed features and the weighted features of the weakened features to obtain the bridge crack step-by-step fusion features.

[0030] Specifically, the network architecture of YOLOv10 is improved in the embodiment, and the network architecture is as shown in Figure 3 The detection of small target information in the bridge crack image by YOLOv10 is not accurate enough, which leads to insufficient feature extraction and expression ability, resulting in high false detection rate and missed detection. Moreover, the feature fusion method is relatively simple, and the information interaction and fusion between different scale feature maps are not sufficient enough, which cannot highlight the key features of the bridge crack and effectively distinguish the crack and non-crack regions. Therefore, by improving the feature fusion method, the imbalance problem of classification and positioning information can be effectively solved, thereby improving the accuracy and efficiency of small target detection. At the same time, by optimizing the feature fusion strategy and reducing redundant feature extraction, the precision and efficiency of crack target detection can be further improved, and the computational complexity can be reduced. For the bridge crack detection task under the unmanned aerial vehicle lens, a large receptive field is usually needed to capture the complete contour of the target and the background information. However, the feature extraction ability of the standard convolution kernel is limited, especially in the low signal-to-noise ratio and complex background conditions, the detection performance is poor. Therefore, an asymmetric convolution kernel is designed, which significantly enhances the feature extraction ability and significantly expands the receptive field while only increasing a small amount of parameters. This method not only improves the detection performance, but also maintains the computational efficiency.

[0031] Specifically, in the backbone part of YOLOv10, preliminary features are extracted through PConv, C2f and Conv modules, and then multi-scale feature fusion is carried out to enhance feature expression; in the neck part, long-distance cross-layer connection and multi-level reconstruction of features are realized; before the detection head of YOLOv10, the features are further focused, and finally target detection is carried out through v10Detect. The method of the application aims at the problems of poor feature correlation, insufficient semantic expression and many redundant features in the traditional feature fusion method, which leads to easy loss of small target information. The feature fusion strategy is proposed, including spatial aggregation and step-by-step fusion. Spatial aggregation enhances the correlation of adjacent features, focuses on spatial context information and improves semantic representation; step-by-step fusion separates strong and weak spatial information, realizes accurate conversion with lightweight convolution, reduces redundant features and preserves small target information. The two cooperate to improve the quality of feature fusion and the effect of small target detection.

[0032] Referring to Figure 5 , the single-level feature T i of the bridge crack is added to the single-level feature T i-1 to obtain a rough feature T c . Then, the feature is compressed into a single channel through 1x1 convolution to aggregate spatial information, and then Sigmoid activation is performed to generate a spatial aggregation weight , and then a spatial feature T s containing spatial information is obtained, and the calculation formula is:

[0033] ; In the formula, ω s is the spatial aggregation weight, Sigmoid is the activation function, Conv is the 1x1 convolution, is the element sum, T s is the spatial feature, T c is the rough feature, and is the element multiplication.

[0034] The spatial feature T s is divided into n groups, and specifically, each group uses a convolution module to refine the feature information of adjacent channels, and different channel global features are converted to generate an attention mask for refining the features. Finally, the features of each group are concatenated to form aggregated and highly correlated adjacent features T g , and the grouped aggregated features T g are embedded into a normalized layer with multi-layer original feature fusion to normalize the mean and standard deviation of T g , combine more spatial position information of small targets to obtain the feature T f , and the calculation process is:

[0035] ; ; where T gi is the refined feature corresponding to the i-th spatial feature, i = 1, 2,..., n, Softmax is the activation function, F and N are the fusion interaction operation and the convolution interaction operation, respectively, is the i-th spatial feature after grouping, T g is the refined feature after merging, T f is the spatial aggregated feature, mean and std represent the mean and standard deviation, respectively, T i and T i-1 are the i-th and i-1-th bridge crack single-level features, respectively.

[0036] ; where mean(·) and std(·) represent the mean and standard deviation, respectively. By using this method, the semantic information from adjacent layers is fully utilized, and the overall feature representation is enhanced.

[0037] From the different stage features T c , the average pooling and Sigmoid function are used to generate feature weight threshold P c , the bridge crack single-level features T i and T i-1 are processed by BN and Sigmoid activation, and ω i and ω i-1 are generated at each spatial position. The weight information is compared with the threshold to separate strong and weak feature information, and the expression is as follows: ; where P c is the feature weight threshold, T c is the rough feature, and avg represents the average pooling. Among them, avg represents the average pooling.

[0038] The formula is: ; where BN is the batch normalization operation, ω i and ω i-1 are the weight parameters of the i-th and i-1-th time steps, respectively.

[0039] ; where ω i up and ω i-1 up are the strong information features corresponding to the i-th bridge crack single-level feature and the i-1-th bridge crack single-level feature, respectively, ω i down and ω i-1down respectively, and Threshold represents the feature weight threshold.

[0040] Strong attention map and mapped to T c , and strong information features T up are generated by fusion. c Weak attention map is mapped to T low to generate weak information features T ; In the formula, T up is a strong information feature, and T low is a weak information feature. T up is input into a 1x1 convolution to generate a more detailed feature map T ; T low is input into a feature transformation unit (FTU) to perform weighted mapping to enhance the information flow between channels after a depth separable convolution, and the weighted features of a are processed by an adaptive average pooling and a convolution layer, and the expression is as follows: ; ; In the formula, a is the weighting coefficient of the weakened features, represents a convolution transformation layer, represents an adaptive average pooling layer, and T m is a feature fused step by step.

[0041] The feature fusion strategy of the method replaces the traditional direct fusion mode, and the output T f focuses on the perception of the spatial background, and the generated P M retains small target information and enhances semantic expression, and the finally generated T k is a high-level integration of T f and T m , ensuring coherent representation of crack-related spatial and semantic information.

[0042] Optionally, the bridge crack fusion features are mixed down-sampled to obtain bridge crack global features, specifically including: performing a convolution operation on the bridge crack features to extract an initial bridge crack feature map; the bridge crack features are segmented into multiple image blocks and grouped by channels, an attention mask reflecting channel correlation in each group is calculated by a self-attention mechanism and added to the image blocks to obtain image blocks processed by the self-attention mechanism; the image blocks processed by the self-attention mechanism are spliced and fused with the initial bridge crack feature map through a C2F module to obtain bridge crack global features.

[0043] Optionally, the bridge crack fusion feature is mixed up-sampling to obtain a bridge crack local detail feature, specifically comprising: performing multiple convolutions on the bridge crack features respectively to obtain bridge crack features of different resolutions; performing global attention up-sampling on the bridge crack features of different resolutions respectively to obtain up-sampled bridge crack features; dividing the up-sampled bridge crack features into multiple image blocks through a self-attention mechanism, and converting each image block into a query vector, a key vector and a value vector; aggregating the query vector, the key vector and the value vector corresponding to each image block into a detail feature through a multi-layer perceptron, and copying each detail feature until the size of all detail features reaches the size of the up-sampled bridge crack features to obtain the bridge crack local detail feature.

[0044] Specifically, in the detection of bridge cracks, problems such as imbalance between classification and positioning information, feature fusion redundancy, and insufficient use of high-resolution features often occur. The FMSA framework includes a multi-scale feature fusion (FDS) and a long-distance cross-layer connection (FUS) module, which can solve the above problems. FMSA fuses multi-resolution features, balances classification and positioning information, combines shallow high-resolution and deep semantic features, and improves detection capability and accuracy. FDS introduces local attention down-sampling to capture shallow spatial details; FUS uses global self-attention up-sampling of low-resolution feature maps to retain semantic information and enhance the distinction between targets and backgrounds.

[0045] Specifically, the FMSA module addresses the imbalance between classification and positioning information in deep networks, especially the conflict between semantic information and spatial information in small target detection. The FMSA module uses global multi-head self-attention (MHSA) to aggregate multi-resolution features at different depths. It combines shallow high-resolution features and deep semantic-rich low-resolution features to improve small target detection capability while maintaining positioning accuracy. The FMSA module consists of a multi-head self-attention (MSA) block, a multi-layer perceptron (MLP), and a residual connection, which preserves position information through standard sinusoidal position embedding. Referring to Figure 6 , the input feature map x is reshaped into an HxW label with a dimension of C, and its output calculation formula is:

[0046] ; In the formula, x p represents a pixel value in the feature map, and p is the position index of the pixel in the feature map.

[0047] Specifically, in CNN-based target detection, shallow layers contain high-resolution spatial information that is crucial for small target detection, but are often overlooked. The FDS module introduces local attention down-sampling (LADS), which more effectively captures spatial details than traditional convolution down-sampling. Referring to Figure 7LADS is self-attention on 2x2 color blocks, aggregated output by MLP and reshaped to reduced resolution. Its output formula is:

[0048] ; Specifically, the FUS module is crucial for small target detection in high-resolution feature maps. The FUS module combines a global self-attention mechanism and up-samples low-resolution feature maps through a global attention up-sampling module to preserve global semantic information. Referring to Figure 8 , which is defined as:

[0049] ; In the formula, Replicate is used to replicate and expand the input data to the specified dimension, MLP is a multi-layer perceptron, MSA is a multi-head self-attention, LN(x) represents a layer normalization operation, which is used to standardize the input data to have zero mean and unit variance, which helps to speed up the training process and improve the stability of the model. x∈R(H×W)×C, x represents the input feature map, R represents the real set, H and W represent the height and width of the feature map respectively, and C represents the number of channels; Z∈R(H×W×stride×stride)×strideC, the dimension of the output Z, stride is the stride or pooling step, which is used to determine the spatial dimension of the output feature map strideC represents the number of output channels divided by the pooling stride.

[0050] Specifically, the FUS module fuses the converted feature map with the FDS output and the original network features in the FMSA module, enhances the distinction between the target and the background, improves the detection accuracy, effectively handles occlusion, and improves the consistency of feature representation and small target positioning ability.

[0051] S104: Classify the target bridge crack features through the classifier to obtain a bridge crack classification result.

[0052] Optionally, the bridge crack classification result is obtained through a bridge crack detection model; the bridge crack detection model comprises a windmill convolution layer, a lightweight feature fusion layer, a fusion multi-head self-attention layer, and a feature output layer, and a network structure diagram of the bridge crack detection model is shown in Figure 3 .

[0053] Specifically, the existing bridge crack dataset is collected from more than one hundred bridges in various regions of a city. The bridge cracks in these bridges can be roughly divided into: 1) horizontal cracks; 2) vertical cracks; 3) diagonal cracks; and 4) cross cracks. The four crack conditions are shown in Figure 1Among them, the transverse cracks often appear in the area with large bending moment (such as the bottom of the midspan or the top of the support), which may indicate insufficient bending resistance, accelerate the corrosion of the reinforcement, and reduce the bearing capacity. The longitudinal cracks are cracks along the direction of the main reinforcement, which may be caused by the expansion of the reinforcement corrosion, and the occurrence of the cracks will further aggravate the spalling of the protective layer, forming a vicious cycle. The diagonal cracks appear in the area with large shear force, and the formation of the cracks will directly threaten the shear resistance of the structure, leading to sudden destruction of the bridge. The intersecting cracks are often caused by insufficient durability of the concrete or repeated overloading, and the intersecting cracks will significantly reduce the integrity of the concrete. Therefore, for bridge cracks, when the model identifies longitudinal cracks, we consider it as a light danger, when the model identifies transverse cracks and intersecting cracks, we consider it as a moderate danger, and when the model identifies diagonal cracks, we consider it as a severe danger.

[0054] S105: According to the bridge crack classification result, the bridge damage situation is evaluated.

[0055] Optionally, according to the crack classification result, the bridge crack damage situation is evaluated, specifically including: If the crack classification result is longitudinal cracks, the evaluation result of the bridge damage situation is determined as light danger; If the crack classification result is transverse cracks and intersecting cracks, the evaluation result of the bridge damage situation is determined as moderate danger; If the crack classification result is diagonal cracks, the evaluation result of the bridge damage situation is determined as severe danger.

[0056] In summary, the present application improves the attention mechanism and multi-scale feature fusion on the basis of the yolov10 network architecture, and is applied to the iron tower bolt loosening detection classification task, and its main features are: 1) Through a real data set, the bridge crack situation is classified, the cracks are divided into transverse cracks, longitudinal cracks, diagonal cracks and intersecting cracks, and then different degrees of early warning are given to different types of cracks and corresponding treatment schemes are given; a feature fusion strategy is proposed, which can solve the problems of insufficient feature representation, feature redundancy and limited computing resources of bridge cracks. A high-efficiency feature fusion framework is proposed, which realizes cross-layer connection and multi-scale feature fusion by fusing mixed down-sampling (FDS), mixed up-sampling (FUS) and fusion multi-head self-attention (FMSA) modules, thereby enhancing the feature representation and detection performance of small targets. 4) A novel pinwheel convolution (PConv) is proposed to replace the standard convolution in the low layer of the backbone network, which is more suitable for the pixel Gaussian spatial distribution of dark small targets, enhances the feature extraction capability, significantly expands the receptive field, and only introduces a small amount of parameter increment.

[0057] The technical features of the above embodiments can be combined in any manner. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not contradict each other, they should be considered to be within the scope of the present application.

Claims

1. A bridge damage assessment method based on bridge crack identification, characterized by, The method comprises the following steps: acquiring bridge image data; extracting initial features of the bridge image data, and sequentially performing parallel convolution through two windmill convolution layers and normalization to obtain a plurality of bridge crack single-level features; spatially aggregating the plurality of bridge crack single-level features to obtain bridge crack spatial aggregation features; gradually fusing the plurality of bridge crack single-level features to obtain bridge crack gradual fusion features; fusing the bridge crack spatial aggregation features and the bridge crack gradual fusion features to obtain bridge crack fusion features; and performing mixed down-sampling and mixed up-sampling on the bridge crack fusion features to obtain bridge crack global features and bridge crack local detail features, and sequentially performing splicing and multi-head self-attention mechanism processing and fusing with the bridge crack fusion features to obtain target bridge crack features; classifying the target bridge crack features through a classifier to obtain a bridge crack classification result; evaluating a bridge damage condition according to the bridge crack classification result.

2. The bridge damage assessment method based on bridge crack identification according to claim 1, wherein, The bridge crack classification result is obtained through a bridge crack detection model; the bridge crack detection model comprises a windmill convolution layer, a lightweight feature fusion layer and a feature output layer; The windmill convolution layer comprises two windmill convolution layers, an ordinary convolution layer and a feature fusion layer connected in cascade, and is configured to extract a plurality of bridge crack single-level features; The lightweight feature fusion layer comprises a multi-scale feature fusion module, a lightweight feature fusion module and a convolution layer, and is configured to extract target bridge crack features according to the bridge crack single-level features; The feature output layer is configured to classify the target bridge crack features to obtain a bridge crack classification result.

3. The bridge damage assessment method based on bridge crack identification according to claim 1, wherein, The spatial aggregation of the plurality of bridge crack single-level features to obtain the bridge crack spatial aggregation features specifically comprises: adding each adjacent two bridge crack single-level features to obtain rough features; performing convolution and activation operations on the rough features in sequence to obtain spatial aggregation weights; and performing element multiplication on the spatial aggregation weights and the rough features to obtain spatial features; dividing the spatial features into n groups, and performing fusion interaction and convolution interaction operations on each group of spatial features, and then performing activation, series joint, normalization in sequence to obtain the bridge crack spatial aggregation features.

4. The bridge damage assessment method based on bridge crack identification according to claim 3, wherein, The gradual fusion of the plurality of bridge crack single-level features to obtain the bridge crack gradual fusion features specifically comprises: performing average pooling and activation operations on the rough features to obtain a feature weight threshold; performing batch normalization on the plurality of bridge crack single-level features respectively, and dividing the plurality of bridge crack single-level features into strong attention features and weak attention features according to the feature weight threshold; mapping the strong attention features and the weak attention features corresponding to the bridge crack single-level features to the rough features respectively and fusing them to obtain rich features and weakened features; performing 1x1 convolution operation on the rich features to obtain detailed features; and performing depth separable convolution operation on the weakened features to obtain weighted features of the weakened features; combining the detailed features and the weighted features of the weakened features to obtain the bridge crack gradual fusion features.

5. The bridge damage assessment method based on bridge crack identification according to claim 1, wherein, The mixed down-sampling of the bridge crack fusion features to obtain the bridge crack global features specifically comprises: The bridge crack features are subjected to a convolution operation to extract an initial bridge crack feature map; The bridge crack features are segmented into multiple image blocks and grouped by channels, an attention mask reflecting channel correlation in each group is calculated through a self-attention mechanism and added to the image blocks to obtain self-attention mechanism processed image blocks; The self-attention mechanism processed image blocks are spliced and fused with the initial bridge crack feature map through a C2F module to obtain bridge crack global features.

6. The bridge damage assessment method based on bridge crack identification according to claim 1, wherein, The bridge crack fusion features are subjected to mixed up-sampling to obtain bridge crack local detail features, specifically including: The bridge crack features are subjected to multiple convolutions to obtain bridge crack features of different resolutions; The bridge crack features of different resolutions are subjected to global attention up-sampling to obtain up-sampled bridge crack features; The up-sampled bridge crack features are segmented into multiple image blocks through a self-attention mechanism, and each image block is converted into a query vector, a key vector and a value vector; The query vector, the key vector and the value vector corresponding to each image block are aggregated into a detail feature through a multi-layer perceptron, and each detail feature is copied until the size of all detail features reaches the size of the up-sampled bridge crack features to obtain bridge crack local detail features.

7. The bridge damage assessment method based on bridge crack identification according to claim 1, wherein, The crack classification of the bridge image data includes transverse cracks, longitudinal cracks, diagonal cracks and cross cracks.

8. The bridge damage assessment method based on bridge crack identification according to claim 7, wherein, The bridge crack damage condition is evaluated according to the crack classification result, specifically including: If the crack classification result is a longitudinal crack, the evaluation result of the bridge damage condition is determined to be light danger; If the crack classification result is a transverse crack or a cross crack, the evaluation result of the bridge damage condition is determined to be moderate danger; If the crack classification result is a diagonal crack, the evaluation result of the bridge damage condition is determined to be severe danger.

Citation Information

Cited By

  • Bridge flood risk position identification method and system

    CN122153611A