Bridge corrosion detection method

By constructing a deep semantic segmentation model to perform semantic segmentation on bridge surface and cable images, the problem of insufficient accuracy in existing bridge corrosion detection is solved, and high-precision identification and intelligent detection of multi-form corrosion on bridge surfaces and cable corrosion are achieved.

CN121883375APending Publication Date: 2026-04-17GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGXI UNIV
Filing Date
2025-12-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing bridge corrosion detection methods are insufficient in terms of accuracy, resulting in low accuracy in detecting corroded areas. In particular, under complex lighting conditions, it is difficult to accurately distinguish between the various forms of corrosion on the bridge surface and the corrosion of bridge cables.

Method used

A first-depth semantic segmentation model is used to perform semantic segmentation of multi-morphological corrosion on bridge surface images. A second-depth semantic segmentation model is combined to perform semantic segmentation of corrosion and protective layer discoloration on bridge cable images. By constructing multi-morphological bridge surface corrosion recognition, cable corrosion and protective layer discoloration recognition, the detection accuracy and precision are improved.

Benefits of technology

It has achieved comprehensive identification of various forms of corrosion on bridge surfaces and cable corrosion, improving detection accuracy and precision. Furthermore, it generates customized maintenance recommendations through AI evaluation algorithms, enhancing the intelligence level of bridge health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883375A_ABST
    Figure CN121883375A_ABST
Patent Text Reader

Abstract

The invention provides a bridge corrosion detection method, which comprises the steps of acquiring a bridge surface image of a to-be-detected bridge and a bridge cable image, performing polymorphic corrosion semantic segmentation on the bridge surface image by adopting a first deep semantic segmentation model to obtain a bridge surface semantic segmentation image, and performing multi-form corrosion detection on the bridge surface semantic segmentation image. The bridge surface semantic segmentation image is marked with a corrosion area of at least one form, and a second deep semantic segmentation model is adopted to perform semantic segmentation of inhaul cable corrosion and protective layer discoloration on the bridge inhaul cable image to obtain a bridge inhaul cable semantic segmentation image, the bridge cable semantic segmentation image is marked with a cable corrosion area and / or a protective layer color change area, and obtaining a detection result of the to-be-detected bridge according to the bridge surface semantic segmentation image and the bridge cable semantic segmentation image. Therefore, the semantic segmentation model is adopted to realize bridge surface polymorphic corrosion identification, inhaul cable corrosion and protective layer color change identification, the corrosion detection is more comprehensive, and the corrosion detection precision and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a method for detecting bridge corrosion. Background Technology

[0002] With the rapid development of computer vision technology, it is commonly used to identify bridge corrosion. Bridge corrosion not only seriously threatens public safety but can also lead to significant losses in bridge construction investment, making bridge corrosion detection particularly important.

[0003] In related technologies, a binary classification strategy is used to construct a convolutional neural network, and a sliding window mechanism is used to detect rusted areas, or a sliding window recursive localization method is used to visualize and select rusted areas.

[0004] However, the above methods still have shortcomings in terms of detection accuracy, resulting in low accuracy in detecting rusted areas. Summary of the Invention

[0005] In view of this, the present application provides a bridge corrosion detection method to solve the problem of insufficient detection accuracy in existing methods, which leads to low detection accuracy of corrosion areas.

[0006] In a first aspect, embodiments of this application provide a method for detecting bridge corrosion, including: Acquire images of the bridge surface and the bridge cables of the bridge to be inspected; Using a first-depth semantic segmentation model, the bridge surface image is semantically segmented for multi-morphological corrosion to obtain a bridge surface semantic segmentation image, in which at least one type of corrosion region is marked. A second-depth semantic segmentation model is used to perform semantic segmentation of cable corrosion and protective layer discoloration on the bridge cable image to obtain a bridge cable semantic segmentation image, in which the cable corrosion area and / or protective layer discoloration area are marked. The detection result of the bridge to be detected is obtained based on the semantic segmentation image of the bridge surface and the semantic segmentation image of the bridge cables.

[0007] This application provides a method for detecting bridge corrosion, comprising: acquiring images of the bridge surface and bridge cables; using a first deep semantic segmentation model to perform semantic segmentation of multi-morphological corrosion on the bridge surface image, obtaining a semantic segmentation image of the bridge surface, in which at least one type of corrosion region is marked; using a second deep semantic segmentation model to perform semantic segmentation of cable corrosion and protective layer discoloration on the bridge cable image, obtaining a semantic segmentation image of the bridge cables, in which cable corrosion regions and / or protective layer discoloration regions are marked; and obtaining the detection result of the bridge under test based on the semantic segmentation images of the bridge surface and bridge cables. Thus, by employing a semantic segmentation model, multi-morphological corrosion recognition of the bridge surface, cable corrosion, and protective layer discoloration are achieved, resulting in more comprehensive corrosion detection and improved accuracy and precision. Attached Figure Description

[0008] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a system architecture diagram for bridge corrosion detection provided in an embodiment of this application; Figure 2 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 1 ; Figure 3 A schematic diagram of the first deep semantic segmentation model provided in the embodiments of this application; Figure 4 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 2 ; Figure 5 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 3 ; Figure 6 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 4 ; Figure 7 A schematic diagram of the channel attention module provided in an embodiment of this application; Figure 8 A schematic diagram of the spatial attention module provided in an embodiment of this application; Figure 9 A schematic diagram of the second deep semantic segmentation model provided in the embodiments of this application; Figure 10 Flowchart of the bridge corrosion detection method provided in this application embodiment Figure 5 ; Figure 11 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 6 ; Figure 12 Flowchart of the bridge corrosion detection method provided in this application embodiment Figure 7 ; Figure 13 This is a schematic diagram of the bridge corrosion detection device provided in the embodiments of this application; Figure 14 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0011] To address the issue of low accuracy in detecting rust areas on bridges, this paper proposes a method to address the significant morphological diversity of bridge surface rust. Bridge rust exhibits notable variations, primarily manifesting as scattered patchy rust and linear rust developing along structural seams. Furthermore, since both cable rust and the altered protective layer exhibit a similar red hue in the visible light spectrum, traditional computer vision methods struggle to accurately distinguish them using color thresholding or simple classification models. While deep learning has shown promise in defect detection, complex lighting conditions in real-world scenarios (such as overcast skies, rain, and dusk) can cause color shifts in the images, further obscuring the chromatic characteristics of the two types of targets. Therefore, this application constructs a first deep semantic segmentation model for identifying multi-morphological bridge surface rust and a second deep semantic segmentation model for identifying cable rust and protective layer discoloration. This enables comprehensive rust detection, improving both accuracy and precision.

[0012] Figure 1 The system architecture diagram for bridge corrosion detection provided in the embodiments of this application is as follows: Figure 1As shown, the detection methods include two aspects: bridge surface corrosion detection and cable corrosion detection. For bridge surface corrosion detection, the bridge surface image of the bridge to be inspected is uploaded, and the first-level deep semantic segmentation model is used to perform semantic segmentation on the bridge surface image. Then, the rusted areas on the bridge surface are marked and the area of ​​the rusted areas is quantified.

[0013] For cable corrosion detection, bridge cable images are uploaded, and a second-depth semantic segmentation model is used to perform semantic segmentation on the bridge cable images. Then, the cable corrosion area, the protective layer discoloration area, and the corrosion area area are marked.

[0014] In some embodiments, an AI evaluation algorithm can also be used to automatically calculate the bridge risk score, analyze the causes of corrosion, and generate customized maintenance recommendations based on the area of ​​the rusted area on the bridge surface and the area of ​​the rusted area on the cables. These recommendations are then presented synchronously on the host computer interface. Users can save the detection time, rusted area area, bridge risk score, causes of corrosion, and customized maintenance recommendations to generate a bridge corrosion detection report. The AI ​​evaluation algorithm can be a pre-trained evaluation model or a large language model; this embodiment does not impose any particular limitation on this.

[0015] Figure 2 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 1 In this embodiment, the executing entity can be a computer device, such as a host computer.

[0016] like Figure 2 As shown, the method may include: S101. Obtain images of the bridge surface and the bridge cables of the bridge to be detected.

[0017] S102. Using the first deep semantic segmentation model, perform semantic segmentation of multi-morphological corrosion on the bridge surface image to obtain a semantic segmentation image of the bridge surface.

[0018] The first deep semantic segmentation model is used to perform semantic segmentation of multi-morphological corrosion on the bridge surface image, resulting in a bridge surface semantic segmentation image. This image is marked with at least one morphology of corrosion regions, which may include, for example, patchy or linear morphologies. The size of the bridge surface semantic segmentation image is the same as the size of the bridge surface image.

[0019] If there is patchy corrosion on the bridge surface, the patchy corrosion area is marked in the semantic segmentation image of the bridge surface. If there is linear corrosion on the bridge surface, the linear corrosion area is marked in the semantic segmentation image of the bridge surface. If there are both patchy and linear corrosion on the bridge surface, the patchy corrosion area and the linear corrosion area are marked in the semantic segmentation image of the bridge surface.

[0020] S103. Using the second-depth semantic segmentation model, semantic segmentation of cable corrosion and protective layer discoloration is performed on the bridge cable image to obtain a semantic segmentation image of the bridge cable.

[0021] The second deep semantic segmentation model is used to perform semantic segmentation of the bridge cable image to identify cable corrosion and protective layer discoloration, resulting in a bridge cable semantic segmentation image. This image is labeled with areas of cable corrosion and / or protective layer discoloration. The size of the bridge cable semantic segmentation image is the same as the size of the bridge cable image.

[0022] S104. Based on the semantic segmentation image of the bridge surface and the semantic segmentation image of the bridge cables, the detection result of the bridge to be detected is obtained.

[0023] Based on the semantic segmentation images of the bridge surface and bridge cables, a risk assessment is performed on the bridge to be inspected, resulting in a bridge risk score. The inspection results for the bridge can include the bridge risk score. Additionally, corresponding maintenance recommendations can be generated.

[0024] In an alternative implementation, the method may further include: The number of pixels corresponding to the preset calibration object is obtained based on the object distance, image distance, horizontal pixel count of the image sensor, horizontal physical scale of the effective imaging area of ​​the image sensor, and actual size of the preset calibration object when the preset calibration object is photographed. Calculate the physical size of the pixels based on the number of pixels corresponding to the preset calibration object and the actual size of the preset calibration object; The actual rust area of ​​the bridge under inspection is calculated based on the total number of pixels in the marked area and the physical size of the pixels.

[0025] The labeled area can be at least one type of corrosion area labeled in the semantic segmentation image of the bridge surface or the corrosion area of ​​the cable labeled in the semantic segmentation image of the bridge cable. Accordingly, the actual corrosion area of ​​the bridge to be detected is the actual patchy corrosion area, the actual linear corrosion area, and the actual cable corrosion area of ​​the bridge to be detected.

[0026] A preset calibration object (such as a high-precision ruler) is photographed to obtain the object distance x and image distance y at the time of shooting. The horizontal pixel count m of the image sensor and the horizontal physical size W of the effective imaging area of ​​the image sensor are also obtained. Here, object distance x is the distance from the photographed object (i.e., the preset calibration object) to the optical center of the shooting lens, image distance y is the distance from the optical center of the shooting lens to the imaging surface, and horizontal pixel count m is the inherent number of photosensitive units in the horizontal direction of the image sensor.

[0027] The imaging magnification is determined based on the object distance and image distance, and is expressed as follows: .

[0028] Let the actual size of the preset calibration object be 'a', then the imaging length of the preset calibration object on the imaging plane is:

[0029] Furthermore, based on the ratio between the horizontal pixel count of the image sensor and the horizontal physical size of the effective imaging area of ​​the image sensor, the pixel count corresponding to the preset calibration object is obtained. , represented as:

[0030] Pixel physical size refers to the physical size of a single pixel. Pixel physical size μ (unit: mm / pixel) is the ratio of the actual size of the preset calibration object to the number of pixels it corresponds to, expressed as:

[0031] Number of pixels Substituting into the above formula, we obtain the expression for the pixel physical size μ as follows:

[0032] Total number of pixels in the labeled area Based on the total number of pixels and the physical size of the pixels in the marked area, the actual corrosion area of ​​the marked area is calculated. , represented as:

[0033] Using the above method, the actual patchy corrosion area, actual linear corrosion area, and actual cable corrosion area of ​​the bridge under inspection can be calculated. In addition, the total number of pixels in the protective layer discoloration area marked in the bridge cable image can be counted, and the actual protective layer discoloration area of ​​the bridge under inspection can be calculated based on the total number of pixels in the protective layer discoloration area and the physical size of the pixels. The specific calculation process is similar to the calculation process of the actual corrosion area, and will not be repeated here.

[0034] In some embodiments, the host computer provides a corrosion detection interface, which includes a bridge surface corrosion detection button, a cable corrosion detection button, an image file operation button (including an open button and a save control), an image display area, a detection result visualization area, and a corrosion parameter information bar.

[0035] When the user clicks the "Open" button, the image of the bridge surface is opened from the pre-stored path and displayed in the image display area. Then, the user clicks the "Bridge Surface Corrosion Detection" button, which calls the first deep semantic segmentation model to perform semantic segmentation of multimorphic corrosion on the bridge surface image, resulting in a semantic segmentation image of the bridge surface. This image is then displayed in the detection result visualization area, and the actual area of ​​the rusted region on the bridge surface is displayed in the corrosion parameter information bar.

[0036] The user clicks the "Open" button, opens the bridge cable image from the pre-stored path, and displays it in the image display area. Then, the user clicks the "Cable Corrosion Detection" button, which calls the second-depth semantic segmentation model to perform semantic segmentation of the bridge cable image for cable corrosion and protective layer discoloration, resulting in a bridge cable semantic segmentation image. This image is then displayed in the detection result visualization area, and the actual area of ​​the cable corrosion area and / or protective layer discoloration area is displayed in the corrosion parameter information bar.

[0037] In addition, the save function can be used to store the semantic segmentation image of the bridge cable, the actual area of ​​the rusted area on the bridge surface, and the actual area of ​​the rusted area of ​​the cable and / or the discolored area of ​​the protective layer to a specified path.

[0038] In this embodiment, addressing the issue of insufficient interactivity in existing detection systems, an integrated host computer is built based on PyQt5 and OpenCV-Python. This computer integrates bridge surface corrosion and cable corrosion detection functions. Through modular design, seamless switching between bridge surface corrosion and cable corrosion detection functions is achieved. A Model-View architecture decouples data management from the interface, thereby constructing a full-process detection platform that integrates image processing, AI assessment, quantification of disease parameters (corrosion area), and visualization report generation. This enhances the intelligence and practicality of bridge health monitoring and provides a full-process intelligent solution for bridge health monitoring.

[0039] Figure 3 A schematic diagram of the first deep semantic segmentation model provided in the embodiments of this application, as shown below. Figure 3 As shown, the first deep semantic segmentation model includes: a backbone network, a first multi-scale corrosion recognition module, a spot corrosion feature recognition module, a linear corrosion feature recognition module, a feature fusion module, and a first decoding module.

[0040] Figure 4 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 2 ,like Figure 4 As shown, in an optional embodiment, step S102 above, employing a first deep semantic segmentation model, performs semantic segmentation of multimorphic corrosion on the bridge surface image to obtain a semantically segmented bridge surface image, including: S201. Use a backbone network to extract features from the bridge surface image to obtain deep image features and shallow image features.

[0041] The backbone network can be the backbone network in the DeepLabV3+ model, used to extract features from the bridge surface image to obtain deep image features and shallow image features. The deep image features are the deep output features of the backbone network, which usually contain high-level visual information such as the overall structure of the object, category semantics, and scene understanding. The shallow image features are the shallow output features of the backbone network, which usually contain basic visual information such as edges, corners, colors, and textures.

[0042] In some embodiments, the backbone network can be, for example, the MobileNetV2 backbone network with an inverse residual model. Since MobileNetV2 has more lightweight network parameters, using a MobileNetV2 backbone network with an inverse residual structure for depthwise separable convolutions reduces the large amount of memory consumed during inference, significantly reduces the number of model parameters, accelerates network convergence, and is more suitable for resource-constrained environments such as mobile devices.

[0043] S202. The first multi-scale corrosion recognition module is used to perform multi-scale corrosion feature recognition on deep image features to obtain multi-scale spliced ​​features.

[0044] Among them, the first multi-scale corrosion identification module can be an enhanced ASPP module derived from Atrous Spatial Pyramid Pooling (ASPP) in the DeepLabV3+ model.

[0045] A first multi-scale corrosion recognition module is used to perform multi-scale corrosion feature recognition on deep image features, and the recognized multi-scale corrosion features are then stitched together to obtain multi-scale stitched features. These multi-scale stitched features contain deeper semantic information. S203. The spot corrosion feature recognition module is used to identify multi-scale spliced ​​features to obtain spot corrosion features.

[0046] Among them, the spot corrosion feature recognition module, which is designed for the small size and discrete distribution of spot corrosion, can integrate a channel attention mechanism (Squeeze-and-Excitation Block, SE Block) to enhance the response of the feature channels related to corrosion. It also adopts a dense convolution strategy of multi-scale small convolution (such as 3x3 and 5x5 in parallel) to focus on the extraction of local subtle features, effectively recognize multi-scale spliced ​​features, and obtain spot corrosion features.

[0047] S204. A linear corrosion feature recognition module is used to identify multi-scale spliced ​​features and obtain linear corrosion features.

[0048] Among them, the linear corrosion branch takes into account the characteristics of long-range continuity and directionality of linear corrosion. It introduces the spatial attention mechanism of Convolutional Block Attention Module (CBAM) to enhance the spatial perception of corrosion contour and direction. It also uses large kernel depth separable convolution (such as 7x7) to expand the receptive field, thereby more accurately identifying the topological structure of linear corrosion such as cracks and edges in multi-scale splicing features, and obtaining linear corrosion features.

[0049] It should be noted that bridge deck corrosion exhibits significant morphological diversity, mainly manifested as scattered patchy corrosion and linear corrosion developing along structural gaps. To overcome the problem of insufficient modeling of such heterogeneous morphological features, this application innovatively introduces a dual-branch structure of a patchy corrosion feature recognition module and a linear corrosion feature recognition module in the back end of the enhanced ASPP module, so as to achieve targeted feature capture and fusion of multi-morphological corrosion.

[0050] S205. Using a feature fusion module, feature fusion is performed on the spot corrosion features and linear corrosion features to obtain fused features.

[0051] The feature fusion module integrates the patchy corrosion features and linear corrosion features output from the two branches: the patchy corrosion feature recognition module and the linear corrosion feature recognition module. The feature fusion module can use channel splicing and 1x1 convolution to achieve feature dimensionality reduction and information complementarity, and finally generate a fusion feature rich in multi-morphological corrosion information (such as a fusion feature map).

[0052] It should be noted that this design can process corrosion features of different shapes in parallel, which not only ensures the detailed recall rate of spot corrosion, but also improves the segmentation accuracy of linear corrosion boundaries. From a mechanistic perspective, it significantly enhances the model's overall ability to recognize multi-scale and multi-shaped corrosion in complex scenes.

[0053] S206. Using the first decoding module, the shallow image features and fusion features are decoded to obtain the semantic segmentation image of the bridge surface.

[0054] The first decoding module is used to perform feature stitching based on shallow image features and fused features to obtain stitched features (such as stitched feature map). Then, 3×3 convolution is used to extract features from the stitched features to obtain convolutional features (such as convolutional feature map). Then, the convolutional features are upsampled (such as 4x upsampling) to obtain the category of each pixel in the bridge surface image. Then, the pixels are classified to obtain the semantic segmentation image of the bridge surface. The semantic segmentation image of the bridge surface is marked with at least one type of rust area.

[0055] It should be noted that 4x upsampling refers to the operation of enlarging the convolutional feature map after 3×3 convolution by 4 times in terms of spatial dimensions (height and width).

[0056] See Figure 3 The first multi-scale corrosion recognition module includes: an image pooling layer, a first convolutional layer, four dilated convolutional layers, a stitching layer, a second convolutional layer, and an attention module. The four dilated convolutional layers correspond to different dilation rates.

[0057] Taking four dilated convolutional layers as an example, the porosity of the four dilated convolutional layers is 4, 8, 12, and 16, respectively.

[0058] It should be noted that, given the irregular distribution of corrosion on bridge surfaces, in order to meet the need for accurate segmentation of small corrosion areas and avoid misidentification of adjacent non-corrosion areas, an additional convolutional kernel was added to the original ASPP void ratio configuration (void ratio of 6, 12, 18) in the DeepLabV3+ model. At the same time, the void ratio values ​​were adjusted (adjusted void ratios of 4, 8, 12, 16) and the interval between each layer was reduced to increase the receptive field of the network and effectively enhance its ability to capture multi-scale corrosion features.

[0059] Figure 5 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 3 ,like Figure 5 As shown, in an optional embodiment, step S202 above, employing a first multi-scale corrosion recognition module, performs multi-scale corrosion feature recognition on deep image features to obtain multi-scale stitched features, including: S301. An image pooling layer is used to perform pooling processing on deep image features to obtain pooled features.

[0060] The pooling process includes: global average pooling, 1×1 convolution, and upsampling. An image pooling layer is used to perform global average pooling on deep image features to obtain global average pooling features. Then, 1×1 convolution is performed on the global average pooling features to obtain convolutional features. Finally, the convolutional features are upsampled to obtain pooling features.

[0061] S302. The first convolutional layer is used to perform convolution processing on the deep image features to obtain the first convolutional features.

[0062] The first convolutional layer is used to perform 1×1 convolution processing on the deep image features to obtain the first convolutional features.

[0063] S303. Four dilated convolutional layers are used to perform dilated convolution processing on deep image features to obtain four sets of dilated convolutional features.

[0064] Each dilated convolutional layer is used to perform dilated convolution processing on deep image features, resulting in a set of dilated convolutional features, where one dilated convolutional layer corresponds to one set of dilated convolutional features.

[0065] The four dilated convolutional layers can all be 3×3 convolutional layers with void ratios of 4, 8, 12, and 16, respectively.

[0066] S304. A splicing layer is used to splice the pooling features, the first convolutional features, and the four sets of dilated convolutional features to obtain the pooling convolutional spliced ​​features.

[0067] A splicing layer is used to splice pooling features, the first convolutional feature, and four sets of dilated convolutional features to obtain pooling convolution spliced ​​features.

[0068] S305. A second convolutional layer is used to perform convolution processing on the pooled convolution splicing features to obtain the second convolutional features.

[0069] A second convolutional layer is used to perform 1×1 convolution processing on the pooling convolution spliced ​​features to obtain the second convolutional features. Since the pooling convolution spliced ​​features are obtained by splicing, the number of channels is large. Performing 1×1 convolution processing on the pooling convolution spliced ​​features can reduce the number of channels.

[0070] S306. An attention module is used to generate multi-scale spliced ​​features based on the second convolutional features.

[0071] The attention module can be CBAM, a deep learning component that combines channel and spatial attention mechanisms, which can enhance the ability of convolutional neural networks to express image features.

[0072] An attention module is used to process the second convolutional features. Channel attention is used to enhance the feature channels related to corrosion, and spatial attention is used to focus on the spatial regions where corrosion occurs, resulting in multi-scale spliced ​​features.

[0073] See Figure 3 The attention module includes: channel attention module and spatial attention module.

[0074] It should be noted that the rusted areas on the bridge surface have significant differences in shape and size and complex contour structure, making accurate segmentation under complex environmental background interference a major challenge. To solve this problem, a CBAM module based on an attention mechanism is adopted. This module integrates channel attention and spatial attention mechanisms, and enhances the network's ability to capture feature channel correlation and key spatial regions in parallel, thereby effectively improving the modeling efficiency of convolutional neural networks for image features.

[0075] Figure 6 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 4 ,like Figure 6 As shown, in an optional implementation, step S306 above, employing an attention module, generates multi-scale concatenated features based on the second convolutional features, including: S401. Using a channel attention module, channel attention processing is applied to the second convolutional features to obtain channel attention features.

[0076] A channel attention module is used to perform channel attention processing on the second convolutional features to enhance the feature channels related to corrosion and weaken the feature channels unrelated to corrosion, thus obtaining channel attention features.

[0077] S402. Using a spatial attention module, spatial attention processing is performed on the channel attention features to obtain spatial attention features, and multi-scale splicing features are generated based on the channel attention features and spatial attention features.

[0078] A spatial attention module is used to perform spatial attention processing on the channel attention features to enhance the response of the corrosion-related spatial regions in the channel attention features and suppress the response of the non-corrosion regions, thus obtaining spatial attention features. Then, the channel attention features and the spatial attention features are multiplied element-wise to obtain multi-scale stitched features.

[0079] It should be noted that the channel attention mechanism enables the network to more effectively focus on corrosion-related feature channels when processing diverse corrosion morphologies, while reducing interference from irrelevant noise, thereby improving the network's adaptability and performance. The spatial attention mechanism, on the other hand, focuses on analyzing the distribution of corrosion areas in bridge surface images, helping the network acquire more accurate spatial features of corrosion. Through the combination of these two attention mechanisms, CBMA can enhance the model's ability to recognize corrosion features on bridge surfaces, ultimately improving the accuracy of corrosion detection.

[0080] Figure 7 This is a schematic diagram of the channel attention module provided in an embodiment of this application. Figure 8 This is a schematic diagram of the spatial attention module provided in an embodiment of this application.

[0081] The expression is as follows:

[0082] Among them, M c (F) represents the channel attention feature map, σ(·) represents the activation operation, and M LP (F) is the shared neural network function, F AP For the feature map after average pooling operation, F MP M is the feature map after max pooling operation. S (F) is the spatial attention feature map, f7×7 (·) represents a convolution operation with a kernel size of 7×7, and [] represents a channel concatenation operation.

[0083] like Figure 7 As shown, the second convolutional feature is feature map F. Feature map F is input, and global max pooling (focusing on the overall feature distribution) and global average pooling (focusing on the most salient features) are performed on it to obtain global max pooling features and global average pooling features. The shared parameters are a shared neural network function. After processing the global max pooling features and global average pooling features by the shared neural network function, feature fusion is performed (the sum of the two provides a more comprehensive assessment of channel importance). Then, after processing by the sigmoid activation function, channel attention feature map M is generated. c (F).

[0084] like Figure 8 As shown, the input channel attention feature map M c (F) After global max pooling and global average pooling, we obtain feature maps after average pooling and max pooling. These feature maps are then concatenated by channel concatenation, followed by a 7×7 convolution operation, and finally activated by a sigmoid activation function to obtain the spatial attention feature map M. S (F), then the channel attention feature map M c (F) and spatial attention feature map M S (F) After element-wise multiplication, the output feature map is used as a multi-scale splicing feature.

[0085] It's important to note that the core function of pooling layers is downsampling (dimensionality reduction, decreasing spatial size). For example, a pooling layer with a stride of 2 will reduce the width and height of the feature map to half of their original size (the area becomes 1 / 4), thus significantly reducing the computational cost and number of parameters in subsequent layers and preventing the network from becoming too large. Simultaneously, this helps to expand the receptive field of subsequent convolutional layers, allowing them to see a larger area in the input image and thus capture a more global picture.

[0086] In this embodiment, the channel attention mechanism adaptively weights feature channels (the network automatically learns which channels are more important for corrosion detection, giving higher weights to corrosion-related feature channels and lower weights to feature channels unrelated to corrosion areas). In complex bridge deck scenes, it prioritizes enhancing corrosion-related features, suppressing irrelevant background interference, and optimizing the network's feature selection and adaptability. The spatial attention mechanism focuses on the spatial distribution patterns of corrosion areas, using spatial weights (higher weights for control positions in corrosion areas and lower weights for spatial positions in non-corrosion areas) to enhance the representation of morphological details, significantly improving the recognition accuracy of subtle changes in corrosion boundaries. The CBAM dual-attention collaborative mechanism enhances the model's multi-dimensional representation of bridge corrosion features through complementary optimization of channel and spatial features, constructing a complete feature perception system. Integrating the CBAM mechanism at the output of the enhanced ASPP module establishes a high-precision feature focusing architecture, effectively improving the detection effect of small-scale corrosion targets through multi-level feature enhancement strategies, while also optimizing the segmentation quality of edge details.

[0087] In summary, this application constructs a deep semantic segmentation network based on multi-shape specialized branches, and improves the accuracy of bridge surface corrosion recognition through an improved DeepLabV3+ architecture. The specific technical path includes: designing an enhanced ASPP module with multiple void ratio configurations, embedding a dual-path attention mechanism (CBAM) to achieve collaborative optimization of feature channels and spatial dimensions, constructing specialized dual-branch structures for patchy and linear corrosion, extracting features for different corrosion morphologies respectively, and integrating multi-shape corrosion representations through a feature fusion mechanism to form a systematic improvement scheme of multi-scale receptive field-dual attention-morphological specialized branches, realizing fine segmentation and recognition of multi-morphological corrosion, and breaking through the technical bottleneck of corrosion feature extraction in complex scenarios.

[0088] Figure 9 A schematic diagram of the second deep semantic segmentation model provided in the embodiments of this application, as shown below. Figure 9 As shown, it includes: a preprocessing module, a color extraction module, a color analysis module, a color confidence module, a weighted feature splicing module, and an output module.

[0089] Figure 10 Flowchart of the bridge corrosion detection method provided in this application embodiment Figure 5 ,like Figure 8 As shown, in an optional embodiment, step S103 above employs a second deep semantic segmentation model to perform semantic segmentation of the bridge cable image to determine cable corrosion and protective layer discoloration, thereby obtaining a semantically segmented image of the bridge cable, including: S501. Convert the bridge cable image from RGB color space to HSV color space to obtain an HSV image.

[0090] Since the HSV color space is more consistent with the physiological characteristics of human eye in observing color, the bridge cable image is converted from the RGB color space to the HSV color space to obtain the HSV image.

[0091] The RGB color space refers to the color space composed of red (R), green (G), and blue (B), while the HSV color space refers to the color space composed of hue (H), saturation (S), and value (V).

[0092] S502. Using a preprocessing module, the HSV image is decoupled from its channels to obtain the H channel feature map, S channel feature map, and V channel feature map.

[0093] To address color constancy processing in the HSV color space, a decoupling characteristic based on hue, saturation, and brightness is constructed. Color constancy processing refers to the mechanism or method in human visual system or image processing technology that maintains the perceptual stability of object color under different lighting conditions.

[0094] Define the three channel components of the output HSV image in HSV space as follows:

[0095] The HSV color constancy optimization model can be expressed as:

[0096] in, , , These represent the H-channel feature map, S-channel feature map, and V-channel feature map, respectively, ΔH ill To estimate the hue shift of the light source relative to the D65 standard light source, λ s This is the saturation decay coefficient (recommended value 0.02 / °). , These represent the mean and standard deviation of the brightness channel in an HSV image, respectively. =0.55, =0.22 is the statistical prior for the D65 light source. , , It represents the initial component values ​​of the three channels of an HSV image after conversion to the HSV color space, when the image is in an observed state, i.e., without any color correction.

[0097] Wherein, λs = 0.02 / ° means that the saturation decreases with the angle at a rate of 0.02 per degree (linear) or about 2% per degree (exponential).

[0098] Estimating the light source refers to estimating or inferring the characteristics of the "real lighting conditions" of the scene at the time of shooting from a given image affected by lighting. In other words, estimating the light source is a mathematical description of the "non-standard light" itself (its color characteristics) by the algorithm.

[0099] In some embodiments, a 64-channel 3×3 convolutional kernel is constructed to extract features from the H-channel, S-channel, and V-channel feature maps. Spatial dimension downsampling is achieved through 3×3 max pooling with a stride of 2. A three-channel independent feature stream processing mechanism is adopted to address the decoupling characteristics of the HSV color space.

[0100] Wherein, H, S, and V represent the H-channel feature map, S-channel feature map, and V-channel feature map, respectively. This indicates pooling processing. This represents a 3×3 convolution kernel.

[0101] S503. Using a color extraction module, color features are extracted from the H-channel feature map, S-channel feature map, and V-channel feature map to obtain H-channel color features, S-channel color features, and V-channel color features.

[0102] A color extraction module is used to extract color features from the H-channel feature map, S-channel feature map, and V-channel feature map respectively, to obtain H-channel color features, S-channel color features, and V-channel color features.

[0103] S504. Using a color analysis module, gradient perception is performed on the color features of the H channel to obtain hue gradient perception features.

[0104] The H channel is the hue channel. Using a color analysis module, gradient sensing is performed on the color features of the H channel to detect the areas where the color features of the H channel change the most, thus obtaining the hue gradient sensing feature. The hue gradient sensing feature is used to indicate the areas in the H channel color features where the hue changes drastically. These areas may correspond to the boundaries of rust or protective layers.

[0105] The hue gradient perception feature is obtained by using the following formula:

[0106] Where H represents the color feature of the H channel. Indicates the color characteristics of the H channel. The gradient in direction is used to detect spatial changes in hue values. It is a learnable weight coefficient, corresponding to direction, This represents a weighted summation of gradient plots over four preset directions (horizontal, 45° diagonal, vertical, and 135° diagonal). It is a comprehensive hue gradient intensity map (hue gradient perceived feature) used to indicate all significant, multidirectional hue change areas in the H channel color feature map, which may correspond to the boundaries of corrosion or protective layers.

[0107] S505. Using a color confidence module, the weights of the H channel and S channel are obtained based on the hue gradient perception features and the S channel color features, respectively.

[0108] See Figure 9 The color confidence module includes a polar coordinate convolutional layer and a weight generation module. The polar coordinate convolutional layer performs circular convolution processing on the hue gradient perception features and S-channel color features to obtain the first circular distribution features and the second circular distribution features. The weight generation module generates H-channel weights and S-channel weights based on the first and second circular distribution features.

[0109] The polar coordinate convolutional layer can be a 6×6 ring convolutional kernel group with 64 channels, and its weight distribution... satisfy:

[0110] Where R is the normalized radius, typically 0.35, and σ is the control radial attenuation coefficient, typically = 0.1. angular frequency, For phase.

[0111] The first annular distribution feature is used to indicate the distribution of hue in the annular space, and the second annular distribution feature is used to indicate the intensity variation of saturation in each direction of hue in the annular space.

[0112] Then, the weight generation module is used to generate H-channel weights based on the first and second ring distribution characteristics. and S-channel weight , represented as:

[0113] in, It exhibits the characteristics of the first ring distribution. This represents a 1×1 convolutional layer, whose function is to map the first ring-shaped feature distribution onto a single-channel feature map. The sigmoid activation function compresses the output values ​​of the convolution to the (0,1) range. The softmax activation function generates absolute confidence, while the softmax activation function generates relative confidence.

[0114] The H channel weights are used to indicate the relative importance distribution of hue, while the S channel weights are used to indicate the relative importance distribution of saturation.

[0115] S506. A weighted feature stitching module is used to stitch features based on H channel weights, hue gradient perception features, S channel weights, S channel color features, and V channel color features to obtain weighted stitched features.

[0116] A weighted feature stitching module is used to calculate the product of the H channel weight and the hue gradient perception feature, the product of the S channel weight and the S channel color feature, and then add the product of the H channel weight and the hue gradient perception feature, the product of the S channel weight and the S channel color feature, and the V channel color feature to perform feature stitching, thus obtaining the weighted stitched feature.

[0117] For example, if the weight of the H channel is w_h, the hue gradient perception feature is H feature, the weight of the S channel is w_s, the color feature of the S channel is S feature, and the color feature of the V channel is V feature, then the weighted concatenation feature is represented as: w_h × H feature + w_s × S feature + V feature.

[0118] It should be noted that since the S channel is a brightness channel, it is relatively unreliable, so the color characteristics of the V channel can be directly transmitted.

[0119] S507. Using the output module, a semantic segmentation image of the bridge cable is generated based on the weighted stitching features.

[0120] In this embodiment, a multi-channel confidence-weighted color constancy algorithm and HSV chromaticity space conversion technology are integrated. Through spectral dimension decoupling and hue gradient perception mechanism, accurate differentiation and quantitative analysis of two types of targets under complex lighting conditions are achieved.

[0121] See Figure 9 The second deep semantic segmentation model also includes a second multi-scale corrosion recognition module. Figure 11 A flowchart illustrating the bridge corrosion detection method provided in this application embodiment. Figure 6 ,like Figure 11 As shown, in an optional implementation, step S507 above, using an output module, generates a semantic segmentation image of bridge cables based on weighted stitching features, including: S601. Using a color analysis module, cross-channel fusion is performed based on hue gradient perception features, S-channel color features, and V-channel color features to obtain cross-channel fusion features.

[0122] The cross-channel fusion feature is obtained by fusing hue gradient perception features, S-channel color features, and V-channel color features using a color analysis module. The cross-channel fusion feature includes color features from the H, S, and V channels.

[0123] S602. The second multi-scale corrosion identification module is used to identify multi-scale corrosion features based on cross-channel fusion features, and multi-scale fusion features are obtained.

[0124] The second multi-scale corrosion identification module can be ASPP from the DeepLabV3+ model. Using this module, multi-scale corrosion features across channel fusion are identified, and the identified multi-scale corrosion features are stitched together to obtain multi-scale fusion features.

[0125] S603. Upsample the multi-scale fusion features to obtain upsampled features.

[0126] Among them, the multi-scale fusion features have high semantic information. Then, the multi-scale fusion features are upsampled to obtain upsampled features, thereby amplifying the multi-scale fusion features in terms of spatial size.

[0127] S604. Using the output module, a semantic segmentation image of the bridge cable is generated based on the weighted stitching features and upsampling features.

[0128] The output module performs element-wise multiplication of the weighted splicing features and the upsampled features to modulate the features and obtain modulated features. Then, the modulated features are upsampled to gradually restore the spatial resolution. Finally, the semantic segmentation image of the bridge cable is output through 1×1 convolution.

[0129] Modulation characteristics Represented as:

[0130] in, For multi-scale fusion features, Indicates upsampling, This indicates weighted splicing features. The symbol ⊙ indicates feature splicing, and ⊙ indicates element-wise multiplication. This design enhances the discriminative power of color features while preserving spatial details.

[0131] See Figure 9 The output module includes: a feature modulation module, a second decoding module, a skip connection module, and a third convolutional layer.

[0132] A feature modulation module is used to perform feature modulation based on weighted concatenation features and upsampled features to obtain modulated features. A second decoding module is used to upsample the modulated features and then perform skip feature concatenation to obtain skip concatenation features. The skip concatenation features are then convolved to obtain decoded features. A third convolutional layer is used to perform convolution processing on the decoded features to generate a semantic segmentation image of bridge cables.

[0133] A feature modulation module is used to perform element-wise multiplication of the weighted concatenated features and upsampled features to obtain modulated features. A second decoding module is then used to upsample the modulated features by a factor of 2 to obtain the first upsampled features. The first upsampled features and the hue gradient perception features are then concatenated using a skip feature concatenation to obtain the first skip concatenated features. The first skip concatenated features are then subjected to a 3×3 convolution to obtain convolutional features. These convolutional features are then upsampled by a factor of 2 to obtain the second upsampled features. Finally, a skip feature concatenation is performed on the second upsampled features, the H-channel color features, the S-channel color features, the V-channel color features, and the hue gradient perception features. The features are concatenated to obtain the second skip concatenated features, and then subjected to 3×3 convolution to obtain convolutional features. The convolutional features are then upsampled by 2 times to obtain the third upsampled features. The third upsampled features, H-channel feature maps, S-channel feature maps, and V-channel feature maps are concatenated to obtain the third skip concatenated features. The third skip concatenated features are then subjected to 3×3 convolution to obtain convolutional features. The convolutional features are then upsampled by 2 times sequentially to obtain the decoded features. A third convolutional layer is then used to perform 3×3 convolution and 1×1 convolution sequentially on the decoded features. The output is a semantic segmentation image of bridge cables.

[0134] In other words, after element-wise feature modulation, the resulting modulated features are fed into the decoding module. The spatial resolution is gradually restored through multi-layer convolution and upsampling operations. The intermediate features of each stage are fused by skip connections. Finally, a two-channel semantic segmentation image of the bridge cable is output through a 1×1 convolutional layer, achieving pixel-level accurate identification of the rusted area and the discolored area of ​​the protective layer.

[0135] In addition, the preprocessing module, color extraction module, color analysis module, color confidence module, and weighted feature stitching module belong to the encoder part. The encoder makes the image smaller and smaller, but the features become more and more obvious and the targeting becomes stronger and stronger. The output module belongs to the decoder part. The decoder enlarges the image by upsampling and fuses it with the intermediate features of the encoder to recover the bridge cable image marked with the cable corrosion area and the protective layer discoloration area.

[0136] See Figure 9 The color extraction module includes: a first-level extraction layer, a first-level pooling layer, a second-level extraction layer, and a second-level pooling layer. Figure 12Flowchart of the bridge corrosion detection method provided in this application embodiment Figure 7 ,like Figure 12 As shown, in an optional embodiment, step S503 above uses a color extraction module to extract color features from the H-channel feature map, S-channel feature map, and V-channel feature map to obtain H-channel color features, S-channel color features, and V-channel color features, including: S701. The first-level extraction layer is used to perform convolution processing on the H-channel feature map, S-channel feature map and V-channel feature map respectively to obtain the first H-channel convolution feature, the first S-channel convolution feature and the first V-channel convolution feature.

[0137] S702. The first pooling layer is used to perform pooling processing on the first H channel convolutional features, the first S channel convolutional features, and the first V channel convolutional features to obtain hue output features, saturation output features, and brightness output features.

[0138] The first-level extraction layer performs a 16-channel 1×1 circular convolution (radius π / 6) on the H-channel feature map to obtain the circular convolution feature. This circular convolution feature is then subjected to a 64-channel 3×3 convolution (e.g., depthwise separable convolution) to obtain the first H-channel convolution feature. Additionally, batch normalization (BN) can be applied to the first H-channel convolution feature to obtain the normalized H-channel feature. Finally, an activation function is used to activate the normalized H-channel feature to obtain the first activated H-channel feature.

[0139] It should be noted that an SE module can be integrated inside or after the depthwise separable convolutional module. The SE module works by first performing global average pooling on the features, then learning the importance weights of each channel through two fully connected layers, and finally using these weights to recalibrate the original feature channels.

[0140] The first-level extraction layer performs a 64-channel 3×3 convolution (e.g., a 3×3 dilated convolution with a dilation rate of 2) on the S-channel feature map to obtain the first S-channel convolutional feature. Additionally, batch normalization (BN) can be applied to the first S-channel convolutional feature to obtain the S-channel normalized feature, and an activation function is then used to activate the S-channel normalized feature to obtain the first S-channel activated feature.

[0141] In some embodiments, for the S channel, the spatial attention mechanism of CBAM can be used to automatically focus on regions where saturation changes significantly, such as the transition zone from high-purity rust to low-purity background, to achieve feature optimization.

[0142] The first extraction layer is used to perform a 64-channel 3×3 convolution on the V-channel feature map to obtain the first V-channel convolutional feature. Additionally, batch normalization (BN) can be applied to the first V-channel convolutional feature to obtain the V-channel normalized feature, and an activation function is then used to activate the V-channel normalized feature to obtain the first V-channel activated feature.

[0143] Then, the first H channel convolutional features, the first S channel convolutional features, and the first V channel convolutional features are subjected to 3×3 max pooling with a stride of 2 to obtain hue output features, saturation output features, and brightness output features.

[0144] In some embodiments, a first pooling layer is used to perform 3×3 max pooling with a step size of 2 on the first H channel activation feature, the first S channel activation feature, and the first V channel activation feature to obtain hue output feature, saturation output feature, and brightness output feature.

[0145] It should be noted that activation processing is used to convert features into intuitive probabilistic representations, which facilitates subsequent thresholding and visualization analysis. Pooling processing reduces the width and height of the feature map, thereby significantly reducing the computational cost and number of parameters in subsequent layers and preventing the network from becoming too large.

[0146] In this embodiment, during the feature extraction stage, the hue channel employs depthwise separable convolution combined with a channel attention mechanism (SE Block). This reduces parameter redundancy and enhances color category discriminative features. The dilated convolution with a dilation rate of 2 expands the receptive field to capture local to global color purity variations, and the spatial attention mechanism of CBAM is introduced to achieve feature optimization. The V channel retains the standard convolution structure to ensure the integrity of luminance information transmission. The outputs of each channel are concatenated to form HSV feature primitives. Its channel dimension is compressed to 2 / 3 of the original network, meeting the requirements of lightweight design. This is represented as:

[0147] Where DSC, DC, and SC represent depthwise separable convolution, dilated convolution, and standard convolution operations, respectively, and H, S, and V represent the H-channel feature map, S-channel feature map, and V-channel feature map, respectively.

[0148] S703. A second-level extraction layer is used to perform convolution processing on the hue output feature, saturation output feature map and brightness output feature respectively to obtain the second H channel convolution feature, the second S channel convolution feature and the second V channel convolution feature.

[0149] S704. A second pooling layer is used to pool the second H-channel convolutional features, the second S-channel convolutional features, and the second V-channel convolutional features to obtain the H-channel color features, S-channel color features, and V-channel color features.

[0150] A second-level extraction layer is used to perform circular convolution on the hue output features to obtain circular convolution features. These circular convolution features are then subjected to a 3×3 convolution to obtain the second H-channel convolution features. Additionally, batch normalization can be applied to the second H-channel convolution features to obtain normalized H-channel features. Finally, an activation function is used to activate the normalized H-channel features to obtain the second activated H-channel features.

[0151] A second-level extraction layer is used to perform a 3×3 convolution on the saturation output features to obtain the second S-channel convolutional features. Additionally, batch normalization can be applied to the second S-channel convolutional features to obtain S-channel normalized features, and an activation function can be used to activate the S-channel normalized features to obtain the second S-channel activated features.

[0152] A second-level extraction layer is used to perform a 3×3 convolution on the brightness output features to obtain the second V-channel convolutional features. Additionally, batch normalization can be applied to the second V-channel convolutional features to obtain the V-channel normalized features, and an activation function can be used to activate the V-channel normalized features to obtain the second V-channel activated features.

[0153] Then, pooling is performed on the second H-channel convolutional features, the second S-channel convolutional features, and the second V-channel convolutional features to obtain the H-channel color features, S-channel color features, and V-channel color features.

[0154] In some embodiments, a second pooling layer is used to pool the second H channel activation feature, the second S channel activation feature, and the second V channel activation feature respectively (such as 3×3 max pooling with a stride of 2) to obtain the H channel color feature, the S channel color feature, and the V channel color feature.

[0155] It should be noted that, due to the mutual interference of cross-channel spectral response features, the confidence weight of high-information channels is easily neutralized by the weight of adjacent low-information channels. This application adopts a channel decoupling weighted architecture: by establishing a channel-independent spectral confidence evaluation model, targeted enhancement is implemented for feature channels with high illumination information entropy to improve the spectral resolution robustness of the algorithm in variable environments.

[0156] In this embodiment, both the first-level extraction layer and the second-level extraction layer can be HSV-CC modules. By extracting layer by layer through the first-level and second-level extraction layers, the accuracy of color feature extraction is improved.

[0157] It should be noted that the size of the bridge cable image is H×W. The first H-channel convolutional feature, the first S-channel convolutional feature, and the first V-channel convolutional feature can be a 64×H / 2×W / 2 feature map, indicating 32 channels and a size of H / 2×W / 2. The second H-channel convolutional feature, the second S-channel convolutional feature, and the second V-channel convolutional feature can be a 32×H / 4×W / 4 feature map, and the hue gradient perception feature can be a 64×H / 8×W / 8 feature map.

[0158] In summary, due to the mutual interference of cross-channel spectral response features, the confidence weight of high-information channels is easily neutralized by the weights of adjacent low-information channels. This application develops a second-depth semantic segmentation model that integrates a multi-channel confidence-weighted color constancy algorithm and HSV chromaticity space transformation through a channel-decoupled weighted architecture. This model effectively distinguishes between cable corrosion and the altered protective layer through chromaticity decoupling and a confidence mechanism. Simultaneously, an independent channel weighted architecture is introduced to optimize cross-channel noise suppression, achieving accurate differentiation and chromaticity quantitative analysis of the two types of targets under illumination shift scenarios, overcoming the chromaticity confusion problem under complex illumination. Regarding multi-technology integration, this project organically combines lightweight networks, multi-scale feature optimization, morphological specialization branches, and chromaticity constancy algorithms, forming a collaborative technology system from feature enhancement and morphological recognition to chromaticity analysis, constructing a complete detection scheme covering bridge surface corrosion and cable corrosion.

[0159] Figure 13 This is a schematic diagram of the structure of the bridge corrosion detection device provided in the embodiments of this application. The device can be integrated into a computer device.

[0160] like Figure 13 As shown, the device may include: The acquisition module 801 is used to acquire images of the bridge surface and the bridge cables of the bridge to be detected. The processing module 802 is used to perform semantic segmentation of multi-morphological corrosion on the bridge surface image using a first deep semantic segmentation model to obtain a semantic segmentation image of the bridge surface, wherein at least one type of corrosion region is marked in the semantic segmentation image of the bridge surface. The processing module 802 is also used to perform semantic segmentation of the bridge cable image on cable corrosion and protective layer discoloration using a second deep semantic segmentation model, so as to obtain a bridge cable semantic segmentation image, in which the cable corrosion area and / or protective layer discoloration area are marked. The processing module 802 is also used to obtain the detection result of the bridge to be detected based on the semantic segmentation image of the bridge surface and the semantic segmentation image of the bridge cables.

[0161] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0162] Figure 14 A schematic diagram of the structure of the computer device provided in the embodiments of this application, such as... Figure 14 As shown, the device may include a processor 901, a memory 902, and a bus 903. The memory 902 stores machine-readable instructions that can be executed by the processor 901. When the computer device is running, the processor 901 communicates with the memory 902 through the bus 903, and the processor 901 executes the machine-readable instructions to perform the above-described method.

[0163] This application also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the above-described method.

[0164] In this embodiment, the computer program, when run by the processor, can also execute other machine-readable instructions to perform other methods as described in the embodiments. For details on the specific execution steps and principles, please refer to the description of the embodiments, which will not be repeated here.

[0165] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0166] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0167] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0168] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0169] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0170] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A method for detecting bridge corrosion, characterized in that, include: Acquire images of the bridge surface and the bridge cables of the bridge to be inspected; Using a first-depth semantic segmentation model, the bridge surface image is semantically segmented for multi-morphological corrosion to obtain a bridge surface semantic segmentation image, in which at least one type of corrosion region is marked. A second-depth semantic segmentation model is used to perform semantic segmentation of cable corrosion and protective layer discoloration on the bridge cable image to obtain a bridge cable semantic segmentation image, in which the cable corrosion area and / or protective layer discoloration area are marked. The detection result of the bridge to be detected is obtained based on the semantic segmentation image of the bridge surface and the semantic segmentation image of the bridge cables.

2. The method according to claim 1, characterized in that, The first deep semantic segmentation model includes: a backbone network, a first multi-scale corrosion recognition module, a patchy corrosion feature recognition module, a linear corrosion feature recognition module, a feature fusion module, and a first decoding module; the step of using the first deep semantic segmentation model to perform semantic segmentation of multi-morphological corrosion on the bridge surface image to obtain a semantic segmentation image of the bridge surface includes: The backbone network is used to extract features from the bridge surface image to obtain deep image features and shallow image features. The first multi-scale corrosion recognition module is used to perform multi-scale corrosion feature recognition on the deep image features to obtain multi-scale stitched features. The patchy corrosion feature recognition module is used to identify the multi-scale spliced ​​features to obtain the patchy corrosion features; The linear corrosion feature recognition module is used to identify the multi-scale spliced ​​features to obtain linear corrosion features; The feature fusion module is used to fuse the spot corrosion features and the linear corrosion features to obtain fused features; Using the first decoding module, the shallow image features and the fused features are decoded to obtain the semantic segmentation image of the bridge surface.

3. The method according to claim 2, characterized in that, The first multi-scale corrosion recognition module includes: an image pooling layer, a first convolutional layer, four dilated convolutional layers, a stitching layer, a second convolutional layer, and an attention module; wherein the four dilated convolutional layers correspond to different dilation rates. The first multi-scale corrosion recognition module is used to perform multi-scale corrosion feature recognition on the deep image features to obtain multi-scale stitched features, including: The deep image features are pooled using the image pooling layer to obtain pooled features; The first convolutional layer is used to perform convolution processing on the deep image features to obtain the first convolutional features; Four dilated convolutional layers are used to perform dilated convolution processing on the deep image features respectively, resulting in four sets of dilated convolutional features; The splicing layer is used to splice the pooling feature, the first convolution feature, and the four sets of dilated convolution features to obtain the pooling convolution spliced ​​feature; The pooling convolution concatenation features are processed by the second convolution layer to obtain the second convolution features. The attention module is used to generate the multi-scale spliced ​​features based on the second convolutional features.

4. The method according to claim 3, characterized in that, The attention module includes: a channel attention module and a spatial attention module; The step of using the attention module to generate the multi-scale concatenated features based on the second convolutional features includes: The channel attention module is used to process the second convolutional features with channel attention to obtain channel attention features; The spatial attention module is used to perform spatial attention processing on the channel attention features to obtain spatial attention features, and the multi-scale stitching features are generated based on the channel attention features and the spatial attention features.

5. The method according to claim 1, characterized in that, The second deep semantic segmentation model includes: a preprocessing module, a color extraction module, a color parsing module, a color confidence module, a weighted feature concatenation module, and an output module; The second deep semantic segmentation model is used to perform semantic segmentation of the bridge cable image to identify cable corrosion and protective layer discoloration, resulting in a semantically segmented image of the bridge cable, including: The bridge cable image is converted from the RGB color space to the HSV color space to obtain an HSV image; The preprocessing module is used to decouple the channels of the HSV image to obtain H channel feature map, S channel feature map and V channel feature map; The color extraction module is used to extract color features from the H-channel feature map, the S-channel feature map, and the V-channel feature map to obtain H-channel color features, S-channel color features, and V-channel color features. The color analysis module is used to perform gradient perception on the H channel color features to obtain hue gradient perception features; Using a color confidence module, the H channel weights and S channel weights are obtained based on the hue gradient perception features and the S channel color features, respectively. The weighted feature stitching module is used to stitch features together based on the H channel weight, the hue gradient perception feature, the S channel weight, the S channel color feature, and the V channel color feature to obtain weighted stitched features. The output module is used to generate the semantic segmentation image of the bridge cables based on the weighted stitching features.

6. The method according to claim 5, characterized in that, The second deep semantic segmentation model also includes: a second multi-scale corrosion recognition module; The step of generating the semantic segmentation image of the bridge cable using the output module based on the weighted stitching features includes: Using the color analysis module, cross-channel fusion is performed based on the hue gradient perception feature, the S-channel color feature, and the V-channel color feature to obtain the cross-channel fusion feature; The second multi-scale corrosion identification module is used to identify multi-scale corrosion features based on the cross-channel fusion features, thereby obtaining multi-scale fusion features. The multi-scale fusion features are upsampled to obtain upsampled features; The output module is used to generate the semantic segmentation image of the bridge cable based on the weighted splicing features and the upsampling features.

7. The method according to claim 5, characterized in that, The color extraction module includes: a first-level extraction layer, a first pooling layer, a second-level extraction layer, and a second pooling layer; The color extraction module is used to extract color features from the H-channel feature map, the S-channel feature map, and the V-channel feature map to obtain H-channel color features, S-channel color features, and V-channel color features, including: Using the first extraction layer, the H-channel feature map, the S-channel feature map, and the V-channel feature map are convolved to obtain the first H-channel convolutional feature, the first S-channel convolutional feature, and the first V-channel convolutional feature. The first pooling layer is used to pool the first H channel convolutional features, the first S channel convolutional features, and the first V channel convolutional features to obtain hue output features, saturation output features, and brightness output features. The second-level extraction layer is used to perform convolution processing on the hue output feature, the saturation output feature map and the brightness output feature respectively to obtain the second H-channel convolution feature, the second S-channel convolution feature and the second V-channel convolution feature. A second pooling layer is used to pool the second H-channel convolutional features, the second S-channel convolutional features, and the second V-channel convolutional features to obtain the H-channel color features, the S-channel color features, and the V-channel color features.

8. The method according to claim 5, characterized in that, The color confidence module includes: a polar coordinate convolutional layer and a weight generation module; The color confidence module obtains the H channel weights and S channel weights based on the hue gradient perception features and the S channel color features, including: The polar coordinate convolutional layer is used to perform circular convolution on the hue gradient perception feature and the S-channel color feature to obtain the first circular distribution feature and the second circular distribution feature. The weight generation module generates the H-channel weight and the S-channel weight based on the first ring distribution feature and the second ring distribution feature.

9. The method according to claim 6, characterized in that, The output module includes: a feature modulation module, a second decoding module, and a third convolutional layer; The step of generating the semantic segmentation image of the bridge cable using the output module based on the weighted splicing features and the upsampling features includes: The feature modulation module is used to perform feature modulation based on the weighted splicing features and the upsampled features to obtain the modulation features; The second decoding module is used to upsample the modulation features and then perform jump feature concatenation to obtain jump concatenation features. The jump concatenation features are then convolved to obtain the decoding features. The third convolutional layer is used to perform convolution processing on the decoded features to generate the semantic segmentation image of the bridge cable.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: The number of pixels corresponding to the preset calibration object is obtained based on the object distance, image distance, horizontal pixel count of the image sensor, horizontal physical scale of the effective imaging area of ​​the image sensor, and actual size of the preset calibration object when the preset calibration object is photographed. Calculate the physical size of the pixels based on the number of pixels corresponding to the preset calibration object and the actual size of the preset calibration object; The actual corrosion area of ​​the bridge under inspection is calculated based on the total number of pixels in the marked area and the physical size of the pixels.