Bridge crack image segmentation system and method thereof
Patent Information
- Application Number
- CN202610458264.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-04-09
AI Technical Summary
[0006]本发明针对现有桥梁裂缝分割技术中存在的细小裂缝漏检率高、复杂背景下假阳性误检多、多维度特征协同利用不足的核心痛点,提供一种适配桥梁结构健康监测场景的高精度图像分割方案,通过双维度注意力机制的协同优化与针对性损失函数设计,实现复杂背景下桥梁裂缝的像素级精准分割
1.通过编码器逐层嵌入的SE模块与跳跃连接路径中IIA模块的有机结合,既实现了裂缝相关特征通道的全程强化,又完成了深浅层特征的空间精准对齐,有效提升了对低对比度、纤细、蜿蜒裂缝的检出能力与分割边界精度。
Smart Images

Figure CN121982045B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and artificial intelligence, specifically to an image segmentation system and method, which is particularly suitable for accurately segmenting crack targets from bridge surface images with complex backgrounds. Background Technology
[0002] As critical transportation infrastructure, the structural health of bridges is directly related to public safety. Cracks are one of the most common forms of bridge damage, making timely and accurate detection crucial. Traditional manual inspection methods are inefficient, subjective, and pose safety risks. With technological advancements, image-based automated inspection methods have become a research hotspot.
[0003] Early automatic detection methods mainly relied on traditional image processing techniques, such as thresholding and edge detection. These methods typically require manual parameter setting, are sensitive to changes in ambient lighting and background texture, and have poor robustness, especially for segmenting small cracks with low contrast and complex shapes.
[0004] In recent years, deep learning-based semantic segmentation techniques, especially U-Net and its variants, have achieved significant success in medical and remote sensing image segmentation and have been introduced into the field of bridge crack detection. These methods can automatically learn features, improving detection accuracy to some extent. For example, Chinese patent document CN115187621A (published on October 14, 2022) discloses a U-Net medical image contour automatic extraction network incorporating an attention mechanism, and US patent document US-20220309674-A1 (published on September 29, 2022) discloses a U-Net-based medical image segmentation method. However, bridge crack detection faces unique challenges: crack pixels account for a very small proportion of the entire image (extremely unbalanced categories); cracks are thin, irregular in shape, and often confused with background stains and textures; image acquisition is greatly affected by lighting and viewing angle. When general segmentation networks are directly applied to this scenario, they often suffer from high rates of missed detection of small cracks and high rates of over-detection (false positives) of complex backgrounds, making it difficult to meet engineering accuracy requirements.
[0005] In existing technologies, although some studies have attempted to introduce attention mechanisms into U-Net to improve feature selectivity, most focus on single-dimensional attention (such as channel attention or spatial attention), failing to fully utilize feature information from different dimensions. For example, Chinese patent document CN120276031B (granted on April 17, 2025) discloses a seismic data interpolation method based on the CBAM-Res2Unet network, which introduces a CBAM module in skip connections to achieve channel and spatial attention; Zilong Huang et al., in their paper "CCNet: Criss-Cross Attention for Semantic Segmentation" published in IEEE Transactions on Pattern Analysis and Machine Intelligence (Vol.45, No.6, pp.6896-6908, June 2023), proposed a cross-attention module for semantic segmentation. These methods have not specifically optimized the loss function design to address the key challenge of "suppressing background misclassification as cracks". Therefore, there is an urgent need for a dedicated segmentation scheme that can deeply integrate multi-dimensional attention and has targeted optimization objectives to achieve high-precision and robust segmentation of bridge cracks. Summary of the Invention
[0006] This invention addresses the core pain points of existing bridge crack segmentation technologies, such as high missed detection rate of small cracks, numerous false positives in complex backgrounds, and insufficient collaborative utilization of multi-dimensional features. It provides a high-precision image segmentation scheme adapted to bridge structural health monitoring scenarios. Through the collaborative optimization of a two-dimensional attention mechanism and the design of a targeted loss function, pixel-level accurate segmentation of bridge cracks in complex backgrounds is achieved.
[0007] To achieve the above objectives, the core technical concept adopted by this invention is as follows: At the network architecture level, an SA-UNet segmentation model based on U-Net improvement is constructed, forming a collaborative feature extraction link of "layer-by-layer channel filtering - spatial adaptive alignment". On the one hand, an SE channel attention mechanism module (hereinafter referred to as the SE module) is embedded in each layer of the feature encoder. During the progressive feature extraction process, the key channel features related to cracks are strengthened layer by layer, and the invalid information transmission of background interference channels is suppressed. On the other hand, an IIA interactive spatial attention module (hereinafter referred to as the IIA module) is introduced into the skip connection path between the encoder and the decoder to solve the problem of spatial misalignment of deep and shallow features, realize the adaptive and accurate fusion of encoder detail features and decoder semantic features, and improve the ability to capture the boundaries of thin and irregular cracks.
[0008] At the model optimization level, a dual-weighted focus loss function is designed to resolve the inherent contradiction between class imbalance and false positive suppression. Based on the standard focus loss, this function balances the extreme sample imbalance between cracks and background by adjusting class weights, focusing on feature learning of rare crack pixels. Simultaneously, it introduces a dynamic false positive penalty mechanism linked to the class weights. This improves crack detection sensitivity while simultaneously strengthening the penalty constraint for high-confidence misjudgments of background pixels, thereby reducing the risk of over-detection in complex backgrounds from the root cause of loss optimization.
[0009] At the data processing level, a preprocessing and data augmentation process was designed to fit the actual working conditions of bridge drone inspections. By optimizing parameters to adapt to on-site lighting and shooting angles, the physical rationality of the samples was ensured while expanding the dataset, thereby improving the model's generalization ability and engineering adaptability.
[0010] Furthermore, the technical solution provided by this invention can realize fully automated processing from the input of original bridge images collected by UAVs to the output of crack pixel-level segmentation masks and the quantification of disease parameters, providing stable, accurate and efficient technical support for bridge structural health monitoring.
[0011] Compared with the prior art, the present invention has the following significant advantages: 1. By organically combining the SE module embedded layer by layer in the encoder with the IIA module in the skip connection path, the full-process enhancement of crack-related feature channels is achieved, and the spatial alignment of deep and shallow features is completed, which effectively improves the detection capability and segmentation boundary accuracy of low-contrast, thin, and meandering cracks.
[0012] 2. The designed dual-weighted focus loss function, through the dynamic linkage mechanism of false positive penalty factor and class weight, fundamentally alleviates the trade-off between "improving crack recall rate" and "suppressing background false detection". Even under complex background stains and texture interference, it can still significantly reduce the false positive rate while maintaining an extremely low crack false detection rate.
[0013] 3. The complete solution covers the entire chain of data preprocessing, model training, and inference application. The data augmentation parameters are closely aligned with the actual working conditions of bridge inspection, and the model inference efficiency is on par with existing mainstream U-Net variants. No additional hardware costs are required, and it can be directly integrated into the bridge UAV intelligent inspection system, demonstrating strong engineering practicality. Attached Figure Description
[0014] Figure 1 This is an overall flowchart of the bridge crack image segmentation method provided by the present invention; Figure 2 This is a schematic diagram of the SA-UNet segmentation network model in this invention; Figure 3This is a flowchart illustrating the internal structure and data processing of the IIA module in this invention. Figure 4 This is a block diagram of the dual-weighted focus loss function in this invention and the calculation process of its dynamic false positive penalty factor. Detailed Implementation
[0015] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings. It should be noted that these descriptions are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0016] Overall process overview The bridge crack image segmentation scheme provided by this invention, based on the engineering needs of bridge inspection, constructs a fully automated processing chain from raw image input to quantitative output of crack defects. The core implementation process is divided into four core stages: dataset construction and standardized preprocessing, SA-UNet segmentation network construction and optimization, model supervised training and weight selection, on-site image inference and application of segmentation results. Each stage is interconnected, ensuring segmentation accuracy and engineering applicability. Figure 1 As shown, the overall process of the bridge crack image segmentation method provided by this invention starts with the acquisition of bridge surface images by UAV, and then proceeds sequentially through image screening and annotation, dataset partitioning and enhancement (including brightness adjustment [0.65, 1.35], random rotation [-90°, 90°], random horizontal / vertical flipping), construction of SA-UNet network (embedding SE and IIA modules), model training (double weighted focus loss), model inference, and finally output of crack segmentation mask. Each step is connected sequentially to form a complete segmentation technology link.
[0017] 1. Specific implementation of the data preparation and preprocessing module The goal of this step is to construct a high-quality, diverse training dataset. Its core tasks include data collection and filtering, precise labeling, and augmentation. (Corresponding...) Figure 1 The entire data processing steps, from UAV acquisition of bridge surface images to brightness adjustment [0.65, 1.35], random rotation [-90°, 90°], and random horizontal / vertical flipping, cover the entire process of image selection and annotation, dataset partitioning and enhancement.
[0018] 1.1 Data Collection and Filtering Drones equipped with high-definition cameras were used to acquire images of key components of the bridge, including the underside, piers, and beams. The acquired raw images underwent rigorous screening: first, an image sharpness evaluation algorithm automatically removed blurry images caused by drone shake; second, manual review removed invalid images with obvious geometric distortion, overexposure, or underexposure. This screening step aims to eliminate images where crack features have been distorted (e.g., loss of edge information due to blurring, or morphological changes due to distortion), ensuring that the training data accurately reflects the crack features to be extracted. This is a prerequisite for the model to achieve accurate feature extraction capabilities.
[0019] 1.2 Data Labeling and Format Conversion Pixel-level semantic annotation was performed on the selected images using professional annotation tools. During annotation, precise polygonal delineations were made along the crack edges, and pixels were categorized as "crack" and "background." The annotation information was then used to generate corresponding annotation files. Subsequently, a conversion program was written to convert the original images and annotation files into a standard dataset format commonly used in semantic segmentation tasks.
[0020] 1.3 Dataset Partitioning and Augmentation The labeled dataset was randomly divided into training, validation, and test sets in an 8:1:1 ratio. To improve the model's adaptability to complex real-world environments, data augmentation was performed on the training set. It is important to note that the data augmentation parameters in this invention are not copied from general computer vision tasks, but are specifically optimized based on statistical analysis of numerous real-world bridge detection scenarios. Existing data augmentation parameters (such as brightness adjustment range [0.5, 1.5], rotation angle range [-180°, 180°], etc.) are general settings and do not consider the specific characteristics of bridge crack detection: In actual engineering, due to equipment and environmental limitations, the illumination changes (brightness factor) of bridge images captured by drones rarely exceed the range [0.65, 1.35]. Augmentation samples exceeding this range are invalid or even interfering data. Similarly, bridge surface cracks have a clear physical direction; if the rotation angle exceeds the reasonable viewing angle range of [-90°, 90°], it will cause unrealistic distortion of the crack morphology, destroying its spatial structural features and misleading the model's learning. Specific augmentation methods and parameters are as follows: Brightness adjustment: The brightness adjustment factor is randomly adjusted within the range of [0.65, 1.35]. This range accurately covers most natural lighting and shadow change scenarios in bridge detection, avoiding interference to model feature extraction caused by the introduction of invalid samples with extreme lighting.
[0021] Geometric transformation: Random horizontal or vertical flipping is performed. At the same time, random rotation is performed within the rotation angle range of [-90°, 90°]. This range strictly matches the common shooting angle changes when drones inspect bridges, ensuring the authenticity and physical rationality of the enhanced crack morphology.
[0022] The combined application of the above enhancement operations aims to simulate the differences in image appearance caused by changes in lighting conditions and shooting angles in actual detection in a way that closely matches real working conditions. This forces the model to learn to maintain consistent crack feature representation under different imaging conditions, thereby improving the effectiveness of the enhanced data and the model's generalization ability.
[0023] 2. Specific implementation of the segmentation network model (SA-UNet) This section is one of the core innovations of this invention, constructing a segmentation network that integrates a dual attention mechanism. This model makes key improvements based on the encoder-decoder symmetric architecture, such as... Figure 2 As shown in the diagram, the SA-UNet segmentation network model of this invention clearly presents the encoder, decoder, skip connections, and the structural design of each layer. The input is a bridge image of 1024×1024×3. The encoder extracts features at four levels: the first stage outputs 512×512×64 features through convolution + SE module + downsampling; the second stage outputs 256×256×128 features through convolution + SE module + downsampling; the third stage outputs 128×128×256 features through convolution + SE module + downsampling; and the fourth stage outputs 64×64×512 features through convolution + SE module + downsampling. Then, the decoder performs four levels of upsampling and fusion. After each upsampling, the features are fused with the corresponding layer features of the encoder through the IIA module. Finally, a 1024×1024×1 crack segmentation mask is output. The core improvement features of this network are the embedding of SE modules in each layer of the encoder and the introduction of IIA modules in the skip connection paths.
[0024] 2.1 Feature Encoder Structure The encoder consists of four feature extraction layers, with input being a 1024×1024 resolution RGB image, corresponding to Figure 2The encoder has a Stage 1 to Stage 4 hierarchical structure. The SE module of this invention is embedded in each layer of the feature encoder, not just at the end of the network. This design is based on the logic of "progressive extraction" of bridge crack features in the encoder: shallow layers mainly capture low-level details such as crack edges and textures; mid-layer layers extract mid-level features such as crack shape and orientation; and deep layers learn the abstract semantic information of the crack. Channel recalibration at each layer allows for layer-by-layer adaptive filtering and enhancement of channels most relevant to the feature extraction target at that layer, while filtering out dominant invalid background features. If the SE module is only introduced in the last layer of the encoder, the unpurified invalid background features accumulated and transmitted from previous layers will interfere with the purity of the final high-level semantic features, causing the crucial crack feature information to be diluted or lost during transmission. Therefore, this "layer-by-layer embedding" design ensures the continuous enhancement of crack-related channels and the layer-by-layer suppression of background interference during feature extraction from low to high levels, providing high-quality feature input for subsequent spatial alignment.
[0025] Level 1: First, two 3×3 convolution operations are performed, with a kernel size of 64. After each convolution, a ReLU activation function is used for non-linear transformation. Then, the resulting feature map is input into the SE module. This module adopts a progressive architecture of "compression-activation-recalibration," and the specific process is as follows: Compression: Global average pooling is performed on the input feature map of size H×W×C, and the H×W spatial information of each channel is aggregated into a single scalar to form a 1×1×C channel descriptor vector.
[0026] Excitation: The vector passes through a bottleneck structure consisting of two fully connected layers to learn the correlation between learning channels. The first fully connected layer reduces the number of channels from C to C / r (where the compression ratio r is set to 16, a value that balances the preservation of key features and reduction of model complexity, given that bridge crack features are concentrated in a few channels). A ReLU activation function is then applied to introduce nonlinearity. The second fully connected layer restores the number of channels to C (the excitation unit uses a bottleneck structure consisting of two fully connected layers to learn the complex, nonlinear inter-channel dependencies between crack features and background features, in order to generate more accurate channel weights). The output is then normalized to the [0,1] interval by the Sigmoid function to generate a weight vector representing the importance of each channel.
[0027] Recalibration: The learned channel weight vector is multiplied with the original feature map channel by channel to achieve adaptive feature recalibration and output the enhanced feature map.
[0028] After the SE module is completed, downsampling is performed using 2×2 max pooling with a step size of 2, reducing the feature map size to 512×512.
[0029] Levels two through four: The network depth increases progressively, with the number of convolutional kernels doubling sequentially to 128, 256, and 512. The structure of each level is the same as the first level: two 3×3 convolutions + ReLU → SE module → max pooling downsampling. After four levels of encoding, the feature map size decreases to 64×64, and the number of channels increases to 512, resulting in feature representations containing rich high-level semantic information. Figure 2 The bottleneck layer features of the encoder output are 64x64x512.
[0030] 2.2 Feature Decoder Structure and Feature Fusion The decoder also consists of a four-level structure, symmetrical to the encoder, responsible for restoring high-level semantic features to the original image resolution and achieving accurate localization. Figure 2 The Stage 4 to Stage 1 hierarchical structure of the decoder. This invention introduces an IIA module into the skip connection path, designed to address the spatial feature fusion challenge of bridge cracks, which are characterized by "irregular morphology and multi-scale distribution." Cracks in images often exhibit meandering, discontinuous, and multi-scale morphologies, making it difficult to effectively align and fuse detailed and semantic information through simple skip connection stitching. Specifically, this invention places the IIA module on the skip connection path between the encoder and decoder (…). Figure 2 The placement of all "skip connections + IIA" (not within the decoder) is based on a deep understanding of the functionality and limitations of skip connections. In U-Net-like architectures, skip connections aim to fuse the rich spatial details (detail features) captured by the encoder with the contextual semantics (semantic features) recovered by the decoder. However, for bridge crack segmentation tasks, the crack detail features extracted by the encoder (such as precise edge locations) and the semantic features upsampled by the decoder (such as the overall discrimination of the crack region) are not completely consistent in spatial distribution, exhibiting an inherent "spatial misalignment." Directly concatenating channels would lead to feature blurring and reduce the accuracy of segmentation boundaries. The IIA module proposed in this invention, through its dual-path spatial interaction attention mechanism, can actively learn and correct this spatial misalignment on this critical path: it adaptively evaluates and weights the spatial positions in the encoder feature map that are crucial to the current decoding stage, guiding the detailed features and semantic features to achieve precise alignment and fusion. This is equivalent to adding an "intelligent spatial alignment filter" to the skip connections, a direct and efficient solution to the pain point of feature fusion on this path. This placement design is the optimal choice to maximize its functionality, rather than an obvious deployment method.
[0031] Starting with the deepest feature layer (64×64×512), each level first doubles the feature map size through an upsampling operation (such as transposed convolution). The upsampled features need to be fused with the features of the corresponding layer in the encoder path to recover spatial detail information. Before this fusion, the encoder-side features are processed by the IIA module, such as... Figure 3 As shown in the diagram, the internal structure and data processing flowchart of the IIA module in this invention detail the complete steps from the input feature maps B, C, H, W, through biorthogonal reconstruction along the height / width dimensions, parallel 1D average pooling and 1D max pooling, concatenation, 1D convolution + linear rectification, and sigmoid function activation to generate height and width weight vectors respectively. Then, an outer product is used to generate a 2D spatial attention weight map, followed by element-wise multiplication weighting and residual connections, finally outputting the fused feature map. This visually presents the module's core innovative architecture of "orthogonal decomposition - parallel modeling - outer product fusion." The design of this module is precisely to address the spatial extension and morphological complexity of bridge cracks. Through a dual-path spatial attention mechanism, the network is guided to focus on key spatial locations related to cracks during fusion, achieving adaptive optimization fusion of deep and shallow features, thereby avoiding the loss of details or semantic confusion caused by simple concatenation.
[0032] The SE module and the IIA module form a collaborative optimization link for bridge crack segmentation: "channel filtering → spatial alignment". The SE module performs channel recalibration at each layer of the encoder, filtering out background interference step by step and strengthening crack-related channel features, thus providing the IIA module with "cleaned" high-quality feature input. The IIA module, in turn, utilizes these channel-filtered features in the skip connection path to perform precise interaction and alignment in the spatial dimension, ensuring that the crack features strengthened by the SE module can be accurately located and reconstructed. This collaborative mechanism produces a "1+1>2" effect: if only the SE module is used without introducing the IIA module, the network can highlight crack channel features, but lacks the ability to correct spatial misalignment of features at different depths, resulting in inaccurate feature localization after strengthening; conversely, if only the IIA module is used without the preceding SE channel filtering, the features transmitted through skip connections will contain a large number of redundant and interfering channels, making it difficult for the IIA module to effectively focus on the true crack spatial context. The organic combination of the two methods has solved the dual challenges of extracting features from bridge cracks, namely, the need for channel reinforcement for small cracks and the need for spatial alignment for irregular cracks, thus significantly improving segmentation accuracy and robustness.
[0033] The specific operation procedure for the IIA module is as follows: The feature reconstruction unit reconstructs the input encoder feature map (dimensions B, C, H, W) using orthogonal perspectives. Specifically, during reconstruction along the height dimension, the feature map is replaced from (B, C, H, W) to (B, W, C, H) format, that is, the width dimension W is replaced with the "channel" dimension. This perspective focuses on capturing the dependencies of features along the image height direction (vertical direction), which is crucial for identifying vertically extending crack segments. During reconstruction along the width dimension, the feature map is reconstructed to (B, H, C, W) format. This perspective focuses on capturing the dependencies of features along the image width direction (horizontal direction), which is beneficial for capturing horizontally meandering crack features. This biorthogonal perspective reconstruction mechanism is designed for the irregular spatial morphology and variable extension direction of bridge cracks. It breaks away from the conventional spatial attention approach (such as the spatial attention module in CBAM) which only performs single-weighted graph convolution on a uniform H×W two-dimensional spatial map, and instead decouples and models long-range spatial dependencies from two geometrically orthogonal dimensions.
[0034] Spatial Weight Generation Unit: For the two reconstructed features (height-view features and width-view features), parallel, structurally identical but parameter-independent sub-modules are used to generate corresponding spatial attention weight maps. Each sub-module operates as follows: First, the input feature descriptions are processed simultaneously using one-dimensional global average pooling and one-dimensional global max pooling to aggregate information and generate two one-dimensional feature vectors. Next, these two one-dimensional vectors are concatenated and fed into a lightweight network consisting of two 1D convolutional layers (containing a non-linear activation function, such as ReLU), to learn the spatial importance distribution along this dimension. Finally, a one-dimensional spatial attention weight vector is generated using the Sigmoid function. This process is executed in parallel along both the height and width paths, generating height and width weight vectors respectively. This is fundamentally different from "general spatial attention," which generates weight maps only on a single two-dimensional plane.
[0035] The feature weighting and fusion unit combines the generated height and width weight vectors through an outer product operation to generate a two-dimensional attention matrix that integrates the spatial weights from both perspectives. Then, this two-dimensional attention matrix is element-wise multiplied with the original input feature map to differentially enhance or suppress features at different spatial locations. Finally, the weighted features are added to the original input features through a residual connection to output the fused feature map.
[0036] Structural Differences and Performance Advantages Compared to Existing Spatial Attention Modules (such as CBAM and CCNet): The core mechanism of the IIA module proposed in this invention differs fundamentally from existing general spatial attention modules (such as the spatial attention module in CBAM or the cross-attention module in CCNet). General spatial attention typically performs global convolution or self-attention calculation directly on an H×W two-dimensional spatial graph to generate a single, isotropic spatial weight graph. This approach is effective for blocky or isotropic targets, but it is difficult to accurately model typical linear, irregular, and detailed structures such as bridge cracks. Cracks often meander along a dominant direction (such as vertical, horizontal, or oblique) in space, and their effective features are distributed in narrow local regions. General global attention is easily diluted by large background areas, resulting in insufficient attention to small cracks. The IIA module of this invention adopts an innovative architecture of "orthogonal decomposition-parallel modeling-outer product fusion," whose design principle directly addresses the spatial characteristics of linear targets: by reconstructing features along two geometrically orthogonal dimensions—height and width—and generating a one-dimensional weight vector, it can independently and sensitively capture the long-range spatial dependencies of cracks along different directions. This dual-path decoupling design endows the IIA module with strong directional sensitivity, enabling it to adaptively strengthen key spatial contexts along the crack extension direction while suppressing redundant information in the vertical direction. This is fundamentally different from the "one-size-fits-all" approach of general attention. Therefore, in bridge crack segmentation tasks, the IIA module can more effectively focus on the crack skeleton and edge details, achieving precise spatial alignment and fusion of encoder detail features and decoder semantic features. This results in clearer and more coherent segmentation boundaries in complex contexts, significantly improving the model's detection rate and segmentation accuracy for small, discontinuous cracks. Experimental data shows that the IIA module significantly improves boundary accuracy compared to general spatial attention mechanisms, validating its superiority in spatial modeling of linear and irregular targets.
[0037] 3. Model Training and Specific Implementation of Loss Function This section details the model optimization process and the specially designed loss function, which is another core innovation of this invention. For example... Figure 4 As shown, the block diagram of the dual-weighted focus loss function and the calculation process of its dynamic false positive penalty factor clearly demonstrate the input of the model's predicted probability and the true label. After class weight allocation (high weight for cracks, low weight for background), background pixel branch judgment and corresponding false positive penalty factor calculation, easy and difficult sample adjustment, and log probability calculation, the final loss value L = -W is obtained. t (1-P t )^γ log(P t ) FP The entire process of calculating and outputting the final loss value clarifies the components of the loss function, the calculation logic of each factor, and the linkage between the dynamic false positive penalty factor, class weight, and misjudgment confidence. This is a key diagram for understanding the core innovation of this loss function.
[0038] 3.1 Training Configuration The model is trained end-to-end using the prepared training and validation sets. An adaptive moment estimation algorithm can be used as the optimizer, with an initial learning rate of 1e-4. The total number of training epochs can be set to 300, with performance evaluated on the validation set after each epoch and model weights saved periodically. The training batch size is configured based on computational resources.
[0039] 3.2 Double-weighted focus loss function To address the challenges of extreme class imbalance and false positives in complex backgrounds during bridge crack segmentation, this invention designs a dual-weighted focus loss function. The core innovation of this function lies in its first-ever revelation of the inherent contradiction between "class weight setting" and "false positive penalty": In the bridge crack segmentation scenario, simply increasing the crack class weight (W...) will not solve the problem. crack To mitigate class imbalance, the model becomes overly sensitive to crack features, exacerbating the risk of misclassifying complex background textures as cracks (false positives). Conversely, using only a fixed false positive penalty mechanism fails to adapt to changes in misclassification risk caused by dynamic adjustments in class weights. Existing technologies design class weights and false positive suppression measures independently, failing to resolve this inherent contradiction. The overall structure of the loss function, the calculation logic of each factor, and their interrelationships are as follows: Figure 4 As shown.
[0040] The loss function of this invention incorporates three targeted design features based on the standard focus loss, and its mathematical expression is as follows: L=-W t (1-P t )^γ log(P t ) F P Category weight W t Used to solve the class imbalance problem. Assigns a higher weight W to the crack class. crack The background category has a lower weight W background This emphasizes the importance of rare crack pixels in the loss calculation, forcing the model to learn crack features.
[0041] Difficulty sample adjustment factor (1-P) t )^γ:where P tγ is the model's predicted probability for the true class t, and γ is the focusing parameter, which is set to 2 in this implementation. This value is based on the analysis of the proportion of hard-to-classify samples (such as blurred edges and small cracks) in the bridge crack dataset, aiming to ensure that the training process accurately focuses on the feature learning of such key samples, avoiding underfitting or overfitting due to improper weight settings.
[0042] Dynamic false positive penalty factor F P This is a unique core design feature of this invention, specifically designed to suppress over-detection where the background is misidentified as a crack. Its definition is as follows:
[0043] When the true label t of a pixel is "background", but the model outputs a high confidence p crack This penalty will be triggered when it is predicted as a "crack".
[0044] The penalty intensity varies with the model's misjudgment confidence level p. crack The increase is exponential (r) fp (A penalty index greater than 0). Simultaneously, the penalty intensity is proportional to the category weight W. crack / W background Proportional, which ensures that the model is able to detect cracks (high W) crack At the same time, the constraint on background misjudgment is also strengthened.
[0045] This mechanism can effectively suppress false positive predictions caused by complex background textures, stains, etc., while maintaining the model's high sensitivity to cracks, thus greatly improving the reliability of segmentation results.
[0046] The dynamic false positive penalty factor (F) designed in this invention P ) and category weights (W t The linkage mechanism is the core innovation of this loss function, aiming to solve a key contradiction in bridge crack segmentation: increasing the class weight (W) of sparse crack pixels to improve their recall rate. crack When the model becomes overly sensitive to crack features, it is more likely to misclassify complex background textures (such as stains and shadows) as cracks, leading to an increased false positive rate. In traditional methods, class weight adjustment and false positive suppression are independent strategies, failing to resolve the inverse relationship between them. The innovation of this invention lies in designing the false positive penalty factor to be correlated with the "false positive confidence (p... crack ")" and "Category Weight Ratio (W)" crack / W background The product of W and W is proportional. Specifically, when the model increases W to learn crack features... crack At that time, in the penalty factor (W) crack / W backgroundThe ratio of ) increases synchronously, which means that for any behavior that misclassifies the background as a crack (especially high-confidence misclassification p) crack The penalty intensity automatically increases when the value of the crack is high. This dynamic linkage mechanism allows the model to focus on cracks (through high W) when forced to pay attention to them. crack At the same time, it is also required to more strictly distinguish between background and cracks (through synchronously enhanced F). P This allows the system to automatically seek the optimal balance between recall and precision during the optimization process, fundamentally mitigating the false positive side effect caused by simply increasing category weights.
[0047] To verify the structural innovation and generalization ability of the proposed dynamic linkage loss function, we conducted generalization experiments on other public datasets with similar characteristics of "extreme class imbalance" and "complex background interference." For example, on the DRIVE retinal vessel segmentation dataset, the pixel ratio of vessels and background is also significantly different, and there is interference from uneven brightness; on the Massachusetts Roads aerial imagery road extraction dataset, distinguishing between narrow roads and background also faces similar challenges. Experiments show that applying the dual-weighted focus loss function of this invention (preserving its dynamic linkage mechanism) to benchmark models (such as U-Net) for these tasks, compared to using standard cross-entropy loss or standard focus loss, without adjusting its core parameters (such as the penalty exponent r), results in significantly better performance. fp Even under conditions where the false positive rate (FPR) is reduced by an average of approximately 18%-25%, the IoU metric is maintained or slightly improved. This demonstrates that the design principle of "dynamic coordination between class weights and false positive penalties" revealed in this invention is not only applicable to bridge crack scenarios, but also has universal guiding significance for a large class of slender / irregular target segmentation tasks with sparse foreground pixels, complex backgrounds, and easy confusion. Its idea of resolving contradictions through a linkage mechanism represents a transferable innovation in loss function structure.
[0048] 4. Model Reasoning and Application After training, the model performance is evaluated on an independent test set. The model weights that achieve the best average intersection-union ratio and other metrics on the validation set are selected for actual deployment. This step corresponds to... Figure 1 The model inference and output crack segmentation mask stages after training the model (double-weighted focus loss).
[0049] For a new image of a bridge to be detected, the same standardized preprocessing (such as resizing and normalization) as during training is first performed. Then, it is fed into a segmentation network model with optimal weights.
[0050] The model performs forward propagation, outputting a probability map of each pixel belonging to a crack. By setting a threshold (e.g., 0.5) for binarization, a clear pixel-level crack segmentation mask map can be obtained. Figure 1 The final output is the crack segmentation mask.
[0051] Based on the generated accurate mask image, further morphological post-processing (such as small connected component removal) can be performed to optimize the results and quantify key geometric parameters such as crack length, average width, maximum width, and area. These quantification results provide objective and accurate data support for the automated assessment of bridge structural health status, safety level classification, and the development of targeted maintenance measures.
[0052] 5. Comparative Experiments and Effect Verification To objectively evaluate the effectiveness of the dynamic false positive penalty factor and overall technical solution proposed in this invention, we designed a rigorous comparative experiment.
[0053] 5.1 Experimental Setup Test dataset: The test set portion of the publicly available bridge crack dataset SDNET2018, which contains complex backgrounds, is used. This dataset includes various bridge deck and component images, with a large amount of dirt, texture, and shadow interference in the background, which can effectively test the robustness and generalization ability of the model.
[0054] Comparison Model: The model of this invention (Ours): a complete SA-UNet model equipped with an SE module, an IIA module, and a dual-weighted focus loss function (including a dynamic false positive penalty factor).
[0055] Comparison Model A (Att-UNet): Employs a public / typical U-Net variant model that integrates spatial and channel attention, representing a typical application of current attention mechanisms in segmentation, and is trained using the standard cross-entropy loss function.
[0056] Comparison Model B (Focal Loss UNet): For a fair comparison, this model uses the exact same network backbone and attention module as the present invention (i.e., the SA-UNet infrastructure, including SE and IIA modules), only replacing the loss function with the standard focal loss (Focal Loss, γ=2), and does not include the dynamic false positive penalty factor and class weight linkage mechanism designed in the present invention. This comparison aims to verify the contribution of the loss function of the present invention separately.
[0057] Training and Evaluation: All models were trained to convergence on the same training set (processed using the preprocessing procedure of this invention) and evaluated on the same SDNET2018 test set. The main evaluation metrics included: Intersection over Union (IoU), False Positive Rate (FPR), False Negative Rate (FNR), and the comprehensive metric F1-Score.
[0058] 5.2 Experimental Results and Analysis The table below shows the quantitative comparison results of each model on the SDNET 2018 test set:
[0059] Results analysis: Advantages in segmentation accuracy (IoU): The model of this invention achieved a crack IoU of 78.2%, significantly higher than the 73.3% of the comparative model A and the 75.8% of the comparative model B. This indicates that the collaborative design of the SE and IIA modules in this invention, combined with a targeted loss function, can more effectively learn and segment complete crack regions.
[0060] Significant effect in suppressing false positives: The model of this invention controls the false positive rate (FPR) at 5.7%, which is approximately 62.5% lower than the contrast model B (15.2%) using standard focus loss (i.e., the present invention reduces false positives by approximately 62.5% relative to the contrast model B). Even compared with the attention mechanism model A, the false positive rate is reduced by approximately 34.5%. This result strongly demonstrates the effectiveness of the dynamic false positive penalty factor (FPR) proposed in this invention. P The effectiveness of this factor is demonstrated by linking the false positive confidence with the category weight, which imposes a strong penalty on situations where "the background is falsely identified as a crack with high confidence," thereby greatly reducing over-detection under complex background interference.
[0061] Balance in overall performance: While significantly reducing the false positive rate, the false negative rate (FNR) of this invention remained at a low 13.5%, indicating that it did not miss real cracks due to suppressing overdetection. Ultimately, the invention achieved an F1-Score of 0.861, comprehensively outperforming both comparative models.
[0062] No significant loss in efficiency: The number of parameters and inference time of the model in this invention are basically the same as those of the structurally similar comparative model B, indicating that the performance improvement does not come from the increase in model complexity, but from the more efficient attention mechanism and loss function design.
[0063] Conclusion: The comparative experimental data above fully demonstrate that the scheme designed in this invention for bridge crack segmentation, which integrates a dual attention mechanism and a dual-weighted focus loss function (especially the dynamic false positive penalty factor), can achieve more accurate segmentation (higher IoU) and significantly suppress background false detections (lower false positive rate) compared to existing attention models and general loss functions. This improves the reliability and practicality of the segmentation results and meets the technical requirements of high accuracy and low false alarms for bridge health monitoring.
[0064] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and these variations still fall within the protection scope of the present invention.
Claims
1. A bridge crack image segmentation system, characterized in that, include: The image acquisition module is used to acquire images of the bridge surface to be segmented. The data preprocessing module is used to preprocess the image to be segmented or the training image; The segmentation network model is the SA-UNet model, which includes a feature encoder, a feature decoder, and skip connection paths connecting their corresponding layers. Each layer of the feature encoder includes an SE channel attention mechanism module, which is used to perform adaptive recalibration of the input feature map in the channel dimension; an IIA interactive spatial attention module is set on the skip connection path, which is used to realize adaptive fusion of encoder features and decoder features in the spatial dimension. The model training module is used to train the segmentation network model using an image dataset containing bridge crack annotations and a dual-weighted focus loss function as the optimization objective, so as to obtain the trained model weights. The segmentation execution module is used to load the trained model weights and perform pixel-level segmentation on the image to be segmented after being processed by the data preprocessing module through the segmentation network model, and output the crack segmentation mask. in: The IIA interactive spatial attention module includes: a feature reconstruction unit, used to reconstruct the input feature map along the height and width dimensions respectively to obtain feature representations from two orthogonal perspectives; a spatial weight generation unit, used to generate corresponding spatial attention weight maps from the feature representations from the two orthogonal perspectives respectively through parallel pooling, convolution, and activation operations; and a feature weighting and fusion unit, used to multiply the spatial attention weight map with the corresponding reconstructed features respectively, then restore the weighted features and add them to the original input features to output the fused feature map. The dual-weighted focus loss function L is defined by the following formula: L=-W t ×(1-P t )^γ×log(P t )×F P Among them, W t The class weights are used to balance the loss contributions of cracks and background pixels; (1-P) t )^γ is the standard focus loss factor, used to reduce the weight of easily classified samples; F P As a false positive penalty factor; The false positive penalty factor F P The definition is as follows: When the real label t is the background, F P =1+(p crack )^{r fp }×(W crack / W background ); When the true label t is a crack, F P =1; Where, p crack r is the confidence level for the model to predict that the pixel is a crack. fp W is a penalty index greater than 0. crack and W background The category weights are for the crack and the background, respectively.
2. The bridge crack image segmentation system according to claim 1, characterized in that, The data preprocessing module includes a data construction submodule for constructing a training dataset. Specifically, it includes: filtering the original images acquired by the image acquisition module and removing images that do not meet the quality requirements; annotating the crack regions of the filtered images to generate semantic segmentation annotation files; and dividing the annotated dataset into training, validation, and test sets according to a set ratio.
3. The bridge crack image segmentation system according to claim 2, characterized in that, The data preprocessing module further includes a data augmentation submodule, used to perform augmentation operations on the training dataset. The augmentation operations include at least one of the following: brightness adjustment, with the adjustment factor range set to [0.65, 1.35]. Rotation operation, with the rotation angle range set to [-90°, 90°]; Flip operation, including random flipping in the horizontal and vertical directions.
4. The bridge crack image segmentation system according to claim 1, characterized in that, The feature encoder comprises a four-layer structure, each layer including: two convolutional layers, the SE channel attention mechanism module, and a downsampling layer; the feature decoder comprises a four-layer structure, each layer including: an upsampling layer and two convolutional layers; the skip connection path fuses the output feature maps of each layer of the encoder via the IIA interactive spatial attention module before upsampling at the corresponding decoder layer.
5. The bridge crack image segmentation system according to claim 4, characterized in that, The SE channel attention mechanism module includes: a compression unit for performing global average pooling on the input feature map to generate channel statistical descriptors; an activation unit containing two fully connected layers for performing nonlinear transformations on the channel statistical descriptors to learn the correlation between channels and generate weight coefficients for each channel; and a recalibration unit for multiplying the weight coefficients with the original input feature map channel by channel to output the enhanced feature map.
6. A method for segmenting bridge crack images, characterized in that, The method, applied to the bridge crack image segmentation system as described in any one of claims 1 to 5, comprises the following steps: We acquire images of the bridge surface, construct and preprocess them to obtain an enhanced semantic segmentation dataset of bridge cracks; A SA-UNet segmentation network model is constructed, in which SE channel attention mechanism modules are embedded in each layer of the feature encoder, and IIA interactive space attention modules are introduced in the skip connection path between the encoder and the decoder. Using the enhanced dataset, the training of the segmentation network model is supervised with a double-weighted focus loss function to obtain the optimal training weights; The trained optimal weights are used to perform forward inference on the new bridge surface image to achieve pixel-level crack segmentation and output the results.
Citation Information
Patent Citations
U-Net medical image contour automatic extraction network fusing attention mechanism
CN115187621A
Seismic data interpolation method and system based on CBAM-Res2Unet network
CN120276031B
Medical image segmentation method based on u-net
US20220309674A1
Concrete crack segmentation method based on SCSEOCUnet
CN111353396A