Accurate micro-crack segmentation method integrating feature fusion and convolution attention

By constructing an encoder-decoder architecture and combining a convolutional block attention module and a feature fusion module, the problems of insufficient long-range dependency modeling and poor robustness in complex backgrounds in existing crack detection technologies are solved. This achieves accurate segmentation of microcracks and integrity of topological structure, thereby improving detection accuracy and robustness.

CN120876869AActive Publication Date: 2025-10-31DALIAN UNIV OF TECH

Patent Information

Application Number
CN202511383286.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing crack detection technologies suffer from insufficient long-range dependency modeling, imprecise characterization of extremely fine cracks, and poor robustness in complex backgrounds. In particular, convolutional neural networks struggle to capture the overall topological structure of narrow cracks and are prone to losing edge information in complex backgrounds, resulting in incomplete segmentation results.

Method used

We employ a micro-crack segmentation method that integrates feature fusion and convolutional attention. By constructing an encoder-decoder architecture and introducing a convolutional block attention module (CBAM) and a feature fusion module (FFM), we enhance multi-scale features at the encoder end and perform cross-layer fusion at the decoder end, thereby achieving accurate capture of crack saliency features and effective suppression of complex background interference.

Benefits of technology

It significantly improves the detection sensitivity and overall segmentation consistency of microcracks, reduces the false detection rate, ensures the continuity and topological integrity of narrow cracks, and enhances the robustness and generalization ability of the model in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876869A_ABST
    Figure CN120876869A_ABST
Patent Text Reader

Abstract

The invention provides a microcrack precise segmentation method integrating feature fusion and convolution attention, and belongs to the field of image processing. According to the method, a crack segmentation network based on an encoder-decoder architecture is constructed, a convolution block attention module is introduced at an encoder end, background noise is adaptively suppressed and obvious characteristics of cracks are enhanced through a channel and space dual attention mechanism, and the method is suitable for the adaptive segmentation of the cracks on the premise of almost not increasing the calculation overhead. The sensitivity of the model to microcracks is improved; a feature fusion module is introduced at a decoder end, and cooperation of low-layer details and high-layer semantics is realized through cross-layer fusion, so that a semantic gap is effectively bridged, detail loss caused by traditional convolution stacking is avoided, and continuity and a complete topological structure of a long and narrow crack are ensured. According to the method, through collaborative optimization of multi-scale feature extraction and an attention mechanism, accurate capture of the saliency features of the crack and effective suppression of complex background interference are realized, and the detection sensitivity and overall segmentation consistency of the micro-crack are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and image processing technology, specifically relating to a method for accurate micro-crack segmentation that integrates feature fusion and convolutional attention. Background Technology

[0002] Cracks are one of the most common forms of damage to civil infrastructure and high-end manufacturing materials during service. Their morphology and evolution patterns are often directly related to the safety and durability of structures. Therefore, timely detection and accurate identification of cracks have significant scientific and engineering application value in the health monitoring of large-scale engineering structures such as bridges, roads, tunnels, and buildings, as well as in high-end manufacturing fields such as composite materials and electronic packaging.

[0003] Existing crack detection technologies mainly fall into two categories: one is traditional manual inspection and image processing methods. Manual inspection relies on the experience of inspectors, which has inherent drawbacks such as low efficiency, high cost, and strong subjectivity, making it difficult to meet the consistency and repeatability requirements of large-scale engineering applications. Detection methods based on traditional image processing, such as threshold segmentation and edge detection, although computationally simple, are prone to inaccurate detection and insufficient robustness in complex backgrounds, lighting changes, or scenarios with small cracks.

[0004] Another category is automated detection methods based on deep learning. In recent years, Convolutional Neural Networks (CNNs), with their superior feature extraction capabilities, have become the mainstream solution for crack detection, and their accuracy has been continuously improved through advancements such as multi-scale feature fusion and attention mechanisms. However, these methods still have the following shortcomings:

[0005] Convolutional neural networks (CNNs) are limited by their local receptive field and lack the ability to model long-range dependencies, making it difficult to capture the overall topological structure of narrow cracks. During layer-by-layer downsampling and convolution stacking, edge and texture information of extremely fine cracks is easily lost, leading to incomplete segmentation results. Under complex background conditions (such as texture interference, illumination variations, and noise pollution), crack features are easily masked, causing false positives and false negatives. Meanwhile, some studies have attempted to introduce Transformers into crack detection tasks to leverage their global modeling advantages. However, since cracks often occupy only a very small proportion of an image, directly using a general Transformer can easily introduce redundant background information, diluting the discriminative signal and negatively impacting detection accuracy.

[0006] In summary, existing technologies still suffer from problems such as insufficient long-range dependency modeling, imprecise characterization of extremely fine cracks, and poor robustness in complex backgrounds in crack detection tasks. Summary of the Invention

[0007] To address the aforementioned problems in existing technologies, this invention proposes a precise micro-crack segmentation method integrating feature fusion and convolutional attention. This method constructs a novel crack segmentation network that combines a Feature Fusion Module (FFM) and a Convolutional Block Attention Module (CBAM). Based on an encoder-decoder architecture, the network introduces a Convolutional Block Attention Module (CBAM) at the encoder end and a Feature Fusion Module (FFM) at the decoder end. Through the synergistic optimization of multi-scale feature extraction and attention mechanisms, it achieves accurate capture of salient crack features and effective suppression of complex background interference, significantly improving the detection sensitivity and overall segmentation consistency of micro-cracks.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0009] A precise micro-crack segmentation method integrating feature fusion and convolutional attention includes the following steps:

[0010] S1. Obtain crack images and preprocess them to construct a crack image dataset;

[0011] S2. Construct a crack segmentation network, wherein the crack segmentation network is an encoder-decoder structure;

[0012] The encoder includes a deep residual network and a convolutional block attention module (CBAM); the deep residual network serves as the backbone network to extract multi-scale features from the input image, and the convolutional block attention module (CBAM) performs saliency enhancement on the multi-scale features to obtain enhanced features for each scale; the enhanced features for each scale include high-level semantic features and low-level detail features.

[0013] The decoder includes a feature fusion module (FFM) and a decoder head. The feature fusion module (FFM) performs cross-layer fusion and refinement of the low-level detail features and high-level semantic features output by the encoder to obtain fused features that have both global semantics and local details. The decoder head processes the fused features sequentially through convolution and progressive upsampling to gradually restore spatial resolution and enhance edge detail expression to obtain a crack segmentation prediction result map.

[0014] S3. The crack segmentation network is trained using the crack image dataset constructed in S1.

[0015] S4. Input the detected image to be segmented into the trained crack segmentation network. The encoder extracts multi-scale features and gradually fuses and restores the resolution in the decoder. Finally, the decoder outputs the segmentation result of the crack region.

[0016] Furthermore, in S1, the specific process of constructing the crack image dataset includes:

[0017] Capture crack images of roads, bridges, tunnels, building facades, etc., using cameras or drones, ensuring that cracks of different scales and shapes are included;

[0018] The crack regions in the acquired crack images are annotated at the pixel level to generate a binary mask corresponding to the original crack image.

[0019] The labeled crack images are preprocessed, including image normalization, cropping, flipping, rotating, scaling and brightness adjustment, to enhance data diversity and the generalization ability of the model, forming a crack image dataset.

[0020] The crack image dataset is divided into training, validation, and test sets to ensure the reasonable distribution of scenes.

[0021] Furthermore, the deep residual network includes a ResNet50 model, which takes the preprocessed crack image as input and extracts multi-scale feature maps from shallow to deep by performing layer-by-layer convolution and downsampling operations on the crack image. It includes shallow features Preserve details such as edges and textures, high-level features It contains more abstract global semantic information.

[0022] Furthermore, the Convolutional Block Attention Module (CBAM) aims to enhance the representational ability of convolutional neural networks by explicitly modeling the saliency of channels and spatial dimensions. It consists of a cascaded Channel Attention Module (CAM) and Spatial Attention Module (SAM). This dual attention mechanism achieves synergistic enhancement of key channels and key spatial locations with almost no increase in computational overhead, while suppressing redundant or noisy information, significantly improving the model's sensitivity and robustness in fine-grained tasks such as crack segmentation.

[0023] The CAM extracts multi-scale feature maps from the deep residual network. As input, global average pooling and global max pooling are used to compress information in the spatial dimension, respectively, to obtain two channel-level descriptors. :

[0024]

[0025]

[0026] in, , These are the height and width of the input image, respectively. Number of channels; , These represent the row and column coordinates of the feature map in the spatial dimension, respectively.

[0027] Will , A multilayer perceptron (MLP) with shared parameters is fed into the machine to learn nonlinear interaction relationships; the MLP is then reduced in dimensionality. To control the bottleneck structure, the multilayer perceptron consists of a first fully connected layer, a ReLU activation function, and a second fully connected layer connected sequentially. Its output is then processed... After activation by the activation function, the channel attention map is obtained. The process is represented as:

[0028]

[0029] in, for Activation function;

[0030] The enhanced features are then obtained by multiplying each channel sequentially. :

[0031]

[0032] The SAM enhances the features of the channel. Perform average pooling and max pooling along the channel dimension to generate two two-dimensional spatial descriptors. ;

[0033]

[0034]

[0035] in, Indicates the number of channels. ;

[0036] Will , After being spliced ​​along the channel axis, via a Convolutional layers capture local cross-channel interactions and are then processed... Spatial attention map is obtained by activation function. :

[0037]

[0038] in, This represents a convolutional layer with a kernel size of [size missing]. ;

[0039] Finally, the spatial attention map Compared with enhanced features Element-wise multiplication yields the refined feature tensor. ;

[0040] The shallow features , The feature tensors obtained after feature enhancement by the convolutional block attention module are respectively , Constitutes low-level detailed features , These represent the height, width, and number of channels of the low-level detail feature tensor, respectively; the high-level feature... , The feature tensors obtained after feature enhancement by the convolutional block attention module are respectively , Constitutes high-level semantic features , These represent the height, width, and number of channels of the high-level semantic feature tensor, respectively.

[0041] Furthermore, the Feature Fusion Module (FFM) includes a channel alignment module, a channel attention enhancement module, a correlation enhancement module, and a deep convolutional refinement module;

[0042] The channel alignment module will incorporate high-level semantic features. and low-level details The number of channels is adjusted to a uniform dimension C using 1×1 convolution, and bilinear interpolation is used to apply the high-level semantic features. Upsampling is performed to improve its spatial resolution and low-level detail features. Consistency is achieved, resulting in high-level semantic features after dimension alignment. and low-level detail features ;

[0043] The channel attention enhancement module respectively... and Channel enhancement is performed using the following formula:

[0044]

[0045]

[0046] in, This represents the channel weight vector for high-level semantic features. ; This represents the channel weight vector for low-level detail features. ; This is a global average pooling operation; , These are the enhanced high-level semantic features and low-level detail features, respectively.

[0047] The correlation enhancement module is used to promote deep interaction and complementary fusion between high-level semantic features and low-level detailed features; firstly, the features are... and Expanding in space:

[0048]

[0049] Calculate the interaction matrix:

[0050]

[0051] in, The interaction matrix represents the similarity between high-level semantic features and low-level detail features in the channel dimension.

[0052] Then the interaction matrix After activation processing, it is used as the fusion weight. :

[0053]

[0054] For fusion weights Complementary weighted fusion is performed to obtain the fusion features after cross-domain interaction. :

[0055]

[0056] in, To convert the matrix shape into a feature map shape; for identity matrix, guaranteeing and Complementary;

[0057] The deep convolutional refining module refines the fused features. After a series of lightweight convolutional refinements, the final fused features are obtained, which preserve both low-level edge information and high-level semantic global information. This process can be represented as follows:

[0058]

[0059] in, For the final fusion feature; For depthwise separable convolution, the receptive field is expanded while the number of parameters is reduced; For batch normalization; It is a non-linear activation.

[0060] Furthermore, the decoding head includes a convolutional block attention module (CBAM), an upsampling module, and a prediction head;

[0061] The Convolutional Block Attention Module (CBAM) has the same architecture as the Convolutional Block Attention Module in the encoder to fuse features. As input, the data is processed sequentially by a cascaded channel attention submodule (CAM) and spatial attention submodule (SAM) to obtain attention-enhanced features. The process is represented as:

[0062]

[0063] The upsampling module applies the attention enhancement feature. Two bilinear upsampling operations with a 2x magnification are performed, and after each upsampling, a 3×3 convolution is applied for refinement to obtain the prediction head;

[0064] The prediction head is subjected to pre-channel compression (Channel Reduction) to obtain features. ;

[0065] The features are processed using a 1×1 convolution. The number of channels is compressed to 1, resulting in a two-dimensional tensor. Then, for two-dimensional tensors... use Activation function activation yields crack segmentation prediction result map, which can be represented as follows:

[0066]

[0067]

[0068] in, This is a crack segmentation prediction result image. The value of each pixel in the image represents the probability that the model identifies that pixel as a crack.

[0069] Furthermore, during the training of the crack segmentation network, a loss function based on binary cross-entropy and Dice loss is adopted:

[0070]

[0071] in, and These are hyperparameters used to balance the weights of the binary cross-entropy loss and the Dice loss; For binary cross-entropy loss, The formulas for Dice loss are as follows:

[0072]

[0073]

[0074] in, Total number of pixels; For the first The actual label of each pixel, 0 or 1; For the first The predicted probability of each pixel, i.e. The corresponding value in; This is the smoothing constant.

[0075] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, the electronic device performs the micro-slit precision segmentation method integrating feature fusion and convolutional attention.

[0076] A storage medium comprising a computer program that, when run on an electronic device, causes the electronic device to perform the micro-crack precision segmentation method integrating feature fusion and convolutional attention.

[0077] The beneficial effects of this invention are as follows: The crack segmentation network proposed in this invention adopts an encoding and decoding architecture of Convolutional Block Attention (CBAM) and Feature Fusion (FFM). Through the synergistic effect of CBAM and FFM modules, a deep integration of local details and global semantics is achieved, resulting in a significant technological breakthrough in the fine segmentation of microcracks on the surface of concrete structures (such as roads, bridges, and tunnels). In the encoding stage, the CBAM module adaptively suppresses background noise and enhances the salient features of cracks through a dual attention mechanism of channels and space. This mechanism improves the model's sensitivity to microcracks with almost no increase in computational overhead, especially in low-contrast or texture interference scenarios, reducing the false detection rate by about 0.5%. In the decoding stage, the FFM module achieves the synergy of low-level details (such as edge textures) and high-level semantics (such as the overall shape of cracks) through cross-layer fusion, effectively bridging the semantic gap, avoiding the loss of details caused by traditional convolution stacking, and ensuring the continuity and complete topological structure of narrow cracks (width < 0.2 mm). This invention has validated its performance on public datasets, achieving an optimal balance in accuracy, recall, and edge consistency. It also exhibits strong generalization ability and high sensitivity in detecting microcracks in complex scenarios, providing a highly reliable and deployable intelligent crack detection tool for the health monitoring of large infrastructure such as bridges and tunnels. Attached Figure Description

[0078] Figure 1 This is a model framework diagram of the micro-crack precise segmentation method integrating feature fusion and convolutional attention in an embodiment of the present invention.

[0079] Figure 2 This is a schematic diagram of the CBAM module in an embodiment of the present invention: it shows the cascading of channel attention and spatial attention to suppress noise and highlight key areas of the crack, thereby achieving adaptive feature enhancement.

[0080] Figure 3 The following is a schematic diagram of the FFM module in an embodiment of the present invention: Low-level and high-level features are aligned, weighted, and interactively fused to output an enhanced joint representation, ensuring the consistency between crack details and the global representation.

[0081] Figure 4 The following is a comparison of ablation experiments in this embodiment of the invention: showing the comparison of segmentation results of different module combinations and backbone networks, verifying that CBAM+FFM synergy brings the best performance. Detailed Implementation

[0082] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0083] Crack segmentation is a core task in structural health monitoring. Its goal is to automatically and accurately extract and characterize crack regions in images using computer vision techniques, thereby providing crucial information for service status assessment and maintenance decisions for large concrete structures such as roads, bridges, and tunnel walls. Addressing the inherent limitations of traditional convolutional networks in long-range dependency modeling, preserving fine crack information, and suppressing complex backgrounds, this invention provides a method for accurate micro-crack segmentation that integrates feature fusion and convolutional attention. This method combines a Feature Fusion Module (FFM) and a Convolutional Block Attention Module (CBI). This paper describes a novel encoding and decoding architecture (CBAM) consisting of two main stages: an encoder and a decoder. The encoding stage first extracts multi-level primary features using a ResNet50 backbone network pre-trained on ImageNet. Then, CBAM is embedded within this backbone, and adaptive recalibration of the feature maps is performed along the channel and spatial dimensions using a cascaded approach of channel attention and spatial attention submodules. This enhances sensitivity to salient channels and regions of cracks while simultaneously suppressing irrelevant background interference. The decoding stage uses FFM as its core. After aligning cross-layer feature dimensions through 1×1 convolutions, a channel attention mechanism is introduced to further refine the semantic expression. Then, a correlation enhancement strategy is used to achieve deep interaction and complementary fusion of high- and low-level features. Finally, spatial resolution is gradually restored through progressive upsampling and convolutional restoration operations, outputting a refined crack segmentation result. Figure 1 As shown, the method specifically includes the following steps:

[0084] S1. Obtain crack images and perform preprocessing to construct a crack image dataset.

[0085] This embodiment utilizes the DeepCrack dataset, proposed by Liu et al. in 2019 for crack segmentation and is one of the most representative and widely cited public benchmarks to date. The dataset covers typical concrete structures such as urban asphalt pavements, high-speed railway bridge beams, tunnel linings, and building facades. Imaging conditions span various scenarios including natural lighting, indoor artificial lighting, backlighting, and shadow interference. Crack types range from minute hairline cracks less than 0.2 mm wide to main cracks penetrating the entire field of view, resulting in rich scale differences and topological morphology. A total of 537 high-resolution RGB images were released, with original sizes ranging from 1600×1200 to 2000×1500 pixels. All images were captured vertically using a professional DSLR camera to ensure clear texture details. Each image is equipped with a manually annotated pixel-by-pixel binary mask, with crack areas labeled as 255 and the background as 0. Annotation consistency was cross-validated by three experienced annotators to ensure edge errors are controlled within 2 pixels. In addition, the dataset is pre-divided into training, validation and test sets in an 8:1:1 ratio to avoid evaluation bias caused by different scenario distributions.

[0086] S2. Construct and train the crack segmentation network.

[0087] The experiments were conducted on a workstation equipped with an Intel® Xeon® Gold 6248R CPU (20 cores, 2.50 GHz), 64 GB DDR4 memory, and an NVIDIA RTX 4090 24 GB GPU. The operating system was Ubuntu 22.04LTS, CUDA version 11.8, cuDNN 8.7, and the PyTorch 2.0 framework. The network model was trained using the SGD optimizer with a momentum of 0.9 and a weight decay of 0.0001. The initial learning rate was set to 0.01 and smoothly decayed to 0.00001 over 200 epochs using cosine annealing. The batch size was set to 8 (4 randomly cropped 512×512 GPUs per GPU), and the gradient clipping norm was 5. The loss function was a combination of weighted binary cross-entropy and Dice loss. During the inference phase, the original resolution is maintained, and a horizontally flipped TTA is used with the average output. Evaluation metrics include precision, recall, harmonic mean F1-score, and intersection-union ratio (mIoU).

[0088] To comprehensively and objectively verify the superiority of the proposed crack segmentation network model, a systematic comparative experiment was conducted on the DeepCrack public benchmark against eleven of the most representative crack segmentation algorithms. The reference models included classic architectures from 2018 to 2020 (UNet-ResNet34, DeepLabv3+, CENet, DeepCrack), as well as the latest Transformer and CNN hybrid paradigms proposed from 2021 to 2024 (TransFuse, UTNet, FAT-Net, DscNet, DTrC-Net, DECS-Net), and variants incorporating attention mechanisms (Attn-UNet), ensuring broad representativeness of the comparison results in terms of time period, structure, and methodology. All experiments were reproduced under the same hardware environment and unified hyperparameter configuration to eliminate interference from implementation differences.

[0089] Table 1 Comparative Experimental Results

[0090]

[0091] As shown in the quantitative results in Table 1, the method of this invention achieves the best performance in the four core indicators of Precision, Recall, F1-score, and mIoU, reaching 99.45%, 99.38%, 99.41%, and 80.64%, respectively. Compared with the second-best DECS-Net (mIoU 75.23%), the mIoU is improved by 5.41 percentage points; compared with TransFuse (mIoU 60.76%), a representative Transformer structure, the improvement is 19.88 percentage points; and compared with the classic DeepLabv3+ (mIoU 68.18%), the advantage further expands to 12.46 percentage points. This significant gain fully demonstrates that the present invention outperforms existing mainstream solutions in terms of overall coverage and edge consistency in the crack region.

[0092] Further analysis of the balance between precision and recall reveals that most comparative algorithms often experience a significant decrease in precision when achieving high recall (e.g., DECS-Net, despite achieving a recall of 92.70%, only has a precision of 82.87%, resulting in an F1 score limited to 87.51%). In contrast, the method of this invention maintains a recall close to 99% while retaining a precision as high as 99.45%, achieving dual suppression of errors and missed detections. This advantage directly translates into maintaining the integrity of the crack topology: visualization results show that this invention can effectively avoid breakage and artifacts in narrow, weak cracks and complex texture backgrounds, significantly reducing missed pixels and fragmented false detection areas.

[0093] In summary, the experimental results quantitatively and qualitatively validate the sophistication of the network model presented in this invention. Its performance improvement is primarily attributed to the precise enhancement of salient crack features by the CBAM attention module and the synergistic utilization of local details and global semantics by the FFM multi-scale feature fusion strategy. These two factors work together to enable the model to robustly capture the fine structure of slender cracks even in complex scenarios, achieving an optimal balance between accuracy, recall, and region consistency, further solidifying its leading position in crack segmentation tasks.

[0094] To systematically analyze the contribution of the convolutional block attention module (CBAM) and the feature fusion module (FFM) to the overall network performance, this invention designed and executed multiple controlled ablation experiments on the DeepCrack benchmark, and the results are summarized in Table 2.

[0095] Table 2 Ablation Experiment Results

[0096]

[0097] First, the "Base" model with ResNet50 as its backbone has demonstrated a highly competitive baseline performance without introducing any additional modules: Precision of 99.42%, Recall of 99.39%, F1-score of 99.40%, and mIoU of 80.23%, which fully demonstrates that the pre-trained convolutional backbone itself has strong representation capabilities.

[0098] Building upon this, embedding CBAM (denoted as Base_cb) separately slightly increased Precision and Recall to 99.42% and 99.41%, respectively, and mIoU to 80.43%. Although the improvement was only 0.20 percentage points, it still showed statistical significance in the high baseline range above 80%. Visualization analysis showed that the channel-spatial dual attention mechanism, while suppressing complex texture backgrounds, enhanced the high-frequency response to narrow crack edges, thereby effectively compressing false-detection pixels.

[0099] Furthermore, when only FFM (1) is enabled, denoted as Base_ffm (1), the mIoU reaches 80.55%, the highest among all single-module settings. This result confirms that the cross-layer feature fusion strategy can bridge the semantic gap between low-level details and high-level semantics, enabling the model to obtain a more complete perception of crack regions while maintaining edge refinement. It is worth noting that, due to the different fusion weight initialization used in FFM (2), the mIoU drops back to 80.17%, suggesting that the design of the fusion path needs to be optimized in conjunction with parameter initialization.

[0100] When CBAM and FFM are co-embedded, i.e., the Base_cb_ffm-resnet50 model, all four metrics are optimal: Precision 99.45%, Recall 99.38%, F1-score 99.41%, and mIoU 80.64%. Compared to Base, mIoU is improved by an additional 0.41 percentage points. Although the absolute increase seems limited, in the high plateau region of 99% Precision and Recall, this gain reflects the complementary effect between attention selection and multi-scale fusion, significantly enhancing the topological continuity of crack edges and reducing fragmented false detections.

[0101] The backbone network sensitivity experiment further validated the above conclusions: when ResNet18 was used as the backbone, the mIoU dropped to 78.95%, indicating that insufficient capacity limited the potential of the attention and fusion modules; ResNet34 achieved 80.02%, which is quite close to the result of ResNet50; however, further deepening to ResNet101 led to overfitting due to the surge in parameters, and the mIoU dropped to 75.32%. Overall, ResNet50 achieved the best trade-off between representational ability and training stability, and also provided the most suitable backbone support for the synergistic gains of CBAM and FFM. Specific comparison results are shown below. Figure 4 As shown, the segmentation results of Base_cb_ffm-ResNet50 are closest to the actual labeled data.

[0102] In summary, the lightweight codec network integrating CBAM and FFM proposed in this invention, through systematic experiments on the DeepCrack dataset, verifies that the network still achieves state-of-the-art performance with an F1 score of 99.41% and a mIoU of 80.64% even under complex lighting and texture interference. Ablation results further demonstrate that the spatial-channel dual attention of the CBAM module effectively suppresses background noise, while the cross-layer fusion mechanism of the FFM module significantly improves the topological integrity of narrow cracks. Comprehensive comparison and visualization analysis both show that this invention achieves the best balance in accuracy, recall, and edge consistency, providing a highly reliable and deployable intelligent crack detection tool for health monitoring of large infrastructure such as bridges and tunnels.

[0103] Finally, it should be noted that the above embodiments are intended to illustrate the technical solutions of the present invention and do not constitute any limitation on the present invention. Those skilled in the art should fully understand that modifications to the technical solutions described in the foregoing embodiments or equivalent substitutions for any part or all of the technical features are entirely feasible. Such modifications or substitutions, as long as they do not depart from the scope of protection defined by the claims of the present invention, should be considered reasonable extensions of the present invention.

Claims

1. A precise micro-crack segmentation method integrating feature fusion and convolutional attention, characterized in that, include: S1. Obtain crack images and preprocess them to construct a crack image dataset; S2. Construct a crack segmentation network, wherein the crack segmentation network is an encoder-decoder structure; The encoder includes a deep residual network and a convolutional block attention module; the deep residual network serves as the backbone network to extract multi-scale features of the input image, and the convolutional block attention module performs saliency enhancement on the multi-scale features to obtain enhanced features for each scale. The enhanced features at each scale include high-level semantic features and low-level detail features; The decoder includes a feature fusion module and a decoding head; the feature fusion module performs cross-layer fusion and refinement of the low-level detail features and high-level semantic features output by the encoder to obtain fused features that have both global semantics and local details; the decoding head processes the fused features sequentially through convolution and progressive upsampling to gradually restore spatial resolution and enhance edge detail expression to obtain a crack segmentation prediction result map. S3. The crack segmentation network is trained using the crack image dataset constructed in S1. S4. Input the detected image to be segmented into the trained crack segmentation network. The encoder extracts multi-scale features and gradually fuses and restores the resolution in the decoder. Finally, the decoder outputs the segmentation result of the crack region.

2. The micro-crack precise segmentation method integrating feature fusion and convolutional attention as described in claim 1, characterized in that, The specific process for constructing the crack image dataset includes: Acquire crack images, ensuring they include cracks of different scales and shapes; The crack regions in the acquired crack images are annotated at the pixel level to generate a binary mask corresponding to the original crack image. The labeled crack images are preprocessed, including image normalization, cropping, flipping, rotating, scaling and brightness adjustment, to enhance data diversity and the generalization ability of the model, forming a crack image dataset. The crack image dataset was divided into training, validation, and test sets.

3. The micro-crack precise segmentation method integrating feature fusion and convolutional attention as described in claim 1, characterized in that, The deep residual network includes a ResNet50 model, which takes the preprocessed crack image as input and extracts multi-scale feature maps from shallow to deep by performing layer-by-layer convolution and downsampling operations on the crack image. It includes shallow features Retain detailed information, high-level features It contains global semantic information.

4. The micro-crack precise segmentation method integrating feature fusion and convolutional attention as described in claim 3, characterized in that, The convolutional block attention module consists of cascaded channel attention submodules and spatial attention submodules; The channel attention submodule extracts multi-scale feature maps from the deep residual network. As input, global average pooling and global max pooling are used to compress information in the spatial dimension, respectively, to obtain two channel-level descriptors. : in, , These are the height and width of the input image, respectively. Number of channels; , These represent the row and column coordinates of the feature map in the spatial dimension, respectively. Will , A multilayer perceptron (MLP) with shared parameters is fed into the machine to learn nonlinear interaction relationships; the MLP is then subjected to dimensionality reduction scaling. To control the bottleneck structure, the multilayer perceptron consists of a first fully connected layer, a ReLU activation function, and a second fully connected layer connected sequentially. Its output is then processed... After activation by the activation function, the channel attention map is obtained. The process is represented as: in, for Activation function; The enhanced features are then obtained by multiplying each channel sequentially. : The spatial attention submodule enhances the features of the channel. Perform average pooling and max pooling along the channel dimension to generate two two-dimensional spatial descriptors. ; in, Indicates the number of channels. ; Will , After being spliced ​​along the channel axis, via a Convolutional layers capture local cross-channel interactions and are then processed... Spatial attention map is obtained by activation function. : in, This represents a convolutional layer with a kernel size of [size missing]. ; Finally, the spatial attention map Compared with enhanced features Element-wise multiplication yields the refined feature tensor. ; The shallow features , The feature tensors obtained after feature enhancement by the convolutional block attention module are respectively , Constitutes low-level detailed features , These represent the height, width, and number of channels of the low-level detail feature tensor, respectively; the high-level feature... , The feature tensors obtained after feature enhancement by the convolutional block attention module are respectively , Constitutes high-level semantic features , These represent the height, width, and number of channels of the high-level semantic feature tensor, respectively.

5. The micro-crack precise segmentation method integrating feature fusion and convolutional attention according to claim 4, characterized in that, The feature fusion module includes a channel alignment module, a channel attention enhancement module, a correlation enhancement module, and a deep convolution refinement module; The channel alignment module will incorporate high-level semantic features. and low-level details The number of channels is adjusted to a uniform dimension C using 1×1 convolution, and bilinear interpolation is used to apply the high-level semantic features. Upsampling is performed to improve its spatial resolution and low-level detail features. Consistency is achieved, resulting in high-level semantic features after dimension alignment. and low-level detail features ; The channel attention enhancement module respectively... and Channel enhancement is performed using the following formula: in, This represents the channel weight vector for high-level semantic features. ; This represents the channel weight vector for low-level detail features. ; This is a global average pooling operation; , These are the enhanced high-level semantic features and low-level detail features, respectively. The correlation enhancement module is used to promote deep interaction and complementary fusion between high-level semantic features and low-level detailed features; firstly, the features are... and Expanding in space: Calculate the interaction matrix: in, The interaction matrix represents the similarity between high-level semantic features and low-level detail features in the channel dimension. Then the interaction matrix After activation processing, it is used as the fusion weight. : For fusion weights Complementary weighted fusion is performed to obtain the fusion features after cross-domain interaction. : in, To convert the matrix shape into a feature map shape; for identity matrix; The deep convolutional refining module refines the fused features. After a series of lightweight convolutional refinements, the final fused features are obtained, which preserve both low-level edge information and high-level semantic global information. This process can be represented as follows: in, For the final fusion feature; For depthwise separable convolution; For batch normalization; It is a non-linear activation.

6. The micro-crack precise segmentation method integrating feature fusion and convolutional attention as described in claim 5, characterized in that, The decoding head includes a convolutional block attention module and an upsampling module; The convolutional block attention module in the decoder has the same architecture as the convolutional block attention module in the encoder to fuse features. As input, the data is processed sequentially by a cascaded channel attention submodule and a spatial attention submodule to obtain attention-enhanced features. The process is represented as: The upsampling module applies the attention enhancement feature. Two bilinear upsampling operations with a 2x magnification are performed, and after each upsampling, a 3×3 convolution is applied for refinement to obtain the prediction head; Prechannel compression is performed on the prediction head to obtain features. ; The features are processed using a 1×1 convolution. The number of channels is compressed to 1, resulting in a two-dimensional tensor. Then, for two-dimensional tensors... use Activation function activation yields crack segmentation prediction result map, which can be represented as follows: in, This is a crack segmentation prediction result image. The value of each pixel in the image represents the probability that the model identifies that pixel as a crack.

7. The micro-crack precise segmentation method integrating feature fusion and convolutional attention according to claim 6, characterized in that, During the training of the crack segmentation network, a loss function based on binary cross-entropy and Dice loss is used: in, and These are hyperparameters used to balance the weights of the binary cross-entropy loss and the Dice loss; For binary cross-entropy loss, The formulas for Dice loss are as follows: in, Total number of pixels; For the first The actual label of each pixel, 0 or 1; For the first The predicted probability of each pixel, i.e. The corresponding value in; This is the smoothing constant.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; characterized in that, When the processor executes the computer program, it causes the electronic device to perform the micro-crack precision segmentation method integrating feature fusion and convolutional attention as described in any one of claims 1-7.

9. A storage medium comprising a computer program, characterized in that, When the computer program is run on an electronic device, the electronic device performs the micro-crack precision segmentation method integrating feature fusion and convolutional attention as described in any one of claims 1-7.

Citation Information

Patent Citations

  • High-precision crack detection method

    CN111222580A

  • Crack image segmentation method based on high-resolution network

    CN117455933A

  • Unmanned aerial vehicle image pavement crack segmentation method fusing multi-scale feature extraction and attention mechanism

    CN118247690A

  • Bridge crack detection method based on multi-scale feature fusion and multi-layer attention

    CN118941542A

  • Concrete crack detection method based on mixed Mama attention segmentation model

    CN120543492A

Cited By

  • Infrared and visible light image fusion method and device based on lightweight model

    CN121329792A

  • Multi-scale mutual feedback attention crack segmentation method and device and electronic equipment

    CN121544643A

  • High-precision crack detection method based on multi-scale feature fusion and direction distribution

    CN121600319A

  • Lightweight target detection edge deployment method based on high-frequency detail enhancement

    CN121746687A

  • Lightweight target detection edge deployment method based on high-frequency detail enhancement

    CN121746687B