Pavement crack edge perception segmentation method and system based on multi-modal cross fusion

CN122453844BActive Publication Date: 2026-09-15CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610920988.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-15
Estimated Expiration
2046-06-25

AI Technical Summary

Technical Problem

[0004]本发明的主要目的是提出一种基于多模态交叉融合的路面裂缝边缘感知分割方法及系统,旨在解决现有多模态路面裂缝分割中模态特征对齐难、边缘细节丢失以及复杂场景鲁棒性差的技术问题

Benefits of technology

[0041] The above-described technical solution of the present invention provides a road surface crack edge perception segmentation method based on multimodal cross-fusion, comprising the following steps: constructing a dual-encoder architecture, inputting the registered RGB image and at least one auxiliary modal image into the corresponding encoders respectively, each encoder employing a hierarchical downsampling structure, extracting multi-scale features through stacked general feature enhancement modules, and outputting RGB features and auxiliary modal features at different levels; at the end of each encoding stage, inputting the RGB features and all auxiliary modal features into a cross-attention fusion module, performing cross-modal information interaction and adaptive weighted fusion to generate the fusion features at the current level; fusing and reconstructing the multi-level fusion features from the encoders to obtain a high-resolution feature map, and outputting the final road surface crack segmentation result. This invention solves the technical problems of difficult modal feature alignment, loss of edge details, and poor robustness in complex scenes in existing multimodal road surface crack segmentation methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122453844B_ABST
    Figure CN122453844B_ABST
Patent Text Reader

Abstract

The application discloses a pavement crack edge perception segmentation method and system based on multi-modal cross fusion, wherein the method comprises the following steps: constructing a double-encoder architecture, inputting a registered RGB image and at least one auxiliary modal image into corresponding encoders respectively, each encoder adopts a hierarchical down-sampling structure, multi-scale features are extracted through a stacked general-purpose feature enhancement module, and different levels of RGB features and auxiliary modal features are output; at the end of each encoding stage, the RGB features and all auxiliary modal features are input into a cross-attention fusion module to perform cross-modal information interaction and adaptive weighted fusion to generate fusion features of the current level; and the multi-level fusion features are fused and reconstructed to obtain a high-resolution feature map, and finally, a pavement crack segmentation result is output. The application solves the technical problems of difficulty in aligning modal features, loss of edge details and poor robustness in complex scenes in the existing multi-modal pavement crack segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a method and system for perceiving and segmenting road surface crack edges based on multimodal cross-fusion. Background Technology

[0002] With the acceleration of urbanization, the demand for intelligent maintenance of road infrastructure is becoming increasingly urgent. Crack detection, as a core component of preventive maintenance, plays a crucial role in ensuring the safety and sustainability of public transportation. Traditional manual detection methods suffer from efficiency bottlenecks and subjective variability, failing to meet the needs of large-scale management of modern road networks. Semantic segmentation based on deep learning, with its powerful feature extraction and pixel-level classification capabilities, has become the mainstream method for pavement crack segmentation and has achieved excellent results. However, most existing deep learning methods rely solely on single-modal RGB images, making their performance highly dependent on lighting conditions and image quality. They are easily affected by shadows, noise, complex background textures, and crack-like interference, and their performance significantly degrades in complex scenarios such as rain, fog, and low illumination.

[0003] With the advancement of sensor technology, an increasing number of modular sensors are being applied to intelligent transportation systems. Different types of sensors can provide rich complementary information for RGB images to correct detection errors and adapt to various detection environments. For example, the RGB-depth mode provides spatial distance information of road surface cracks; the RGB-thermal mode captures thermal anomalies on and beneath the road surface; and the RGB-polarization mode provides radiation characteristics that help enhance the distinction between targets and backgrounds in complex scenes, especially for fine targets such as road surface cracks. Therefore, utilizing the complementary information of multimodal images overcomes the limitations of single-modal RGB data in complex scenes, providing a feasible solution for all-weather, highly robust road surface crack detection in practical engineering environments. Although multimodal segmentation is highly anticipated, the physical characteristics of RGB, infrared, and polarization data differ significantly, making accurate data registration and synchronization difficult to achieve in software. Furthermore, road surface cracks exhibit large differences across modes, making it difficult to define unified discrimination features. These factors combined result in low accuracy and reliability of existing multimodal road surface crack segmentation methods. Therefore, to address these technical problems, it is urgent to propose a road surface crack edge perception segmentation method and system based on multimodal cross-fusion. Summary of the Invention

[0004] The main objective of this invention is to propose a method and system for perceptual segmentation of road crack edges based on multimodal cross-fusion, aiming to solve the technical problems of difficult modal feature alignment, loss of edge details, and poor robustness in complex scenarios in existing multimodal road crack segmentation.

[0005] To achieve the above objectives, this invention provides a method for perceiving and segmenting road surface crack edges based on multimodal cross-fusion, wherein the method includes the following steps:

[0006] S1. Construct a dual encoder architecture, input the registered RGB image and at least one auxiliary modal image into the corresponding encoder respectively. Each encoder adopts a hierarchical downsampling structure, extracts multi-scale features through stacked general feature enhancement modules, and outputs RGB features and auxiliary modal features at different levels.

[0007] S2. At the end of each encoding stage, the RGB features and all auxiliary modal features are input into the cross-attention fusion module to perform cross-modal information interaction and adaptive weighted fusion to generate the fusion features of the current level.

[0008] S3. The multi-level fusion features from the encoder are fused and reconstructed to obtain a high-resolution feature map, and the final road surface crack segmentation result is output.

[0009] In one preferred embodiment, the general feature enhancement module consists of a frequency domain enhancement Mamba module, a shifted edge sensing module, and an adaptive residual connection.

[0010] In one preferred embodiment, the frequency domain enhancement Mamba module adopts a three-branch parallel structure;

[0011] The first branch is the selective state-space modeling branch, which extracts global features through selective state-space models in four directions;

[0012] The second branch is the Fourier channel attention branch, which transforms the features from the spatial domain to the frequency domain through Fourier transform to perform global channel attention modeling.

[0013] The third branch is a strip-gated convolutional branch, which sequentially passes through a 5×5 depth convolutional layer, a 1×7 depth separable convolutional layer, a 7×1 depth separable convolutional layer, a gating layer, and a 1×1 convolutional layer to extract horizontal and vertical contextual features.

[0014] The outputs of the first, second, and third branches are gated and fused, and then adaptively connected to the input features using residual connections.

[0015] One preferred embodiment is the Fourier channel attention branch, which is specifically:

[0016]

[0017]

[0018]

[0019]

[0020]

[0021] in, respectively for input features Perform Fourier transform and inverse transform. The feature map is obtained after Fourier transform. and These are the real and imaginary features in the frequency domain, multiplied by the learnable frequency domain weights, respectively. and These are the learnable frequency domain weights. and These are the true and imaginary values ​​taken from the feature map obtained after Fourier transform. The final generated channel attention weight vector has dimensions of , It is the Sigmoid activation function. The output characteristics of the Fourier channel attention branch, For element-wise multiplication, To enhance the original input features of the Mamba module in the frequency domain, It is a 1×1 convolutional layer.

[0022] In one preferred embodiment, the shift edge sensing module divides the input features into two groups along the channel;

[0023] The first group is multiplied and fused with the input features by shifting windows in the height and width directions, and then processed by a 3×3 depth convolution.

[0024] The second group is processed sequentially through 1×1 convolution, 3×3 depthwise convolution, and Fourier channel attention;

[0025] The outputs of the first and second groups are concatenated and then fused using a lightweight convolution before being output.

[0026] In one preferred embodiment, when the number of auxiliary modal features is greater than 1 in step S2, the most discriminative auxiliary modal features are selected through a learnable scoring network before entering the cross-attention fusion module. Specifically, each auxiliary modal feature is generated into a single-channel sub-map through lightweight convolution, multiplied with itself and added together, and then the most discriminative auxiliary modal feature is output through max pooling.

[0027] One preferred embodiment, step S2, specifically comprises:

[0028] Layer normalization is performed on RGB features and auxiliary modal features respectively;

[0029] Directional context is extracted using multi-scale band convolution kernels;

[0030] Using auxiliary modality features as the query and RGB features as the key / value pair, calculate the first cross-attention and output the first cross-attention result; simultaneously, using RGB features as the query and auxiliary modality features as the key / value pair, calculate the second cross-attention and output the second cross-attention result.

[0031] The first and second cross-attention results are concatenated along the channels, and after channel averaging and max pooling, a two-channel spatial weight map is generated through 7×7 convolution. The first and second cross-attention results are weighted and summed, and after 1×1 convolution, they are added to the original input features to output the fused features.

[0032] One preferred embodiment, step S3, specifically includes:

[0033] The decoder adopts a top-down progressive upsampling strategy, upsampling the highest-level fusion features step by step, and then performing cross-scale stitching and weighted fusion with the lower-level fusion features in turn. After each level of fusion, features are purified through convolution and activation operations, and finally the segmentation result with the same size as the input image is output.

[0034] One preferred embodiment of the road surface crack edge sensing and segmentation method based on multimodal cross-fusion further includes:

[0035] Model training involves end-to-end training of the dual encoder, cross-attention fusion module, and decoder using a labeled multimodal pavement crack sample set until the model converges.

[0036] The model training employs a hybrid loss function, which is a weighted sum of binary cross-entropy loss and Dice loss.

[0037] A system including the aforementioned road surface crack edge sensing and segmentation method based on multimodal cross-fusion, characterized in that it comprises a dual encoder module, a cross-modal fusion module, and a decoding and reconstruction module connected in sequence:

[0038] The dual encoder module is used to perform hierarchical downsampling on the registered RGB image and at least one auxiliary modality image respectively, and extract multi-scale features through the general feature enhancement module;

[0039] The cross-modal fusion module is used to perform bidirectional cross-attention fusion on RGB features and auxiliary modal features at the end of each encoding stage, generating fused features that are fed back to the encoding and decoding stages;

[0040] The decoding and reconstruction module is used to progressively upsample and stitch together multi-level fused features to output pixel-level segmentation results of road surface cracks.

[0041] The above-described technical solution of the present invention provides a road surface crack edge perception segmentation method based on multimodal cross-fusion, comprising the following steps: constructing a dual-encoder architecture, inputting the registered RGB image and at least one auxiliary modal image into the corresponding encoders respectively, each encoder employing a hierarchical downsampling structure, extracting multi-scale features through stacked general feature enhancement modules, and outputting RGB features and auxiliary modal features at different levels; at the end of each encoding stage, inputting the RGB features and all auxiliary modal features into a cross-attention fusion module, performing cross-modal information interaction and adaptive weighted fusion to generate the fusion features at the current level; fusing and reconstructing the multi-level fusion features from the encoders to obtain a high-resolution feature map, and outputting the final road surface crack segmentation result. This invention solves the technical problems of difficult modal feature alignment, loss of edge details, and poor robustness in complex scenes in existing multimodal road surface crack segmentation methods.

[0042] In this invention, a multimodal cross-fusion-based pavement crack edge perception and segmentation method is employed. The encoding end utilizes a frequency-domain enhanced Mamba module, significantly improving the extraction capability of global features. Simultaneously, the proposed shifted edge perception module effectively enhances the model's target localization accuracy and refines the segmentation effect of crack edges. Furthermore, a cross-attention fusion module is used to fuse features among multimodal data, fully learning the interdependencies between different modalities through multi-scale cross-attention, promoting efficient interaction of complementary features. This invention demonstrates excellent performance in automated multimodal crack segmentation tasks, significantly improving detection accuracy and segmentation quality. It can accurately identify pavement cracks, providing more reliable information support for road maintenance and traffic monitoring, and possesses significant practical application value. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0044] Figure 1 This is a first schematic diagram of a road surface crack edge sensing and segmentation method based on multimodal cross-fusion according to an embodiment of the present invention;

[0045] Figure 2 This is a second schematic diagram of the pavement crack edge sensing and segmentation method based on multimodal cross-fusion according to an embodiment of the present invention;

[0046] Figure 3This is a schematic diagram of the general feature enhancement module according to an embodiment of the present invention;

[0047] Figure 4 This is a schematic diagram of a selective state-space modeling branch according to an embodiment of the present invention;

[0048] Figure 5 This is a third schematic diagram of the pavement crack edge sensing and segmentation method based on multimodal cross-fusion according to an embodiment of the present invention.

[0049] The realization of the objective, functional characteristics and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] It should be noted that all directional indicators (such as up, down, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.

[0052] Furthermore, in this invention, descriptions involving "first," "second," etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.

[0053] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0054] See Figures 1-5 According to one aspect of the present invention, the present invention provides a method for perceiving and segmenting road surface crack edges based on multimodal cross-fusion, wherein the method for perceiving and segmenting road surface crack edges based on multimodal cross-fusion includes the following steps:

[0055] S1. Construct a dual encoder architecture, input the registered RGB image and at least one auxiliary modal image into the corresponding encoder respectively. Each encoder adopts a hierarchical downsampling structure, extracts multi-scale features through stacked general feature enhancement modules, and outputs RGB features and auxiliary modal features at different levels.

[0056] S2. At the end of each encoding stage, the RGB features and all auxiliary modal features are input into the cross-attention fusion module to perform cross-modal information interaction and adaptive weighted fusion to generate the fusion features of the current level.

[0057] S3. The multi-level fusion features from the encoder are fused and reconstructed to obtain a high-resolution feature map, and the final road surface crack segmentation result is output.

[0058] Specifically, in this embodiment, step S1 is as follows:

[0059] Registered RGB image and Auxiliary modal images Encoders with the same structure but different weights are input separately to obtain rich road surface crack features using multimodal data input. Each encoder adopts a hierarchical downsampling structure with four stages. Each stage consists of downsampling and multiple basic processing blocks and general feature enhancement modules stacked together. The general feature enhancement module consists of a frequency domain enhancement Mamba module, a shifted edge sensing module, and an adaptive residual connection.

[0060] Specifically, in this embodiment, in order to fully extract crack information from different modalities, the present invention constructs the general feature enhancement module, see [link to relevant documentation]. Figure 3 ;

[0061] First, the input features After layer normalization, we get , To enhance the original input features of the Mamba module in the frequency domain;

[0062] Next, the frequency domain enhancement Mamba module, consisting of three parallel branches, is used to enhance the two-dimensional spatial global perception capability of frequency domain features and reduce noise interference introduced by inter-modal differences.

[0063] The frequency domain enhancement Mamba module adopts a three-branch parallel structure;

[0064] The first branch is the selective state-space modeling branch, which extracts global features through selective state-space models in four directions; see [link / reference]. Figure 4 The original input features are flattened into a sequence, and global features are output along four directions using a selective state-space model. This module improves computational efficiency by introducing a scanning mechanism and variable parameter modeling, while maintaining the global receptive field, thus achieving a good balance between modeling accuracy and inference speed.

[0065] The second branch is the Fourier channel attention branch, which transforms the features from the spatial domain to the frequency domain through Fourier transform and performs global channel attention modeling; it then performs global average pooling on the original input features to obtain the input features z. Next, the features are transformed to the frequency domain using Fast Fourier Transform, and learnable frequency domain weights are introduced to enhance the channel feature representation; specifically:

[0066]

[0067]

[0068]

[0069]

[0070] in, respectively for input features Perform Fourier transform and inverse transform. The feature map is obtained after Fourier transform. and These are the real and imaginary features in the frequency domain, multiplied by the learnable frequency domain weights, respectively. and These are the learnable frequency domain weights. and These are the true and imaginary values ​​taken from the feature map obtained after Fourier transform. The final generated channel attention weight vector has dimensions of , It is the Sigmoid activation function. The output characteristics of the Fourier channel attention branch, For element-wise multiplication, To enhance the original input features of the Mamba module in the frequency domain, It is a 1×1 convolutional layer;

[0071] The third branch is a strip-gated convolutional branch, which sequentially passes through a 5×5 depth convolutional layer, a 1×7 depth separable convolutional layer, a 7×1 depth separable convolutional layer, a gating layer, and a 1×1 convolutional layer to extract horizontal and vertical contextual features; specifically:

[0072] A 5×5 depthwise convolutional layer with a square kernel is applied to the original input features to extract local contextual features. After the initial depthwise convolution, two separable convolutional layers with large strip kernels (1×7 and 7×1) are used to capture high aspect ratio crack features. Unlike standard convolutions that extract features from square regions each time, large strip convolutions allow the network to focus more on features along the horizontal or vertical axis. The combination of horizontal and vertical large strip convolutions enables the network to collect directional features on two spatial axes, thereby enhancing the spatial representation of slender or narrow structures. Finally, pointwise convolutions further enhance the interaction of cross-channel dimensional features. In this way, the resulting feature map... Each location encodes horizontal and vertical features over a broad spatial region, specifically:

[0073]

[0074]

[0075] in, For size Depth convolution, Indicates size is Pointwise convolution, Original input features The output feature map obtained after passing through three depth-separated convolutional layers in sequence—a 5×5 depth-separated convolutional layer, a 1×7 depth-separated convolutional layer, and a 7×1 depth-separated convolutional layer—contains the extracted multi-directional (horizontal and vertical) contextual information. This is the branch output for the third branch;

[0076] The outputs of the first, second, and third branches are gated and fused, and then adaptively connected to the input features using residual connections.

[0077] Specifically, in this embodiment, to compensate for the shortcomings of the frequency domain enhancement Mamba module in local detail extraction, a shifted edge perception module is constructed. This module significantly improves the model's target localization capability and refines segmentation edges by introducing local enhancement and window shifting operations; specifically:

[0078] The output characteristics obtained from the frequency domain enhancement Mamba module After layer normalization, the channels are divided into two groups. ;

[0079] The first group of features is multiplied and fused with the input features through window shifting in both height and width directions, and then processed by a 3×3 depthwise convolution. To preserve the basic spatial information of the input features, zero-padding is used to prevent information loss. The window size is used to divide the features along the channels, and each group of features is shifted in height or width. The shifted features are then concatenated, and padding is removed through a shrinking operation, effectively reducing the loss of spatial information during local feature extraction. The window shifting operation is as follows:

[0080]

[0081] in, The output features are the first set of features after window shifting, fusion, and convolution processing. It is a 3×3 depth convolutional layer. This is a window shift operation in the width direction. This is a window shift operation in the height direction. This is the first set of input features;

[0082] The second group sequentially undergoes 1×1 convolution, 3×3 depthwise convolution, and Fourier channel attention processing to enhance channel labeling; this operation effectively introduces local features, specifically:

[0083]

[0084] in, The output features of the second set of features are processed by 1×1 convolution, 3×3 depthwise convolution, and Fourier channel attention. For Fourier channel attention modules, It is a 1×1 pointwise convolutional layer;

[0085] The outputs of the first and second groups are concatenated and then fused using a lightweight convolution. It is worth noting that the RGB features and auxiliary modal features output at each encoding stage are saved for subsequent fusion.

[0086] Specifically, in this embodiment, when the number of auxiliary modal features is greater than 1 in step S2, the most discriminative auxiliary modal features are selected through a learnable scoring network before entering the cross-attention fusion module. Specifically, each auxiliary modal feature is generated into a single-channel sub-map through lightweight convolution, multiplied with itself and added together, and then the most discriminative auxiliary modal feature is output through max pooling to reduce redundant calculations.

[0087] Specifically, in this embodiment, in order to improve the complementary representation capability of multimodal features, at the end of each encoding stage, the RGB branch output features are... With auxiliary modal features Input is fed into the cross-attention fusion module, and the output is the fused feature. The fused features are fed back to the next stage input of the RGB branch and the auxiliary branch on the one hand, and saved as a skip connection for use by the decoder on the other hand.

[0088] Step S2 specifically includes:

[0089] Layer normalization is performed on RGB features and auxiliary modal features respectively;

[0090] Directional context is extracted using multi-scale strip convolution kernels; the multi-scale strip convolution kernels include depthwise separable convolutions of 1×7, 1×11, 1×21 and 7×1, 11×1, 21×1.

[0091] Using auxiliary modality features as the query and RGB features as the key / value pair, calculate the first cross-attention and output the first cross-attention result; simultaneously, using RGB features as the query and auxiliary modality features as the key / value pair, calculate the second cross-attention and output the second cross-attention result.

[0092] The first and second cross-attention results are concatenated along the channels, and after channel averaging and max pooling, a two-channel spatial weight map is generated through 7×7 convolution. The first and second cross-attention results are weighted and summed, and after 1×1 convolution, they are added to the original input features to output the fused features.

[0093] Specifically, in this embodiment, scale mapping is used to map RGB features and auxiliary modality features to a unified scale. Specifically, vertical bar convolutions of different scales are used to process each feature, which are then concatenated and mapped using a 1×1 convolution to form matrices Q, K, and L of a unified scale as input for the next stage. RGB features and auxiliary modality features each obtain two sets of matrices (Q1, K1, V1) and (Q2, K2, V2). Then, cross-attention is used to calculate attention by querying key-value pairs of the corresponding counterpart and itself, followed by feature weighting to ultimately achieve feature selection. This operation ensures that the model focuses more on features with similar semantics during feature selection, thereby achieving better semantic alignment. Specifically:

[0094]

[0095]

[0096]

[0097]

[0098] in, These represent RGB image features and auxiliary modal features, respectively. For attention, each head dimension, For the first The output features of each modality at the end of the encoding stage This is the set of sizes for the bar convolution kernels, with values ​​of 7, 11, and 21. 1× Horizontal bar convolution, It is a vertical bar convolution. For the first The fused features obtained by extracting modal features through multi-scale strip convolution For the first The query matrix corresponding to each modality For the first The bond matrix corresponding to each mode For the first The value matrix corresponding to each mode This is the first cross-attention output. This is the output of the second cross-attention;

[0099] The complementary features are combined through spatial adaptive fusion. and The method dynamically learns an optimal fusion weight for each spatial location, rather than simply adding or cascading them. Therefore, spatial adaptive fusion is a key step in achieving high-quality cross-modal fusion within the cross-attention fusion module, directly improving the accuracy and completeness of crack segmentation in variable environments. Specifically:

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107] in, For splicing features, This represents the channel average pooling characteristic. This represents the channel max pooling feature. To concatenate the channel average pooling features and the channel max pooling features along the channel, It is the Sigmoid activation function. This is a two-channel spatial weighting map. , Two-channel spatial weighting map The first channel weight map after splitting, Two-channel spatial weighting map The second channel weight map after splitting, , For weighted fusion features, For the final fusion feature, The original output features of the RGB branch at the end of the encoding stage. The original output features of the auxiliary modal branch at the end of the encoding stage;

[0108] The above output This is the fusion feature of the current encoding stage, which is fed back to the next stage of the two branches and saved as a skip connection. It is worth noting that if there are multiple auxiliary modalities, the most discriminative modal feature is selected through a learnable scoring network before entering the cross-attention fusion module (specifically: each auxiliary modal feature is generated into a single-channel score map through lightweight convolution, multiplied with itself and then added, and then output through max pooling).

[0109] Specifically, in this embodiment, step S3 is as follows:

[0110] The decoder adopts a top-down progressive upsampling strategy, upsampling the highest-level fusion features step by step, and then performing cross-scale stitching and weighted fusion with the lower-level fusion features in turn. After each level of fusion, features are purified through convolution and activation operations, and finally the segmentation result with the same size as the input image is output.

[0111] Specifically, in this embodiment, to fully utilize the multi-level fusion features generated during the encoding process, the decoder gradually integrates deep semantic information with shallow detail information to achieve high-precision crack image reconstruction; the decoder receives the fusion feature list output by the encoder in four stages, denoted as:

[0112] ;

[0113] The number of channels is [C1, C2, C3, C4] = [64, 128, 320, 512], and the spatial size is halved at each level. To fully exploit the multi-level fusion features generated during the encoding stage, the decoder progressively fuses deep semantic information with shallow detail information to achieve high-precision reconstruction of the crack image. The decoder receives the fusion feature sequence output from the four stages of the encoder, denoted as:

[0114] ;

[0115] The corresponding channel numbers are [C1, C2, C3, C4] = [64, 128, 320, 512], and the spatial resolution decreases by half with each level. Based on this, the decoder employs a top-down upsampling and skip connection strategy to extract high-dimensional semantic features. Start gradually upsampling and compare it with the shallow features from the previous stage. Cross-scale stitching and weighted fusion are performed to restore the spatial size while preserving fine-grained information such as crack edges and textures. Each level of decoding completes feature purification through convolution and activation operations, effectively suppressing background noise and artifact interference. The final output is a crack segmentation result with the same size as the input image, clear semantics, and complete details.

[0116] Specifically, in this embodiment, the road surface crack edge sensing and segmentation method based on multimodal cross-fusion further includes:

[0117] Model training involves end-to-end training of the dual encoder, cross-attention fusion module, and decoder using a labeled multimodal pavement crack sample set until the model converges.

[0118] The model training employs a hybrid loss function; the hybrid loss function is a weighted sum of binary cross-entropy loss and Dice loss;

[0119] The choice of loss function directly affects training efficiency and segmentation performance. Road crack segmentation can be represented as a pixel-wise binary classification problem, where each pixel is classified as either a crack or background. A key challenge of this task is the severe class imbalance between crack and background regions. To address this issue, a hybrid loss function is designed, combining binary cross-entropy loss and Dice loss. Specifically, the binary cross-entropy loss focuses on pixel-level binary classification and provides stable gradient supervision, while the Dice loss emphasizes region-level overlap and effectively mitigates the impact of class imbalance. By leveraging the complementary advantages of these two losses, the loss function designed in this invention effectively alleviates the class imbalance problem in crack segmentation.

[0120] The loss function is:

[0121]

[0122]

[0123]

[0124] in, The total number of pixels in the image. It is a pixel The true value, Represents pixels The predicted value; in addition, This represents the smoothing factor set to 1 for numerical stability. and It is a hyperparameter that controls the weights of the two loss components; For binary cross-entropy loss, For Dice's loss, The total loss; in this invention, , This invention does not impose specific limitations, and can be set as needed; this hybrid formula can ensure accurate pixel-level prediction while maintaining the continuity and integrity of the crack structure under class imbalance.

[0125] Specifically, in this embodiment, the registered RGB image and X (X≥1) auxiliary modal images are input into the corresponding encoders. Each encoder adopts a hierarchical downsampling structure and extracts multi-scale features through stacked general feature enhancement modules. The feature map size is halved after each stage, outputting RGB features and auxiliary modal features at different levels. At the end of each encoding stage, the RGB features and all auxiliary modal features are input into the cross-attention fusion module to perform cross-modal information interaction and adaptive weighted fusion, generating the fused features of the current level. These fused features are fed back to the next stage input of the RGB branch and the auxiliary modal branch to continuously enhance subsequent encoding, and are also saved as skip connections for use by the corresponding level of the decoder. The invention fuses and reconstructs the multi-level fusion features from the encoder to obtain a high-resolution feature map, and outputs the final road crack segmentation result. The invention consists of an encoder constructed by a general feature enhancement module and a cross-attention fusion module integrated for multimodal feature fusion, achieving global context feature extraction and the complementarity and fusion of multimodal features. The general feature enhancement module aggregates global context space information through selective state-space modeling and channel Fourier enhancement, and utilizes operations such as window shifting and local convolution to enhance local modeling capabilities, improving robustness to cracks of different scales and weak continuity. The cross-attention fusion module uses multi-scale cross-attention to semantically align the input multimodal feature maps to ensure the consistency and complementarity of the represented images. This operation effectively achieves semantic alignment and feature selection of RGB and auxiliary modal features and promotes their complementarity and fusion, thereby ensuring the accuracy of the segmentation result.

[0126] According to another aspect of the present invention, a road surface crack edge sensing and segmentation system based on multimodal cross-fusion is provided, comprising a dual encoder module, a cross-modal fusion module, and a decoding and reconstruction module connected in sequence:

[0127] The dual encoder module is used to perform hierarchical downsampling on the registered RGB image and at least one auxiliary modality image respectively, and extract multi-scale features through the general feature enhancement module;

[0128] The cross-modal fusion module is used to perform bidirectional cross-attention fusion on RGB features and auxiliary modal features at the end of each encoding stage, generating fused features that are fed back to the encoding and decoding stages;

[0129] The decoding and reconstruction module is used to progressively upsample and stitch together multi-level fused features to output pixel-level segmentation results of road surface cracks.

[0130] To facilitate understanding of the terminology related to this invention, the following explanations are provided:

[0131] CEM is a general feature enhancement module.

[0132] CAF stands for Cross-Attention Fusion Module

[0133] FEM is a frequency domain enhancement Mamba module.

[0134] SEM is a shift edge sensing module.

[0135] FCA is the Fourier channel attention branch.

[0136] SGConv is a strip-gated convolutional branch.

[0137] VSS Block is a branch for selective state-space modeling.

[0138] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. All equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A method for perceptual segmentation of road surface crack edges based on multimodal cross-fusion, characterized in that, Includes the following steps: S1. Construct a dual encoder architecture, inputting the registered RGB image and at least one auxiliary modal image into the corresponding encoders respectively. Each encoder adopts a hierarchical downsampling structure, extracts multi-scale features through stacked general feature enhancement modules, and outputs RGB features and auxiliary modal features at different levels. The general feature enhancement module consists of a frequency domain enhancement Mamba module, a shifted edge sensing module, and an adaptive residual connection. The shift edge sensing module divides the input features into two groups along the channel; The first group is multiplied and fused with the input features by shifting windows in the height and width directions, and then processed by a 3×3 depth convolution. The second group is processed sequentially through 1×1 convolution, 3×3 depthwise convolution, and Fourier channel attention; The outputs of the first and second groups are concatenated and then fused using a lightweight convolution before being output. S2. At the end of each encoding stage, the RGB features and all auxiliary modal features are input into the cross-attention fusion module to perform cross-modal information interaction and adaptive weighted fusion to generate the fusion features of the current level. S3. The multi-level fusion features from the encoder are fused and reconstructed to obtain a high-resolution feature map, and the final road surface crack segmentation result is output.

2. The method for road surface crack edge perception and segmentation based on multimodal cross-fusion according to claim 1, characterized in that, The frequency domain enhancement Mamba module adopts a three-branch parallel structure; The first branch is the selective state-space modeling branch, which extracts global features through selective state-space models in four directions; The second branch is the Fourier channel attention branch, which transforms the features from the spatial domain to the frequency domain through Fourier transform to perform global channel attention modeling. The third branch is a strip-gated convolutional branch, which sequentially passes through a 5×5 depth convolutional layer, a 1×7 depth separable convolutional layer, a 7×1 depth separable convolutional layer, a gating layer, and a 1×1 convolutional layer to extract horizontal and vertical contextual features. The outputs of the first, second, and third branches are gated and fused, and then adaptively connected to the input features using residual connections.

3. The method for road surface crack edge perception and segmentation based on multimodal cross-fusion according to claim 2, characterized in that, The Fourier channel attention branch is specifically as follows: in, These are respectively input features Perform Fourier transform and inverse transform. The feature map is obtained after Fourier transform. and These are the real and imaginary features in the frequency domain, multiplied by the learnable frequency domain weights, respectively. and These are the learnable frequency domain weights. and These are the true and imaginary values ​​taken from the feature map obtained after Fourier transform. The final generated channel attention weight vector has dimensions of , It is the Sigmoid activation function. The output characteristics of the Fourier channel attention branch, For element-wise multiplication, To enhance the original input features of the Mamba module in the frequency domain, It is a 1×1 convolutional layer.

4. The method for road surface crack edge sensing and segmentation based on multimodal cross-fusion according to any one of claims 1-3, characterized in that, In step S2, when the number of auxiliary modal features is greater than 1, the most discriminative auxiliary modal features are selected through a learnable scoring network before entering the cross-attention fusion module. Specifically, each auxiliary modal feature is generated into a single-channel sub-map through lightweight convolution, multiplied with itself and added together, and then the most discriminative auxiliary modal feature is output through max pooling.

5. The method for road surface crack edge sensing and segmentation based on multimodal cross-fusion according to any one of claims 1-3, characterized in that, Step S2 specifically includes: Layer normalization is performed on RGB features and auxiliary modal features respectively; Directional context is extracted using multi-scale band convolution kernels; Using the auxiliary modality feature as the query and the RGB feature as the key / value pair, calculate the first cross-attention and output the first cross-attention result; simultaneously, using the RGB feature as the query and the auxiliary modality feature as the key / value pair, calculate the second cross-attention and output the second cross-attention result. The first and second cross-attention results are concatenated along the channels, and after channel averaging and max pooling, a two-channel spatial weight map is generated through 7×7 convolution. The first and second cross-attention results are weighted and summed, and after 1×1 convolution, they are added to the original input features to output the fused features.

6. The method for road surface crack edge perception and segmentation based on multimodal cross-fusion according to any one of claims 1-3, characterized in that, Step S3 specifically includes: The decoder adopts a top-down progressive upsampling strategy, upsampling the highest-level fusion features step by step, and then performing cross-scale stitching and weighted fusion with the lower-level fusion features in turn. After each level of fusion, features are purified through convolution and activation operations, and finally the segmentation result with the same size as the input image is output.

7. The method for road surface crack edge sensing and segmentation based on multimodal cross-fusion according to any one of claims 1-3, characterized in that, The road surface crack edge sensing and segmentation method based on multimodal cross-fusion also includes: Model training involves end-to-end training of the dual encoder, cross-attention fusion module, and decoder using a labeled multimodal pavement crack sample set until the model converges. The model training employs a hybrid loss function, which is a weighted sum of binary cross-entropy loss and Dice loss.

8. A system comprising the pavement crack edge sensing and segmentation method based on multimodal cross-fusion as described in any one of claims 1-7, characterized in that, It includes a dual encoder module, a cross-modal fusion module, and a decoding and reconstruction module connected in sequence: The dual encoder module is used to perform hierarchical downsampling on the registered RGB image and at least one auxiliary modality image respectively, and extract multi-scale features through the general feature enhancement module; The cross-modal fusion module is used to perform bidirectional cross-attention fusion on RGB features and auxiliary modal features at the end of each encoding stage, generating fused features that are fed back to the encoding and decoding stages; The decoding and reconstruction module is used to progressively upsample and stitch together multi-level fused features to output pixel-level segmentation results of road surface cracks.

Citation Information

Patent Citations

  • Crack image segmentation method based on double encoders in complex environment

    CN117058382A

  • Highway pavement crack segmentation method and system based on multi-modal fusion and hypergraph convolution

    CN121982511A