Low-light-level-SAR (Synthetic Aperture Radar)-based cross-modal integrated feature fusion and road segmentation method
By combining the feature fusion and road segmentation of low-light-level image and SAR images with a method based on spatial-spectral perception network, the problem of independent image fusion and road segmentation tasks in the existing technology is solved, and high-quality fused images and accurate road segmentation effects are achieved.
Patent Information
- Application Number
- CN202510674075.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Existing low-light-level and SAR remote sensing image fusion methods fail to effectively combine image fusion and road segmentation tasks, resulting in the generated fused images possibly retaining irrelevant redundant features or losing key discriminative semantic information, making it difficult to meet the requirements of advanced vision tasks.
A method based on spatial-spectral perception network is adopted. A dual-branch structure is used to process low-light level and SAR images respectively. The CBS convolutional block and the cross-modal shared feature extraction module SFEM are used to share features. The segmentation results and fused images are obtained through the segmentation feature decoder and the fusion feature decoder. A semantic-driven strategy with dynamic factors is designed to optimize the loss function.
It achieves the coordinated optimization of image fusion and road segmentation, enhances the edge feature extraction capability, adaptively adjusts the weights of semantic segmentation loss and fusion loss, improves the quality of fused images and road segmentation accuracy, and adapts to the generalization of different application scenarios.
Smart Images

Figure CN120599463A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method for cross-modal integrated feature fusion and road segmentation based on low-light-level radar (LSL) and SAR. Background Art
[0002] Currently, multi-source remote sensing data used in road segmentation methods include low-light-level imagery and synthetic aperture radar (SAR) imagery. These two types of imagery have distinct characteristics and advantages. Low-light-level imagery refers to the passive capture of light signals by low-light detectors at night or in low-light conditions, which are converted into visible images. It can display the illumination conditions of urban lights at night and is therefore widely used in fields such as road extraction and urban planning. However, low-light-level imagery is highly dependent on lighting and weather conditions and is susceptible to interference from haze, resulting in a low signal-to-noise ratio (SNR) in the generated images. In contrast, SAR imagery is generated using active microwave imaging technology, which can penetrate clouds, smog, and some vegetation, providing highly accurate topographic information. It operates in all weather conditions and is unrestricted by lighting and weather conditions. However, the grayscale or pseudo-color display of SAR images differs significantly from human visual perception and is susceptible to noise such as coherent speckle, making image interpretation difficult.
[0003] Currently, relatively little research exists on low-light-level and SAR remote sensing image fusion methods and application scenarios. Existing image fusion techniques tend to produce fused images that conform to human visual perception and demonstrate high performance metrics, but they neglect the contribution of the fused image information to the road segmentation task, resulting in overly independent image fusion and road segmentation tasks. A few methods attempt to link image fusion with higher-level vision tasks, such as object detection and segmentation, using these higher-level vision tasks to guide the image fusion process. However, these methods are unable to complete both image fusion and higher-level vision tasks in an end-to-end manner, and struggle to achieve a performance balance between the two during training.
[0004] There is still a lack of corresponding patents for cross-modal fusion of low-light-level data with low signal-to-noise ratio (SNR) and synthetic aperture radar (SAR) data. Existing patents focus excessively on improving the visual quality indicators of the fused image (such as clarity and contrast), while neglecting to optimize the features required for subsequent high-level vision tasks. This results in the image fusion task being too independent of the high-level vision task. Although the Transformer encoder effectively improves the quality of the fused image, the optimization goal of this method is limited to improving the visual effect or local feature complementarity of the fused image, and fails to establish correlation constraints with high-level vision tasks such as semantic segmentation and object detection. As a result, the generated fused image may retain redundant features irrelevant to the task or lose key discriminative semantic information, making it difficult to meet the requirements of high-level vision tasks for high-level abstract semantic features. Summary of the Invention
[0005] In view of the existing technical defects in the background technology, the present invention proposes a method for cross-modal integrated feature fusion and road segmentation based on low-light-level radar (LLL-SAR).
[0006] The technical solution adopted by the present invention to solve this problem is: an agricultural hyperspectral image classification network based on a spatial-spectral perception network. The present invention firstly comprises the following steps:
[0007] Step S1, acquiring a low-light-level image and a SAR image, and preprocessing the low-light-level image;
[0008] Step S2: Process the low-light image and SAR image separately using a dual-branch structure: use the convolutional layer Conv2D to extract low-light detail features and SAR detail features, and use the Sobel operator and 1×1 convolution to enhance the edge feature information of the road in the low-light image.
[0009] Step S3: Based on the low-light level detail features, SAR detail features, and edge feature information, the low-light level coding features and SAR coding features are obtained by means of the CBS convolution block, and then the features are shared by the cross-modal shared feature extraction module SFEM.
[0010] Step S4: Obtain the segmentation results and fused image through the segmentation feature decoder and the fusion feature decoder, and construct a low-light-level SAR cross-modal integrated feature fusion and road segmentation method model, namely, the low-light-level SAR model;
[0011] Step S5, training the overall loss function of the low-light-level SAR model;
[0012] Step S6: Evaluate the performance of the low-light-level SAR model.
[0013] Preferably, in step S1, the specific process of acquiring the low-light-level image and the SAR image and preprocessing the low-light-level image is as follows:
[0014] (1) The process of acquiring low-light-level images and SAR images is as follows:
[0015] Obtain panchromatic low-light images with a resolution of 10 meters and SAR images with a resolution of 10 meters from remote sensing images;
[0016] (3) The process of preprocessing low-light images is as follows:
[0017] The low-light image is threshold segmented with the segmentation threshold set to (5, 255), and a binary image is output. Then, the binary image is processed using the dilation method, with a kernel of size (5, 5) and three dilation iterations. Finally, the distance transform method is used to process the disconnected roads with larger intervals, and the processed low-light image is output.
[0018] Preferably, in step S2, the specific process of the dual-branch structure processing the low-light image and the SAR image respectively is as follows:
[0019] The low-light image and SAR image are used as input data. The low-light image is input into the low-light feature extraction branch, and the SAR image is input into the SAR feature extraction branch. A two-dimensional convolutional layer Conv2D is used to extract the detail features of the low-light image and SAR image, and the number of channels of the low-light image and SAR image is modified to obtain low-light detail features and SAR detail features. At the same time, the low-light image is input into the edge feature guidance structure, and the edge feature information of the road is extracted using the Sobel operator. Then, a 1×1 convolutional layer Conv is used to adjust the extracted edge feature information.
[0020] Preferably, in step S3, the CBS convolution block is used to obtain the low-light level coding features and the SAR coding features, and the specific process of sharing the features through the cross-modal shared feature extraction module SFEM includes:
[0021] In step S3.1, the low-light feature extraction branch uses the CBS convolutional block to perform feature encoding on the low-light detail features to obtain low-light coded features. The specific process is as follows: first, convolution is used to extract the low-light detail features, then batch normalization (BN) is performed, and finally, the SiLU activation function is used to activate the feature information to obtain low-light feature information. The low-light feature information is then concatenated with the corresponding edge feature information in the channel dimension to obtain the low-light coded features.
[0022] For SAR detail features, the SAR feature extraction branch uses CBS convolution blocks to encode SAR detail features. The specific process is as follows: first, SAR detail features are extracted using convolution, then batch normalization (BN) is performed, and finally, the SiLU activation function is used to activate the feature information to obtain SAR coded features.
[0023] Step S3.2: In the cross-modal shared feature extraction module SFEM, the low-light level coded features and SAR coded features are input into the SFEM module. First, the input low-light level coded features and SAR coded features are irregularly extracted using deformable convolution (DeConv). The local spatial correspondence between the low-light level coded features and SAR coded features is adaptively adjusted to obtain the low-light level features and SAR features. Then, the low-light level features and SAR features are separated in the channel and spatial dimensions, and global information interaction between the low-light level features and SAR features is achieved in the channel and spatial dimensions, respectively. Finally, the low-light level features and SAR features output in the channel and spatial dimensions are aggregated to obtain cross-modal long-range dependency features, namely the shared feature identifier. The shared feature identifier is spliced with the low-light level coded features and SAR coded features in the channel dimension to obtain the spliced features.
[0024] In step S3.3, the concatenated features are input into different CBS convolution blocks of the next layer to obtain the low-light coding features and SAR coding features of the next layer respectively. Then, the features are input into the new SFEM module to repeat the processing flow of step S3.2. After the multi-layer processing flow, the shared features are realized.
[0025] Preferably, in step S4, the specific process of obtaining the segmentation result and the fused image through the segmentation feature decoder and the fusion feature decoder, and constructing the low-light-level light-SAR cross-modal integrated feature fusion and road segmentation method model is as follows:
[0026] After processing by the CBS convolutional block and the cross-modal shared feature extraction module (SFEM), the feature extraction results of the low-light feature extraction branch and the SAR feature extraction branch are concatenated in the channel dimension to obtain fused features. 1×1 convolution is then used to reduce the dimensionality of the fused features. The reduced fused features are then fed into the segmentation feature decoder and fusion feature decoder in the feature decoding stage to obtain the segmentation result and fused image.
[0027] The feature decoding stage includes segmentation feature decoding and fusion feature decoding, wherein the segmentation feature decoder consists of four segmentation feature upsampling modules SDM, which are used to aggregate cross-modal semantic shared features and fusion features at the same level, and upsample the fusion features after dimensionality reduction to obtain segmentation results; the fusion feature decoder consists of four fusion feature upsampling modules FDM, which are used to aggregate cross-modal semantic shared features and fusion features at the same level, and through residual connections, avoid the loss of pixel-level features after multi-layer convolution extraction to obtain aggregated features. The aggregated features are then upsampled to reconstruct the fused image, and a low-light-level SAR cross-modal integrated feature fusion and road segmentation method model, namely the low-light-level SAR model, is constructed.
[0028] Preferably, in step S5, the process of training the overall loss function of the low-light-level SAR model is as follows:
[0029] Design an overall loss function, which consists of two parts: fusion loss and semantic-driven loss. The expression of the overall loss function is shown in Formula 1:
[0030] L total =L fu +αL seg_dri (1)
[0031] Among them, L total represents the overall loss of the proposed method, L fu represents the fusion loss, L seg_dri represents semantic driven loss, α represents dynamic weight factor;
[0032] The fusion loss consists of pixel intensity loss and texture loss. The fusion loss expression is shown in Formula 2:
[0033] L fu =L int +λL texture (2)
[0034] Among them, L int represents the pixel intensity loss, L texture represents texture loss, λ represents weight factor;
[0035] The pixel intensity loss expression is shown in Formula 3:
[0036]
[0037] Among them, H, W represent the image size, I f represents the fusion image reconstructed by the fusion network, I w ,I s denote a pair of low-light-level images and SAR images, respectively. ||·||1 denotes L1 regularization, and max(·) denotes element-wise maximum value calculation.
[0038] The texture loss expression is shown in Formula 4:
[0039]
[0040] Where, ▽ represents the Sobel gradient operator;
[0041] A semantic-driven strategy based on dynamic factors is designed. By designing a dynamically changing weighting factor, the coordinated optimization of the image fusion task and the illuminated road segmentation task is ensured. The weight of the semantic-driven loss in the total loss is adjusted by the dynamic weight factor α, so that the algorithm tends to rely on the fusion loss to update the network parameters in the initial training stage. As the training progresses, the dynamic weight factor gradually increases, and the parameters of the illuminated road segmentation part are effectively updated, gradually injecting semantic feature information into the fusion feature. The definition of α is shown in Formula 5:
[0042]
[0043] Among them, t∈[0,1] represents the ratio between the number of iterations of the current training process and the total number of iterations, η is a hyperparameter of the semantic-driven loss size, and e represents a natural constant, which is a mathematical constant.
[0044] Preferably, in step S6, the process of evaluating the performance of the low-light-level SAR model is as follows:
[0045] The mean intersection-over-union (MIoU) is used as the segmentation evaluation metric to calculate the degree of overlap between the predicted segmentation and the true segmentation label. The expression is shown in Formula 6:
[0046]
[0047] Among them, N represents the number of segmentation categories, TP i Indicates the number of pixels predicted correctly for the i-th category, FP i Indicates the number of pixels predicted incorrectly for the i-th category, FN i Indicates the number of pixels that actually belong to the i-th category but are not correctly predicted.
[0048] Compared with the existing technology, the present invention proposes a method for cross-modal integration of low-light-level light-source SAR (LLL-SAR) feature fusion and road segmentation. The present invention has the following beneficial effects:
[0049] (1) Integrated cross-modal feature sharing architecture to achieve collaborative optimization of image fusion and road segmentation
[0050] This paper couples the low-light-level light and SAR image fusion task and the illuminated road segmentation task into the same framework and designs a cross-modal shared feature extraction module (SFEM). This allows the two tasks to share cross-modal features, prompting the fusion network to generate feature representations that meet both visual quality requirements and semantic segmentation needs. Through joint optimization, the fusion network can enhance semantic features related to road segmentation while maintaining high-quality images, avoiding the generation of irrelevant redundant information, thereby improving the performance of the overall task.
[0051] (2) Edge feature guided structure enhances road edge perception in low-light environments
[0052] This paper introduces an edge feature guidance structure into the low-light-level image feature extraction branch to explicitly enhance the ability to extract edge information. This structure can highlight the contrast and structural features of road edges during the fusion process, thereby improving the recognition ability of the segmentation network. In particular, it can maintain robust road extraction in low signal-to-noise ratio scenarios and reduce edge blur and breakage issues.
[0053] (3) Semantic-driven strategy based on dynamic factors to achieve adaptive weight balance between tasks
[0054] This paper proposes a semantic-driven strategy based on dynamic factors to adaptively adjust the weights of semantic segmentation loss and fusion loss. In the early stages of training, this strategy focuses on optimizing the fusion task to ensure the robustness of the fused features. As training progresses, the influence of the semantic-driven loss is gradually increased, allowing the fused features to continuously adapt to the needs of the road segmentation task. This dynamic adjustment mechanism effectively avoids the interference of semantic tasks on early feature learning, while ensuring that the final model is able to balance fusion quality and segmentation accuracy, thereby improving the generalization and practicality of the method in different application scenarios.
[0055] (4) Pioneering low-light level and SAR cross-modal fusion technology
[0056] This method proposes a dedicated framework for the cross-modal fusion of low-light and SAR data for the first time. Through network design, it deeply combines the detailed texture of low-light images with the strong penetrating features of SAR, solving the technical problem of multimodal data fusion in low-light environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a flow chart of the method of the present invention;
[0058] Figure 2 Detailed network structure diagram of the method of the present invention;
[0059] Figure 3 This is a flowchart of the preprocessing of low-light-level images according to the present invention;
[0060] Figure 4 Detailed structural diagram of the cross-modal shared feature extraction module SFEM of the present invention;
[0061] Figure 5 These are structural diagrams of the segmentation upsampling module SDM and the fusion upsampling module FDM of the present invention, where (a) is the structural diagram of the segmentation upsampling module SDM, and (b) is the structural diagram of the fusion upsampling module FDM. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of this application to further clearly and completely describe the technical solutions in the embodiments of this application. It should be noted that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making any creative work shall fall within the scope of protection of this application.
[0063] In order to make the invention objectives, technical solutions and advantages of this application clearer, the embodiments of this application are further described in detail in conjunction with the drawings in the specification: In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the advantages of the present invention will be further illustrated by comparing the embodiments in conjunction with the drawings and specific implementation methods.
[0064] The present invention proposes a method for cross-modal integration of low-light-level SAR (LLL-SAR) and feature fusion and road segmentation. The flowchart of the method of the present invention is as follows: Figure 1 The detailed network structure diagram of the method of the present invention is shown in FIG. Figure 2 As shown in the figure, a pair of registered low-light image and SAR image are first input into the network. Secondly, the network uses the low-light feature extraction branch and the SAR feature extraction branch to extract features from them respectively. In particular, the low-light image needs to be subjected to the edge feature guided structure for feature extraction, and then it is subjected to the CBS convolution block to obtain low-light coding features and SAR coding features. Then, the cross-modal shared feature extraction module SFEM is used to fuse them and output them to the segmentation decoder and fusion decoder respectively. Finally, the segmentation decoder decodes the cross-modal shared features to obtain the road segmentation result, and the fusion decoder reconstructs the cross-modal shared features and outputs the fused image.
[0065] The process of the present invention specifically includes the following steps:
[0066] Step S1, acquiring a low-light-level image and a SAR image, and preprocessing the low-light-level image;
[0067] Specifically, in step S1, the specific process of acquiring the low-light-level image and the SAR image and preprocessing the low-light-level image is as follows:
[0068] (1) The process of acquiring low-light-level images and SAR images is as follows:
[0069] Obtain panchromatic low-light images with a resolution of 10 meters and SAR images with a resolution of 10 meters from remote sensing images;
[0070] (2) The process of preprocessing low-light images is as follows:
[0071] The preprocessing flow chart of low-light image is as follows Figure 3As shown in the figure, since the nighttime roads are blocked by trees, clouds and fog, the low-light data detected by satellites have road disconnection problems. Morphological methods and distance transform methods are used to preprocess the low-light images. In order to balance the interference of uneven illumination in the low-light images on the subsequent road segmentation task, the low-light images are threshold segmented with the segmentation threshold set to (5, 255) and a binary image is output. Then, in order to connect the disconnected roads with small intervals, the binary image is processed using the dilation method. The kernel size is (5, 5) and the number of dilation iterations is 3. Finally, the distance transform method is used to process the disconnected roads with large intervals and the processed low-light image is output. This method can better connect the disconnected roads in the low-light image and reduce the impact of occlusion such as trees and clouds on the quality of the low-light image. At the same time, the road pixel values of the processed low-light image are consistent, avoiding the influence of uneven illumination on the subsequent algorithms.
[0072] Step S2: Process the low-light image and SAR image separately using a dual-branch structure: use the convolutional layer Conv2D to extract low-light detail features and SAR detail features, and use the Sobel operator and 1×1 convolution to enhance the edge feature information of the road in the low-light image.
[0073] Specifically, in step S2, the specific process of the dual-branch structure processing the low-light image and the SAR image respectively is as follows:
[0074] The low-light image and SAR image are used as input data. The low-light image is input into the low-light feature extraction branch, and the SAR image is input into the SAR feature extraction branch. The input low-light image and SAR image must be paired, that is, their geographical locations correspond to each other. A 2D convolutional layer Conv2D is used to extract detail features such as texture and edges from the low-light image and SAR image. The number of channels of the low-light image and SAR image is modified to obtain low-light detail features and SAR detail features, thereby enhancing the feature expression of the image in the channel dimension. At the same time, the low-light image is input into the edge feature guidance structure, and the edge feature information of the road is extracted using the Sobel operator. Then, a 1×1 convolutional layer Conv is used to adjust the extracted edge feature information.
[0075] Step S3: Based on the low-light level detail features, SAR detail features, and edge feature information, the low-light level coding features and SAR coding features are obtained by means of the CBS convolution block, and then the features are shared by the cross-modal shared feature extraction module SFEM.
[0076] Specifically, in step S3, the CBS convolution block is used to obtain the low-light level coding features and the SAR coding features, and the specific process of sharing the features through the cross-modal shared feature extraction module SFEM includes:
[0077] In step S3.1, the low-light feature extraction branch uses the CBS convolution block (CBS1) to encode the low-light detail features to obtain low-light coded features. The specific process is as follows: First, convolution is used to extract the low-light detail features. Then, batch normalization (BN) is performed to ensure that the features have a similar distribution across different batches. Finally, the SiLU activation function is used to activate the feature information. The SiLU activation function introduces nonlinear factors to enhance the expression ability and obtain low-light feature information. The low-light feature information is then concatenated with the corresponding edge feature information in the channel dimension to obtain the low-light coded features.
[0078] For SAR detail features, the SAR feature extraction branch uses CBS convolution block (CBS5) to encode SAR detail features. The specific process is as follows: SAR detail features are first extracted using convolution, then batch normalization (BN) processing is performed, and finally the SiLU activation function is used to activate the feature information to obtain SAR encoding features.
[0079] Step S3.2: Design a cross-modal shared feature extraction module (SFEM) to extract shared feature identifiers from low-light-level images and SAR images. The shared feature identifiers are then gradually injected into the low-light-level feature extraction branch and the SAR feature extraction branch to achieve feature information interaction between low-light-level images and SAR images.
[0080] In the cross-modal shared feature extraction module SFEM1, the low-light coding features and SAR coding features are input into the SFEM1 module. The detailed structure diagram of the cross-modal shared feature extraction module SFEM is shown in the figure. Figure 4 As shown in the figure, “DeConv” represents deformable convolution, Q, K, and V represent query vector, key vector, and value vector respectively, and “Softmax” represents the Softmax function, which is used to calculate the correlation weight of low-light coding features and SAR coding features;
[0081] In order to avoid the interference of the subsequent shared feature identification extraction caused by the offset of the position and shape of the same feature in the two images, in the cross-modal shared feature extraction module SFEM, the deformable convolution DeConv is first used to perform irregular extraction on the input low-light coding features and SAR coding features, and the local spatial correspondence between the low-light coding features and SAR coding features is adaptively adjusted to obtain low-light features and SAR features; then, the low-light features and SAR features are separated in the channel and spatial dimensions, and global information interaction between the low-light features and SAR features is realized in the channel and spatial dimensions respectively; finally, the low-light features and SAR features output in the channel and spatial dimensions are aggregated to obtain cross-modal long-range dependency features, namely shared feature identification; the shared feature identification is spliced with the low-light coding features and SAR coding features in the channel dimension to obtain spliced features;
[0082] In step S3.3, the concatenated features are fed into different CBS convolutional blocks (CBS2 and CBS6) of the next layer to obtain the low-light level coded features and SAR coded features of the next layer, respectively. These features are then fed into a new SFEM module and the processing flow of step S3.2 is repeated. After four layers of processing, shared features are achieved.
[0083] In the four-layer processing flow, the four CBS convolution blocks used by the low-light feature extraction branch are defined as CBS1, CBS2, CBS3, and CBS4. Similarly, the four CBS convolution blocks used by the SAR feature extraction branch are defined as CBS5, CBS6, CBS7, and CBS8.
[0084] Since the CBS convolution block continuously processes the input low-light detail features and SAR detail features and inputs the features into the SFEM module, a multi-level feature information interaction channel from shallow features to deep semantic features of low-light images and SAR images is established, processing cross-modal interaction information of different fine-grained levels respectively; based on this information interaction channel, the model can gradually combine the shallow details (such as edges and textures) of low-light images and SAR images with deep semantic information (such as object categories and structural features), thereby constructing a robust feature map that can be applied to segmentation and fusion tasks.
[0085] Step S4, obtaining the segmentation results and fused images through the segmentation feature decoder and the fusion feature decoder, and constructing a low-light-level-SAR cross-modal integrated feature fusion and road segmentation method model;
[0086] Specifically, in step S4, the segmentation result and fused image are obtained through the segmentation feature decoder and the fusion feature decoder, and the specific process of constructing the low-light-level SAR cross-modal integrated feature fusion and road segmentation method model is as follows:
[0087] After processing by the CBS convolutional block and the cross-modal shared feature extraction module (SFEM), the feature extraction results of the low-light feature extraction branch and the SAR feature extraction branch are concatenated in the channel dimension to obtain fused features. The fused features are then reduced in dimension using a 1×1 convolution. The reduced fused features are then fed into the segmentation feature decoder and fusion feature decoder in the feature decoding stage to obtain the final segmentation result and fused image.
[0088] The structure diagram of the segmentation upsampling module SDM is as follows Figure 5 As shown in (a), “BN” represents the batch normalization layer, “ReLU” represents the rectified linear unit, and “Upsample” represents the bilinear interpolation upsampling layer; the structure of the fusion upsampling module FDM is shown in Figure 5 (b) shows that “+” represents residual connection;
[0089] The feature decoding stage includes segmentation feature decoding and fusion feature decoding, wherein the segmentation feature decoder consists of four segmentation upsampling modules SDM, which are used to aggregate cross-modal semantic shared features and fusion features at the same level, and upsample the fusion features after dimensionality reduction to obtain segmentation results; the fusion feature decoder consists of four fusion feature upsampling modules FDM, which are used to aggregate cross-modal semantic shared features and fusion features at the same level, and through residual connections, avoid the loss of pixel-level features after multi-layer convolution extraction to obtain aggregated features. The aggregated features are then upsampled to reconstruct the fused image, and a low-light-level SAR cross-modal integrated feature fusion and road segmentation method model, namely the low-light-level SAR model, is constructed.
[0090] Step S5, training the overall loss function of the low-light-level SAR model;
[0091] Specifically, in step S5, the specific process of training the overall loss function of the low-light-level SAR model is as follows:
[0092] In order to generate high-quality fused images and enhance the semantic information in the fused features to improve the road segmentation accuracy of the model, an overall loss function is designed. The overall loss function consists of two parts: fusion loss and semantic-driven loss. The expression of the overall loss function is shown in Formula 1:
[0093] L total =L fu +αL seg_dri (1)
[0094] Among them, L total represents the overall loss of the proposed method, L fu represents the fusion loss, L seg_dri represents semantic driven loss, α represents dynamic weight factor;
[0095] The fusion loss optimizes the pixel intensity and texture similarity between the fused image and the source image to guide the fusion network to extract the complementary feature information of the low-light image and the SAR image. To achieve this goal, the fusion loss consists of pixel intensity loss and texture loss. The fusion loss expression is shown in Formula 2:
[0096] L fu =L int +λL texture (2)
[0097] Among them, L int represents the pixel intensity loss, L texture represents texture loss, λ represents weight factor;
[0098] Pixel intensity loss is used to evaluate the difference in pixel intensity between the fused image and the source image. The pixel intensity loss expression is shown in Formula 3:
[0099]
[0100] Among them, H, W represent the image size, I f represents the fusion image reconstructed by the fusion network, I w ,I s denote a pair of low-light-level images and SAR images, respectively. ||·||1 denotes L1 regularization, and max(·) denotes element-wise maximum value calculation.
[0101] Texture loss is used to constrain the fusion network to extract complementary texture feature information of low-light images and SAR images. The texture loss expression is shown in Formula 4:
[0102]
[0103] in, represents the Sobel gradient operator;
[0104] A semantic-driven strategy based on dynamic factors is designed. By designing a dynamically changing weighting factor, the coordinated optimization of the image fusion task and the illuminated road segmentation task is ensured. The weight of the semantic-driven loss in the total loss is adjusted by the dynamic weight factor α, so that the algorithm tends to rely on the fusion loss to update the network parameters in the initial training stage. As training progresses, the dynamic weight factor gradually increases, and the parameters of the illuminated road segmentation part are effectively updated, gradually injecting semantic feature information into the fusion feature. The definition of α is shown in Formula 5:
[0105]
[0106] Among them, t∈[0,1] represents the ratio between the number of iterations of the current training process and the total number of iterations, η is a hyperparameter of the semantic-driven loss size, and e represents a natural constant, which is a mathematical constant.
[0107] Step S6, low-light-level SAR model performance evaluation;
[0108] Specifically, in step S6, the process of evaluating the performance of the low-light-level SAR model is as follows:
[0109] The mean intersection over union (MIoU) is used as the segmentation evaluation metric to calculate the degree of overlap between the predicted segmentation and the true segmentation label. The expression is shown in Formula 6:
[0110]
[0111] Among them, N represents the number of segmentation categories, TP i Indicates the number of pixels predicted correctly for the i-th category, FP i Indicates the number of pixels predicted incorrectly for the i-th category, FN i Indicates the number of pixels that actually belong to the i-th category but are not correctly predicted.
[0112] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0113] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. Based on the cross-modal integrated feature fusion and road segmentation method of low-light-level light-SAR, the characteristics are: include: Step S1, acquiring a low-light-level image and a SAR image, and preprocessing the low-light-level image; Step S2: Process the low-light image and SAR image separately using a dual-branch structure: use the convolutional layer Conv2D to extract low-light detail features and SAR detail features, and use the Sobel operator and 1×1 convolution to enhance the edge feature information of the road in the low-light image. Step S3: Based on the low-light level detail features, SAR detail features, and edge feature information, the low-light level coding features and SAR coding features are obtained by means of the CBS convolution block, and then the features are shared by the cross-modal shared feature extraction module SFEM. Step S4: Obtain the segmentation results and fused image through the segmentation feature decoder and the fusion feature decoder, and construct a low-light-level SAR cross-modal integrated feature fusion and road segmentation method model, namely, the low-light-level SAR model; Step S5, training the overall loss function of the low-light-level SAR model; Step S6: Evaluate the performance of the low-light-level SAR model.
2. The method for cross-modal integration of feature fusion and road segmentation based on low-light-level SAR according to claim 1 is characterized in that: In step S1, the specific process of obtaining the low-light-level image and the SAR image and preprocessing the low-light-level image is as follows: (1) The process of acquiring low-light-level images and SAR images is as follows: Obtain panchromatic low-light images with a resolution of 10 meters and SAR images with a resolution of 10 meters from remote sensing images; (2) The process of preprocessing low-light images is as follows: The low-light image is threshold segmented with the segmentation threshold set to (5, 255), and a binary image is output. Then, the binary image is processed using the dilation method, with a kernel of size (5, 5) and three dilation iterations. Finally, the distance transform method is used to process the disconnected roads with larger intervals, and the processed low-light image is output.
3. The method for cross-modal integration of feature fusion and road segmentation based on low-light-level SAR according to claim 1 is characterized in that: In step S2, the specific process of the dual-branch structure processing the low-light image and the SAR image respectively is as follows: The low-light image and SAR image are used as input data. The low-light image is input into the low-light feature extraction branch, and the SAR image is input into the SAR feature extraction branch. A two-dimensional convolutional layer Conv2D is used to extract the detail features of the low-light image and SAR image, and the number of channels of the low-light image and SAR image is modified to obtain low-light detail features and SAR detail features. At the same time, the low-light image is input into the edge feature guidance structure, and the edge feature information of the road is extracted using the Sobel operator. Then, a 1×1 convolutional layer Conv is used to adjust the extracted edge feature information.
4. The method for cross-modal integration of feature fusion and road segmentation based on low-light-level SAR according to claim 1 is characterized in that: In step S3, the CBS convolution block is used to obtain the low-light coding features and the SAR coding features. The specific process of sharing the features through the cross-modal shared feature extraction module SFEM includes: In step S3.1, the low-light feature extraction branch uses the CBS convolutional block to perform feature encoding on the low-light detail features to obtain low-light coded features. The specific process is as follows: first, convolution is used to extract the low-light detail features, then batch normalization (BN) is performed, and finally, the SiLU activation function is used to activate the feature information to obtain low-light feature information. The low-light feature information is then concatenated with the corresponding edge feature information in the channel dimension to obtain the low-light coded features. For SAR detail features, the SAR feature extraction branch uses CBS convolution blocks to encode SAR detail features to obtain SAR coded features. The specific process is as follows: first, SAR detail features are extracted using convolution, then batch normalization (BN) processing is performed, and finally, the SiLU activation function is used to activate the feature information to obtain SAR coded features. Step S3.2: In the cross-modal shared feature extraction module SFEM, the low-light level coded features and SAR coded features are input into the SFEM module. First, the input low-light level coded features and SAR coded features are irregularly extracted using deformable convolution (DeConv). The local spatial correspondence between the low-light level coded features and SAR coded features is adaptively adjusted to obtain the low-light level features and SAR features. Then, the low-light level features and SAR features are separated in the channel and spatial dimensions, and global information interaction between the low-light level features and SAR features is achieved in the channel and spatial dimensions, respectively. Finally, the low-light level features and SAR features output in the channel and spatial dimensions are aggregated to obtain cross-modal long-range dependency features, namely the shared feature identifier. The shared feature identifier is spliced with the low-light level coded features and SAR coded features in the channel dimension to obtain the spliced features. In step S3.3, the concatenated features are input into different CBS convolution blocks of the next layer to obtain the low-light coding features and SAR coding features of the next layer respectively. Then, the features are input into the new SFEM module to repeat the processing flow of step S3.
2. After the multi-layer processing flow, the shared features are realized.
5. The method for cross-modal integration of feature fusion and road segmentation based on low-light-level SAR according to claim 1 is characterized in that: In step S4, the specific process of obtaining the segmentation result and the fused image through the segmentation feature decoder and the fusion feature decoder, and constructing the low-light-level-SAR cross-modal integrated feature fusion and road segmentation method model is as follows: After processing by the CBS convolutional block and the cross-modal shared feature extraction module (SFEM), the feature extraction results of the low-light feature extraction branch and the SAR feature extraction branch are concatenated in the channel dimension to obtain fused features, which are then reduced in dimension using 1×1 convolution. Afterwards, the fusion features after dimension reduction are input into the segmentation feature decoder and fusion feature decoder in the feature decoding stage respectively to obtain the segmentation results and fusion images; The feature decoding stage includes segmentation feature decoding and fusion feature decoding. The segmentation feature decoder consists of four segmentation feature upsampling modules SDM. The SDM module is used to aggregate cross-modal semantic shared features and fusion features at the same level, and upsample the fusion features after dimension reduction to obtain segmentation results. The fusion feature decoder contains four fusion feature upsampling modules FDM. The FDM module is used to aggregate cross-modal semantic shared features and fusion features at the same level. Through residual connection, the aggregated features are obtained. The aggregated features are then upsampled to reconstruct the fused image, constructing a low-light-level SAR cross-modal integrated feature fusion and road segmentation method model, namely the low-light-level SAR model.
6. The method for cross-modal integration of feature fusion and road segmentation based on low-light-level SAR according to claim 1 is characterized in that: In step S5, the process of training the overall loss function of the low-light-level SAR model is as follows: Design an overall loss function, which consists of two parts: fusion loss and semantic-driven loss. The expression of the overall loss function is shown in Formula 1: L total =L fu +αL seg_dri (1) Among them, L total represents the overall loss of the proposed method, L fu represents the fusion loss, L seg_dri represents semantic driven loss, α represents dynamic weight factor; The fusion loss consists of pixel intensity loss and texture loss. The fusion loss expression is shown in Formula 2: THE fu =L int +λL texture (2) Among them, L int represents the pixel intensity loss, L texture represents texture loss, λ represents weight factor; The pixel intensity loss expression is shown in Formula 3: Among them, H, W represent the image size, I f represents the fusion image reconstructed by the fusion network, I w ,I s denote a pair of low-light-level images and SAR images, respectively. ||·||1 denotes L1 regularization, and max(·) denotes element-wise maximum value calculation. The texture loss expression is shown in Formula 4: in, represents the Sobel gradient operator; A semantic-driven strategy based on dynamic factors is designed. By designing a dynamically changing weighting factor, the coordinated optimization of the image fusion task and the illuminated road segmentation task is ensured. The weight of the semantic-driven loss in the total loss is adjusted by the dynamic weight factor α, so that the algorithm tends to rely on the fusion loss to update the network parameters in the initial training stage. As the training progresses, the dynamic weight factor gradually increases, and the parameters of the illuminated road segmentation part are effectively updated, gradually injecting semantic feature information into the fusion feature. The definition of α is shown in Formula 5: Among them, t∈[0,1] represents the ratio between the number of iterations of the current training process and the total number of iterations, η is a hyperparameter of the semantic-driven loss size, and e represents a natural constant, which is a mathematical constant.
7. The method for cross-modal integration of feature fusion and road segmentation based on low-light-level SAR according to claim 1, characterized in that: In step S6, the process of evaluating the performance of the low-light-level SAR model is as follows: The mean intersection-over-union (MIoU) is used as the segmentation evaluation metric to calculate the degree of overlap between the predicted segmentation and the true segmentation label. The expression is shown in Formula 6: Among them, N represents the number of segmentation categories, TP i Indicates the number of pixels predicted correctly for the i-th category, FP i Indicates the number of pixels predicted incorrectly for the i-th category, FN i Indicates the number of pixels that actually belong to the i-th category but are not correctly predicted.
Citation Information
Patent Citations
Multi-modal visual information fusion method and system based on modal feature constraint
CN117456332A
Remote sensing image segmentation method based on optical image-SAR image feature alignment
CN117611813A
RGB-D indoor scene semantic segmentation method of edge threshold based on feature calibration
CN118711184A
Cross-modal medical image generation method and apparatus
WO2024087218A1
Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment
WO2024230038A1