A Remote Sensing Image Segmentation Method and System Based on Hierarchical Multimodal Feature Fusion

By employing a hierarchical multimodal feature fusion method, combined with CNN and Transformer branches, the problems of data redundancy and computational complexity in complex scenarios of hyperspectral and synthetic aperture radar images in remote sensing image processing are solved. This enables more refined information filtering and fusion, improving the accuracy of remote sensing image segmentation and the robustness of the model.

CN120526144BActive Publication Date: 2025-10-31耕宇牧星(北京)空间科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510601969.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-10-31
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In remote sensing image processing, existing technologies suffer from data redundancy, high computational complexity, and weak anti-interference capabilities in complex scenes for hyperspectral and synthetic aperture radar images. Furthermore, traditional methods are ineffective in interactive fusion of cross-modal information, resulting in insufficient accuracy in remote sensing image segmentation.

Method used

A hierarchical multimodal feature fusion method is adopted, which uses CNN branches to extract local texture features of hyperspectral images and Transformer branches to perform global feature modeling. Through layer-by-layer optimization and modal interaction mechanisms, combined with residual connections, feature weighting calculation and information fusion of hyperspectral images and synthetic aperture radar images are realized.

Benefits of technology

It improves the accuracy and generalization ability of remote sensing image segmentation, enhances the robustness of the model, and provides a new solution for multi-source remote sensing image fusion and intelligent analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526144B_ABST
    Figure CN120526144B_ABST
Patent Text Reader

Abstract

This invention discloses a remote sensing image segmentation method and system based on hierarchical multimodal feature fusion, belonging to the field of remote sensing image processing technology. The method includes: acquiring hyperspectral and synthetic aperture radar (SAR) images of the target to be segmented; inputting the hyperspectral image into an HSI feature extraction network to obtain a first aggregated feature; inputting the SAR image into a SAR feature extraction network to obtain a second aggregated feature; inputting the first aggregated feature into a first weight acquisition network to obtain a first correlation weight; fusing the second aggregated feature with the first correlation weight to obtain a modal interaction feature; inputting the modal interaction feature into a second weight acquisition network to obtain a second correlation weight; fusing the first aggregated feature with the second correlation weight to obtain multimodal remote sensing features; and inputting the multimodal remote sensing features into a classification output network to obtain the segmentation result of the target to be segmented. This improves the accuracy of remote sensing image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, and more specifically to a remote sensing image segmentation method and system based on hierarchical multimodal feature fusion. Background Technology

[0002] Hyperspectral imagery (HSI) and synthetic aperture radar (SAR) imagery are two important data sources in the field of remote sensing. Hyperspectral imagery, due to its rich spectral information, is of significant value in target identification, environmental monitoring, and agricultural and forestry applications. However, due to the redundancy of spectral information and the influence of atmospheric and sensor noise, hyperspectral imagery suffers from data redundancy, high computational complexity, and weak anti-interference capabilities in certain complex scenarios. In contrast, synthetic aperture radar (SAR) imagery, acquired through microwave remote sensing, is unaffected by weather and lighting conditions, possesses strong penetration capabilities, and can effectively provide structural information about ground features. However, SAR imagery suffers from speckle noise and lacks rich spectral information, making it difficult to accurately identify certain target categories. Utilizing the complementary characteristics of hyperspectral and SAR imagery for remote sensing image analysis has become a research hotspot.

[0003] However, traditional multimodal remote sensing image processing methods are mostly based on manual feature extraction, relying mainly on statistical methods or simple filters for data fusion. These methods often perform poorly when dealing with complex scenes. In recent years, deep learning technology has made significant progress in remote sensing image segmentation, especially convolutional neural networks (CNNs), which excel in extracting spatial features and pattern recognition. However, the local receptive field of CNNs limits their ability to model long-range dependencies, potentially leading to misclassification when dealing with complex land cover categories. Furthermore, the Transformer architecture, with its global attention mechanism, has achieved breakthroughs in computer vision tasks, capable of capturing long-range dependencies. However, its direct application to remote sensing image processing incurs high computational costs and is prone to losing texture details in the absence of local feature guidance.

[0004] Therefore, how to improve the accuracy of remote sensing image segmentation based on the interactive fusion of cross-modal information is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a remote sensing image segmentation method and system based on hierarchical multimodal feature fusion, which improves the accuracy of remote sensing image segmentation based on interactive fusion of cross-modal information.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A remote sensing image segmentation method based on hierarchical multimodal feature fusion includes:

[0008] Acquire hyperspectral and synthetic aperture radar images of the target to be segmented;

[0009] The first aggregated feature is obtained by inputting the hyperspectral image into the HSI feature extraction network;

[0010] The synthetic aperture radar image is input into the SAR feature extraction network to obtain the second aggregated feature;

[0011] Based on the first aggregated feature, the first weight acquisition network is input to obtain the first relevant weight;

[0012] Based on the fusion of the second aggregation feature and the first relevant weight, modal interaction features are obtained;

[0013] The modal interaction features are input into the second weight acquisition network to obtain the second relevant weights;

[0014] Based on the fusion of the first aggregated features and the second related weights, multimodal remote sensing features are obtained;

[0015] The multimodal remote sensing features are input into the classification output network to obtain the segmentation result of the target to be segmented.

[0016] Preferably, the HSI feature extraction network includes: a location embedding layer, a first hybrid enhancement Transformer unit, a first downsampling layer, a second hybrid enhancement Transformer unit, a second downsampling layer, a third hybrid enhancement Transformer unit, and a first multi-scale aggregation unit;

[0017] The hyperspectral image is input into the location embedding layer to obtain the embedding features;

[0018] The embedded features are sequentially input into the first hybrid enhancement Transformer unit and the first downsampling layer to obtain the first scale features;

[0019] The first scale feature is sequentially input into the second hybrid enhancement Transformer unit and the second downsampling layer to obtain the second scale feature;

[0020] The second-scale feature is input into the third hybrid enhancement Transformer unit to obtain the third-scale feature;

[0021] The first scale feature, the second scale feature, and the third scale feature are input into the first multi-scale aggregation unit to obtain the first aggregated feature.

[0022] Preferably, the first hybrid enhancement Transformer unit, the second hybrid enhancement Transformer unit, and the third hybrid enhancement Transformer unit have the same structure, each including:

[0023] The system consists of a first normalized layer, a pooling layer, a first depthwise separable convolutional layer, a second normalized layer, a first activation layer, a third normalized layer, and a multilayer perceptron.

[0024] The first input feature is input into the first normalization layer to obtain the first processed feature;

[0025] The first processing feature is input to the pooling layer to obtain the second processing feature;

[0026] The first processing feature and the second processing feature are fused together to obtain the first fused feature;

[0027] The first fusion feature is sequentially input into the first depthwise separable convolutional layer, the second normalization layer, and the first activation layer to obtain the third processed feature;

[0028] The first fusion feature and the third processing feature are fused to obtain the second fusion feature;

[0029] The second fusion feature is sequentially input into the third normalization layer and the multilayer perceptron to obtain the fourth processing feature;

[0030] After the second fusion feature and the fourth processing feature are fused, the first output feature is obtained.

[0031] Preferably, the SAR feature extraction network includes: an initial feature extraction unit, a first multi-layer feature extraction unit, a second multi-layer feature extraction unit, and a second multi-scale aggregation unit;

[0032] The synthetic aperture radar image is input to the initial feature extraction unit to obtain fourth-scale features;

[0033] The fourth-scale feature is input into the first multi-layer feature extraction unit to obtain the fifth-scale feature;

[0034] The fifth-scale feature is input into the second multi-layer feature extraction unit to obtain the sixth-scale feature;

[0035] The fourth-scale feature, the fifth-scale feature, and the sixth-scale feature are input into the second multi-scale aggregation unit to obtain the second aggregated feature.

[0036] Preferably, the initial feature extraction unit includes: a first convolutional layer, a fourth normalization layer, a second activation layer, and a third downsampling layer;

[0037] The synthetic aperture radar image is input into the first convolutional layer to obtain the first intermediate feature;

[0038] The first intermediate feature is sequentially input into the fourth normalization layer and the second activation layer to obtain the second intermediate feature;

[0039] The first intermediate feature and the second intermediate feature are fused and then input into the third downsampling layer to obtain the fourth scale feature.

[0040] Preferably, the first multi-layer feature extraction unit and the second multi-layer feature extraction unit have the same structure, both including: a second convolutional layer, a fifth normalization layer, a third activation layer and a fourth downsampling layer;

[0041] The second output feature is sequentially input into the second convolutional layer, the fifth normalization layer, and the third activation layer to obtain the third intermediate feature;

[0042] The third intermediate feature and the second input feature are fused and then input into the fourth downsampling layer to obtain the second output feature;

[0043] The second input feature is either the fourth scale feature or the fifth scale feature, and the second output feature corresponds to either the fifth scale feature or the sixth scale feature.

[0044] Preferably, the first weight acquisition network includes: a second depthwise separable convolutional layer, a sixth normalization layer, a fourth activation layer, a third depthwise separable convolutional layer, a seventh normalization layer, and a fifth activation layer;

[0045] The first aggregated feature is sequentially input into the second depthwise separable convolutional layer, the sixth normalized layer, the fourth activation layer, the third depthwise separable convolutional layer, the seventh normalized layer, and the fifth activation layer to obtain the first related weights;

[0046] The second weight acquisition network includes: a global feature extraction unit, a deep feature extraction unit, and a sixth activation layer;

[0047] The modal interaction features are respectively input to the global feature extraction unit and the deep feature extraction unit to obtain global features and deep features respectively;

[0048] The fused global features and deep features are then input into the sixth activation layer to obtain the second related weights.

[0049] Preferably, the global feature extraction unit includes: a global average pooling layer, a fourth depthwise separable convolutional layer, an eighth normalization layer, a seventh activation layer, a fifth depthwise separable convolutional layer, and a ninth normalization layer;

[0050] The modal interaction features are sequentially input into the global average pooling layer, the fourth depthwise separable convolutional layer, the eighth normalization layer, the seventh activation layer, the fifth depthwise separable convolutional layer, and the ninth normalization layer to obtain the global features;

[0051] The deep feature extraction unit includes: a sixth depthwise separable convolutional layer, an eighth activation layer, a seventh depthwise separable convolutional layer, an eighth depthwise separable convolutional layer, a ninth depthwise separable convolutional layer, and a tenth normalization layer;

[0052] The modal interaction features are sequentially input into the sixth depthwise separable convolutional layer and the eighth activation layer to obtain the first extracted features;

[0053] The first extracted feature is sequentially input into the seventh depth separable convolutional layer and the eighth depth separable convolutional layer to obtain the second extracted feature;

[0054] The first extracted feature and the second extracted feature are fused and then sequentially input into the ninth depth separable convolutional layer and the tenth normalized layer to obtain the deep features.

[0055] Preferably, the classification output network includes: a third convolutional layer, an eleventh normalization layer, a ninth activation layer, a fourth convolutional layer, a twelfth normalization layer, a tenth activation layer, a fifth convolutional layer, and an output layer;

[0056] The multimodal remote sensing features are sequentially input into the third convolutional layer, the eleventh normalized layer, the ninth activation layer, the fourth convolutional layer, the twelfth normalized layer, and the tenth activation layer to obtain segmentation features;

[0057] The segmentation features are sequentially input into the fifth convolutional layer and the output layer to obtain the classification probability map of the target to be segmented;

[0058] The segmentation result is obtained based on the classification probability map.

[0059] A remote sensing image segmentation system based on hierarchical multimodal feature fusion includes: an image acquisition module, an aggregated feature acquisition module, a first feature acquisition module, a second feature acquisition module, and a result output module;

[0060] The image acquisition module is used to acquire hyperspectral images and synthetic aperture radar images of the target to be segmented;

[0061] The aggregated feature acquisition module is used to obtain a first aggregated feature by inputting the hyperspectral image into the HSI feature extraction network; and to obtain a second aggregated feature by inputting the synthetic aperture radar image into the SAR feature extraction network.

[0062] The first feature acquisition module is used to input the first aggregated feature into the first weight acquisition network to obtain the first relevant weight; and to fuse the second aggregated feature with the first relevant weight to obtain the modal interaction feature.

[0063] The second feature acquisition module is used to input the modal interaction features into the second weight acquisition network to obtain the second relevant weights; and to obtain multimodal remote sensing features by fusing the first aggregated features and the second relevant weights.

[0064] The result output module is used to input the multimodal remote sensing features into the classification output network to obtain the segmentation result of the target to be segmented.

[0065] As can be seen from the above technical solution, compared with the prior art, this invention discloses a remote sensing image segmentation method and system based on hierarchical multimodal feature fusion. It utilizes CNN branches to extract local texture features from hyperspectral and synthetic aperture radar (SAR) images, while employing Transformer branches to model global features of the images. This fully exploits the complementary information in the hyperspectral and SAR images. The features of the hyperspectral and SAR images are weighted and calculated through layer-by-layer optimization, combined with modal interaction mechanisms and residual connections to ensure the effectiveness and stability of cross-modal information fusion. Compared with traditional simple stitching or weighted fusion methods, this method can more precisely filter, interact with, and fuse different modal information, thereby improving the accuracy and generalization ability of remote sensing image segmentation. This invention not only improves the accuracy of remote sensing image segmentation but also enhances the robustness of the model, providing a new solution for multi-source remote sensing image fusion and intelligent analysis. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0067] Figure 1 The flowchart of a remote sensing image segmentation method based on hierarchical multimodal feature fusion provided by the present invention is shown.

[0068] Figure 2This is a schematic diagram of the HSI feature extraction network structure provided by the present invention.

[0069] Figure 3 This is a schematic diagram of the hybrid enhanced Transformer unit structure provided by the present invention.

[0070] Figure 4 This is a schematic diagram of the SAR feature extraction network structure provided by the present invention.

[0071] Figure 5 A schematic diagram of the second weight acquisition network structure provided by the present invention.

[0072] Figure 6 This is a schematic diagram of a remote sensing image segmentation system based on hierarchical multimodal feature fusion provided by the present invention. Detailed Implementation

[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0074] Example 1

[0075] like Figure 1 As shown, this embodiment of the invention discloses a remote sensing image segmentation method based on hierarchical multimodal feature fusion, comprising:

[0076] Acquire hyperspectral and synthetic aperture radar images of the target to be segmented;

[0077] The first aggregated feature is obtained by inputting hyperspectral imagery into the HSI feature extraction network;

[0078] The second aggregated feature is obtained by inputting synthetic aperture radar imagery into the SAR feature extraction network.

[0079] The first aggregated feature is input into the first weight acquisition network to obtain the first relevant weight;

[0080] Based on the fusion of the second aggregation feature and the first relevance weight, the modal interaction feature is obtained;

[0081] The second relevant weights are obtained by inputting modal interaction features into the second weight acquisition network.

[0082] Multimodal remote sensing features are obtained by fusing the first aggregation feature and the second correlation weight.

[0083] The segmentation result of the target to be segmented is obtained by inputting multimodal remote sensing features into the classification output network.

[0084] Example 2

[0085] This invention discloses a remote sensing image segmentation method based on hierarchical multimodal feature fusion, comprising:

[0086] Acquire hyperspectral and synthetic aperture radar images of the target to be segmented.

[0087] Preferably, hyperspectral imagery possesses rich spectral information (typically containing dozens to hundreds of bands) and exhibits strong global correlation between different bands. Synthetic aperture radar (SAR) imagery, unlike hyperspectral imagery, primarily contains single-band or multi-polarization data, lacking rich spectral information but possessing strong spatial texture features. By fully leveraging the complementary characteristics of hyperspectral imagery (HSI) and synthetic aperture radar (SAR) imagery, interactive fusion of cross-modal information can be achieved, thereby improving segmentation accuracy and generalization capability.

[0088] The first aggregated feature is obtained by inputting hyperspectral images into the HSI feature extraction network.

[0089] Preferred, such as Figure 2 As shown, the HSI feature extraction network includes: a location embedding layer, a first hybrid enhancement Transformer unit, a first downsampling layer, a second hybrid enhancement Transformer unit, a second downsampling layer, a third hybrid enhancement Transformer unit, and a first multi-scale aggregation unit;

[0090] Hyperspectral images are input into the location embedding layer to obtain embedding features;

[0091] The embedded features are sequentially input into the first hybrid enhancement Transformer unit and the first downsampling layer to obtain the first scale features;

[0092] The first-scale features are sequentially input into the second hybrid enhancement Transformer unit and the second downsampling layer to obtain the second-scale features;

[0093] The second-scale features are input into the third hybrid enhancement Transformer unit to obtain the third-scale features;

[0094] The first-scale feature, the second-scale feature, and the third-scale feature are input into the first multi-scale aggregation unit to obtain the first aggregated feature.

[0095] Preferably, Transformer performs well in modeling long-range dependencies. This invention uses a hybrid enhanced Transformer branch to extract features from hyperspectral images.

[0096] Preferred, such as Figure 3 As shown, the first hybrid enhancement Transformer unit, the second hybrid enhancement Transformer unit, and the third hybrid enhancement Transformer unit have the same structure, all including:

[0097] The system consists of a first normalized layer, a pooling layer, a first depthwise separable convolutional layer, a second normalized layer, a first activation layer, a third normalized layer, and a multilayer perceptron.

[0098] The first input feature is fed into the first normalization layer to obtain the first processed feature;

[0099] The first processed feature is input into the pooling layer to obtain the second processed feature;

[0100] After fusing the first processing feature and the second processing feature, the first fused feature is obtained;

[0101] The first fusion feature is sequentially input into the first depthwise separable convolutional layer, the second normalized layer, and the first activation layer to obtain the third processed feature;

[0102] The first fusion feature and the third processing feature are fused together to obtain the second fusion feature;

[0103] The second fusion feature is sequentially input into the third normalization layer and the multilayer perceptron to obtain the fourth processed feature;

[0104] The first output feature is obtained by fusing the second fusion feature and the fourth processing feature.

[0105] Preferably, the following layers are used: a normalization layer to normalize embedded features and stabilize the training process; a pooling layer to reduce data dimensionality and computational complexity; a residual link to combine input features and transformed features while preserving the original information; a depthwise separable convolutional layer to extract local multi-scale features; a first activation layer using the ReLU activation function to introduce non-linearity and enhance the model's expressive power; and a multilayer perceptron (MLP) to further extract and fuse features.

[0106] Preferably, the hyperspectral image is input into the location embedding layer to obtain the embedding feature X. HSI ;

[0107] The first input feature of the first hybrid enhancement Transformer unit is the embedded feature X. HSI Its first output feature is Z, and features are further extracted through the first downsampling layer to obtain the first scale feature φ1: φ1 = Downsample(Z), where Downsample() represents the downsampling operation;

[0108] The first input feature of the second hybrid enhancement Transformer unit is the first scale feature φ1, and its first output feature is V. The second scale feature φ2 is obtained by further extracting features through the second downsampling layer: φ2 = Downsample(V);

[0109] The first input feature of the third hybrid enhancement Transformer unit is the second scale feature φ2, and its first output feature is the third scale feature φ3.

[0110] Preferably, the first scale feature φ1, the second scale feature φ2, and the third scale feature φ3 are input into the first multi-scale aggregation unit, and trilinear interpolation is performed on φ1, φ2, and φ3 respectively to unify them into the same spatial size, thereby obtaining the first alignment feature. Second alignment feature and third alignment feature

[0111] The aligned features are then processed using pointwise convolution. and The fusion is performed along the channel dimension to obtain the first aggregated feature φ after fusion. HSI :

[0112]

[0113] Among them, f agg () represents the channel dimension aggregation function, which is usually implemented by pointwise convolution.

[0114] Preferably, the present invention avoids directly performing fully connected interactions on all scale features, but instead aligns and then fuses them, which effectively reduces computational overhead, computational complexity, and improves computational efficiency.

[0115] The second aggregated feature is obtained by inputting synthetic aperture radar imagery into the SAR feature extraction network.

[0116] Preferred, such as Figure 4 As shown, the SAR feature extraction network includes: an initial feature extraction unit, a first multi-layer feature extraction unit, a second multi-layer feature extraction unit, and a second multi-scale aggregation unit;

[0117] The synthetic aperture radar image is input into the initial feature extraction unit to obtain the fourth scale feature φ4;

[0118] The fourth-scale feature is input into the first multi-layer feature extraction unit to obtain the fifth-scale feature φ5;

[0119] The fifth-scale feature is input into the second multi-layer feature extraction unit to obtain the sixth-scale feature φ6.

[0120] The fourth-scale feature, the fifth-scale feature, and the sixth-scale feature are input into the second multi-scale aggregation unit to obtain the second aggregation feature.

[0121] Preferably, synthetic aperture radar (SAR) images differ from hyperspectral images, primarily containing single-band or multi-polarization data and lacking rich spectral information, but possessing strong spatial texture features. CNN-based models have significant advantages in extracting local spatial structure and edge information; therefore, this invention employs CNN branches for feature extraction from SAR images.

[0122] Preferably, the initial feature extraction unit includes: a first convolutional layer, a fourth normalization layer, a second activation layer, and a third downsampling layer;

[0123] The synthetic aperture radar image is input into the first convolutional layer to obtain the first intermediate feature;

[0124] The first intermediate feature is sequentially input into the fourth normalization layer and the second activation layer to obtain the second intermediate feature;

[0125] The first and second intermediate features are fused and then input into the third downsampling layer to obtain the fourth scale feature φ4.

[0126] Preferably, after processing by the first convolutional layer, feature extraction is performed using a standard CNN module, including: a fourth normalization layer to improve stability; a second activation layer using the ReLU activation function to provide non-linearity; and residual links to avoid gradient vanishing.

[0127] Preferably, the first multi-layer feature extraction unit and the second multi-layer feature extraction unit have the same structure, both including: a second convolutional layer, a fifth normalization layer, a third activation layer and a fourth downsampling layer;

[0128] The second output feature is sequentially input into the second convolutional layer, the fifth normalized layer, and the third activation layer to obtain the third intermediate feature;

[0129] The third intermediate feature is fused with the second input feature and then fed into the fourth downsampling layer to obtain the second output feature;

[0130] The second input feature is either the fourth-scale feature or the fifth-scale feature, and the second output feature is either the fifth-scale feature or the sixth-scale feature.

[0131] Preferably, the second input feature of the first multi-layer feature extraction unit is the fourth scale feature φ4, and its second output feature is the fifth scale feature φ5; the second input feature of the second multi-layer feature extraction unit is the fifth scale feature φ5, and its second output feature is the sixth scale feature φ6.

[0132] Preferably, the fourth scale feature φ4, the fifth scale feature φ5, and the sixth scale feature φ6 are input into the second multi-scale aggregation unit, and trilinear interpolation is performed on φ4, φ5, and φ6 respectively to unify them to the same spatial size, thus obtaining the corresponding fourth alignment feature. Fifth alignment feature and the sixth alignment feature

[0133] The aligned features φ4, φ5, and φ6 are fused along the channel dimension using pointwise convolution to obtain the fused second aggregated feature φ. SAR :

[0134]

[0135] Preferably, the present invention constructs suitable feature extraction models for hyperspectral images and synthetic aperture radar images respectively through a Transformer+CNN dual-branch architecture: the HSI feature extraction network adopts a hybrid enhanced Transformer to capture global spectral features; the SAR feature extraction network adopts a CNN to extract local texture information; and the multi-scale aggregation unit further fuses features to improve segmentation performance.

[0136] The first aggregated feature is input into the first weight acquisition network to obtain the first relevant weight.

[0137] Preferably, the first weight acquisition network includes: a second depthwise separable convolutional layer, a sixth normalization layer, a fourth activation layer, a third depthwise separable convolutional layer, a seventh normalization layer, and a fifth activation layer;

[0138] The first aggregated feature is sequentially input into the second depthwise separable convolutional layer, the sixth normalization layer, the fourth activation layer, the third depthwise separable convolutional layer, the seventh normalization layer, and the fifth activation layer to obtain the first relevant weight.

[0139] Preferably, in this embodiment, the fourth activation layer uses the ReLU activation function, and the fifth activation layer uses the Sigmoid activation function. First aggregated feature φ HSI Input to a second-depth separable convolutional layer extracts effective features while reducing computational cost.

[0140] Preferably, the first relevant weight W HSI for:

[0141] W HSI =Sigmoid(BN(DWConv(ReLU(BN(DWConv(φ HSI ))))));

[0142] Where Sigmoid represents the activation operation of the fifth activation layer, BN represents the normalization operation, DWConv represents the depthwise separable convolution operation, and ReLU represents the activation operation of the fourth activation layer.

[0143] Modal interaction features are obtained by fusing the second aggregation feature with the first correlation weight.

[0144] Preferably, the second aggregation feature φ SAR The first relevant weight W HSI Preliminary modal interaction is achieved through element-wise multiplication, resulting in modal interaction features.

[0145]

[0146] Here, ⊙ represents the matrix multiplication operation.

[0147] The second relevant weights are obtained by inputting the modal interaction features into the second weight acquisition network.

[0148] Preferred, such as Figure 5 As shown, the second weight acquisition network includes: a global feature extraction unit, a deep feature extraction unit, and a sixth activation layer;

[0149] Modal interaction features are input into the global feature extraction unit and the deep feature extraction unit, respectively, to obtain global features and deep features.

[0150] The second relevant weights are obtained by fusing global and deep features and inputting them into the sixth activation layer.

[0151] Preferably, the global feature extraction unit includes: a global average pooling layer, a fourth depthwise separable convolutional layer, an eighth normalization layer, a seventh activation layer, a fifth depthwise separable convolutional layer, and a ninth normalization layer;

[0152] Modal interaction features are sequentially input into a global average pooling layer, a fourth depthwise separable convolutional layer, an eighth normalization layer, a seventh activation layer, a fifth depthwise separable convolutional layer, and a ninth normalization layer to obtain global features.

[0153] Preferably, in this embodiment, the sixth activation layer uses the Sigmoid activation function, and the seventh activation layer uses the ReLU activation function.

[0154] Preferred, global feature A c for:

[0155]

[0156] GPA stands for Global Average Pooling.

[0157] Preferably, the deep feature extraction unit includes: a sixth depthwise separable convolutional layer, an eighth activation layer, a seventh depthwise separable convolutional layer, an eighth depthwise separable convolutional layer, a ninth depthwise separable convolutional layer, and a tenth normalization layer;

[0158] Modal interaction features are sequentially input into the sixth deep separable convolutional layer and the eighth activation layer to obtain the first extracted features;

[0159] The first extracted features are sequentially input into the seventh and eighth depth separable convolutional layers to obtain the second extracted features;

[0160] The first and second extracted features are fused and then sequentially input into the ninth deep separable convolutional layer and the tenth normalized layer to obtain deep features.

[0161] Preferably, in this embodiment, the eighth activation layer adopts the ReLU activation function and utilizes depthwise separable convolution and residual connections to enhance the cross-modal feature representation capability.

[0162] Preferred, deep feature A s for:

[0163]

[0164] Among them, A t Indicates the first extracted feature. This indicates an addition / merging operation.

[0165] Preferably, based on global feature A c and deep features A s After fusion, the input is fed into the sixth activation layer to obtain the second relevant weight W. SAR :

[0166] W SAR =Sigmoid(A c +A s );

[0167] Here, Sigmoid represents the Sigmoid activation function, which is used to normalize weights.

[0168] Multimodal remote sensing features are obtained by fusing the first aggregated feature and the second correlation weight.

[0169] Preferred multimodal remote sensing features φ +

[0170] φ + =W SAR ⊙φ HSI .

[0171] Preferably, this invention fully integrates the complementary information of hyperspectral imagery and synthetic aperture radar imagery through layer-by-layer feature extraction, modal interaction, weight calculation, and residual connection, achieving a more comprehensive and robust feature representation. The resulting multimodal remote sensing features can be used for remote sensing image segmentation, improving segmentation accuracy and reducing information loss caused by single modalities.

[0172] The segmentation result of the target to be segmented is obtained by inputting multimodal remote sensing features into the classification output network.

[0173] Preferably, the classification output network includes: a third convolutional layer, an eleventh normalization layer, a ninth activation layer, a fourth convolutional layer, a twelfth normalization layer, a tenth activation layer, a fifth convolutional layer, and an output layer;

[0174] Multimodal remote sensing features are sequentially input into the third convolutional layer, the eleventh normalized layer, the ninth activation layer, the fourth convolutional layer, the twelfth normalized layer, and the tenth activation layer to obtain segmentation features;

[0175] The segmentation features are sequentially input into the fifth convolutional layer and the output layer to obtain the classification probability map of the target to be segmented;

[0176] The segmentation result is obtained based on the classification probability map.

[0177] Preferably, in this embodiment, the third and fourth convolutional layers are both 3×3 standard convolutional layers, and the ninth and tenth activation layers are both ReLU activation functions. The multimodal remote sensing feature φ + After a series of convolution-normalization-activation operations to compress channels and extract segmentation information, the segmentation feature ψ is obtained:

[0178] ψ=ReLU(BN(Conv 3×3 (ReLU(BN(Conv 3×3 (φ + ))))));

[0179] Among them, Conv 3×3 : Represents a standard 3×3 convolutional layer, used to capture local features; BN stands for Batch Normalization, used to accelerate training and stabilize the model; ReLU: Represents the ReLU activation function, used for non-linear activation, enhancing feature representation capabilities.

[0180] Preferably, in this embodiment, the fifth convolutional layer is a 1×1 standard convolutional layer, and the output layer uses the Softmax activation function; the segmentation features are mapped to categories through the 1×1 convolutional layer and the Softmax activation function to obtain the classification probability map P of the target to be segmented.

[0181] P = Softmax(Conv) 1×1 (ψ));

[0182] Among them, Conv 1×1 It is responsible for channel mapping, which converts features into C-dimensional class probabilities; the Softmax normalization operation converts the result into a pixel-level class probability distribution. The final P is a pixel-level classification probability map with H×W×C dimensions, representing the probability that each pixel belongs to C classes, where H and W represent the height and width, respectively.

[0183] Preferably, the segmentation result of the target to be segmented is obtained based on the classification probability map P, that is, the pixel-level semantic segmentation map, where each pixel is assigned a clear category label, thus achieving fine segmentation of the input image.

[0184] Preferably, the remote sensing image segmentation model is composed of the HSI feature extraction network, the SAR feature extraction network, the first weight acquisition network, the second weight acquisition network, and the classification output network. The remote sensing image segmentation model is trained based on the total loss function to obtain a trained remote sensing image segmentation model, which is used for subsequent remote sensing image target segmentation.

[0185] Preferably, the total loss function is L total :

[0186] L total =λ1L Focal +λ2L Dice ;

[0187] Where λ1 and λ2 represent the loss weights, both set to 0.5 in this embodiment to balance the class discrimination ability and the boundary preservation ability. Focal L represents the class imbalance loss. Dice This indicates a loss that occurs while maintaining the boundary.

[0188] Preferably, in remote sensing image segmentation, target areas are often small, and class distribution is imbalanced. To address this, this invention introduces Focal Loss to reduce focus on easily classified samples and enhance learning of difficult-to-classify samples. The Focal Loss L... Focal Specifically:

[0189] L Focal =-α(1-P t ) γ log(P t );

[0190] Where α represents the category balancing factor, preventing minor categories from being ignored, and P... t This represents the predicted probability corresponding to the true category, and γ represents the weight that controls the difficulty of the sample, which is usually taken as 2.

[0191] Preferably, to improve the accuracy of the target boundary and reduce ambiguous areas, this invention introduces a boundary preservation loss (Dice Loss) to optimize the overlap between the predicted region and the true target region. The boundary preservation loss L... Dice Specifically:

[0192]

[0193] Among them, P i and G i These represent the predicted value and the actual value, respectively, and N represents the category. This formula calculates the Dice coefficient, which is used to measure the similarity of target regions.

[0194] Preferably, in order to ensure the stable convergence of the remote sensing image segmentation model and obtain optimal segmentation performance, the present invention adopts the following optimization strategy:

[0195] The AdamW optimizer is used for training to reduce the impact of L2 regularization and improve generalization ability.

[0196]

[0197] Where, θ (t) Let θ represent the model parameters at the t-th iteration. (t+1) This represents the updated model parameters, i.e., the parameter values ​​used for forward and backward propagation in the (t+1)th iteration, where η is the learning rate and m is the number of parameters. t and v t These are the first-order moment estimate and the second-order moment estimate, respectively, where ∈ denotes the minimum value (1e). -8 To prevent the denominator from being zero, λ represents the weight decay factor to prevent overfitting.

[0198] Cosine Annealing (LR) learning rate scheduling is employed to ensure rapid initial convergence and stable optimization in later stages.

[0199]

[0200] Where, η t η represents the learning rate in the current training epoch or iteration; π is a mathematical constant, approximately 3.14159, used to control the periodicity of the cosine function, ensuring a smooth decrease in the learning rate; η represents the learning rate. max and η min T represents the maximum learning rate and the minimum learning rate, respectively. cur T represents the current training round number. max This indicates the total number of training rounds.

[0201] Preferably, the present invention uses Focal Loss+Dice Loss to optimize the segmentation model. At the same time, it combines AdamW+cosine annealing learning rate training strategy to improve the segmentation accuracy and stability of the model, ensuring that the obtained segmentation results have clear boundaries and complete details.

[0202] Example 3

[0203] like Figure 6 As shown, a remote sensing image segmentation system based on hierarchical multimodal feature fusion includes: an image acquisition module, an aggregated feature acquisition module, a first feature acquisition module, a second feature acquisition module, and a result output module;

[0204] The image acquisition module is used to acquire hyperspectral images and synthetic aperture radar images of the target to be segmented;

[0205] The aggregated feature acquisition module is used to obtain the first aggregated feature by inputting hyperspectral imagery into the HSI feature extraction network; and to obtain the second aggregated feature by inputting synthetic aperture radar imagery into the SAR feature extraction network.

[0206] The first feature acquisition module is used to input the first aggregated feature into the first weight acquisition network to obtain the first relevant weight; and to obtain the modal interaction feature by fusing the second aggregated feature with the first relevant weight.

[0207] The second feature acquisition module is used to input modal interaction features into the second weight acquisition network to obtain the second correlation weights; and to obtain multimodal remote sensing features by fusing the first aggregated features and the second correlation weights.

[0208] The results output module is used to input multimodal remote sensing features into the classification output network to obtain the segmentation results of the target to be segmented.

[0209] Preferably, in this embodiment, the functional implementation methods of each functional module correspond one-to-one with the above-mentioned methods, and will not be described in detail here.

[0210] Example 4

[0211] Based on the same inventive concept, the present invention also provides a computer device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0212] Memory, used to store computer programs;

[0213] When the processor executes a program stored in memory, it is able to implement a remote sensing image segmentation method based on hierarchical multimodal feature fusion, as in Embodiment 1 or 2.

[0214] The electronic device may include a processor, a communications interface, a memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions in the memory to execute a remote sensing image segmentation method based on hierarchical multimodal feature fusion as described in Embodiment 1 or 2.

[0215] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention.

[0216] The aforementioned storage media include: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media that can store program code.

[0217] As described above, this invention discloses a remote sensing image segmentation method and system based on hierarchical multimodal feature fusion. It utilizes CNN branches to extract local texture features from hyperspectral and synthetic aperture radar (SAR) images, while employing Transformer branches to model global features of the images. This fully leverages the complementary information in the hyperspectral and SAR images. The features of the hyperspectral and SAR images are weighted through layer-by-layer optimization, and combined with modal interaction mechanisms and residual connections, the effectiveness and stability of cross-modal information fusion are ensured. Compared to traditional simple stitching or weighted fusion methods, this approach allows for more precise filtering, interaction, and fusion of different modal information, thereby improving the accuracy and generalization ability of remote sensing image segmentation. This invention not only improves the accuracy of remote sensing image segmentation but also enhances the robustness of the model, providing a new solution for multi-source remote sensing image fusion and intelligent analysis.

[0218] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0219] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method based on hierarchical multimodal feature fusion, characterized in that, include: Acquire hyperspectral and synthetic aperture radar images of the target to be segmented; The first aggregated feature is obtained by inputting the hyperspectral image into the HSI feature extraction network; The HSI feature extraction network includes: a location embedding layer, a first hybrid enhancement Transformer unit, a first downsampling layer, a second hybrid enhancement Transformer unit, a second downsampling layer, a third hybrid enhancement Transformer unit, and a first multi-scale aggregation unit; The hyperspectral image is input into the location embedding layer to obtain the embedding features; The embedded features are sequentially input into the first hybrid enhancement Transformer unit and the first downsampling layer to obtain the first scale features; The first scale feature is sequentially input into the second hybrid enhancement Transformer unit and the second downsampling layer to obtain the second scale feature; The second-scale feature is input into the third hybrid enhancement Transformer unit to obtain the third-scale feature; The first scale feature, the second scale feature, and the third scale feature are input into the first multi-scale aggregation unit to obtain the first aggregated feature; The synthetic aperture radar image is input into the SAR feature extraction network to obtain the second aggregated feature; The SAR feature extraction network includes: an initial feature extraction unit, a first multi-layer feature extraction unit, a second multi-layer feature extraction unit, and a second multi-scale aggregation unit; The synthetic aperture radar image is input to the initial feature extraction unit to obtain fourth-scale features; The fourth-scale feature is input into the first multi-layer feature extraction unit to obtain the fifth-scale feature; The fifth-scale feature is input into the second multi-layer feature extraction unit to obtain the sixth-scale feature; The fourth-scale feature, the fifth-scale feature, and the sixth-scale feature are input into the second multi-scale aggregation unit to obtain the second aggregated feature; Based on the first aggregated feature, the first weight acquisition network is input to obtain the first relevant weight; Based on the fusion of the second aggregation feature and the first relevant weight, modal interaction features are obtained; The modal interaction features are input into the second weight acquisition network to obtain the second relevant weights; Based on the fusion of the first aggregated features and the second related weights, multimodal remote sensing features are obtained; The multimodal remote sensing features are input into the classification output network to obtain the segmentation result of the target to be segmented.

2. The remote sensing image segmentation method based on hierarchical multimodal feature fusion according to claim 1, characterized in that, The first hybrid enhancement Transformer unit, the second hybrid enhancement Transformer unit, and the third hybrid enhancement Transformer unit have the same structure, each including: The system consists of a first normalized layer, a pooling layer, a first depthwise separable convolutional layer, a second normalized layer, a first activation layer, a third normalized layer, and a multilayer perceptron. The first input feature is input into the first normalization layer to obtain the first processed feature; The first processing feature is input to the pooling layer to obtain the second processing feature; The first processing feature and the second processing feature are fused together to obtain the first fused feature; The first fusion feature is sequentially input into the first depthwise separable convolutional layer, the second normalization layer, and the first activation layer to obtain the third processed feature; The first fusion feature and the third processing feature are fused to obtain the second fusion feature; The second fusion feature is sequentially input into the third normalization layer and the multilayer perceptron to obtain the fourth processing feature; After the second fusion feature and the fourth processing feature are fused, the first output feature is obtained.

3. The remote sensing image segmentation method based on hierarchical multimodal feature fusion according to claim 1, characterized in that, The initial feature extraction unit includes: a first convolutional layer, a fourth normalization layer, a second activation layer, and a third downsampling layer; The synthetic aperture radar image is input into the first convolutional layer to obtain the first intermediate feature; The first intermediate feature is sequentially input into the fourth normalization layer and the second activation layer to obtain the second intermediate feature; The first intermediate feature and the second intermediate feature are fused and then input into the third downsampling layer to obtain the fourth scale feature.

4. The remote sensing image segmentation method based on hierarchical multimodal feature fusion according to claim 1, characterized in that, The first multi-layer feature extraction unit and the second multi-layer feature extraction unit have the same structure, both including: a second convolutional layer, a fifth normalization layer, a third activation layer and a fourth downsampling layer; The second input feature is sequentially input into the second convolutional layer, the fifth normalization layer, and the third activation layer to obtain the third intermediate feature; The third intermediate feature and the second input feature are fused and then input into the fourth downsampling layer to obtain the second output feature; The second input feature is either the fourth scale feature or the fifth scale feature, and the second output feature corresponds to either the fifth scale feature or the sixth scale feature.

5. The remote sensing image segmentation method based on hierarchical multimodal feature fusion according to claim 1, characterized in that, The first weight acquisition network includes: a second depthwise separable convolutional layer, a sixth normalization layer, a fourth activation layer, a third depthwise separable convolutional layer, a seventh normalization layer, and a fifth activation layer; The first aggregated feature is sequentially input into the second depthwise separable convolutional layer, the sixth normalized layer, the fourth activation layer, the third depthwise separable convolutional layer, the seventh normalized layer, and the fifth activation layer to obtain the first related weights; The second weight acquisition network includes: a global feature extraction unit, a deep feature extraction unit, and a sixth activation layer; The modal interaction features are respectively input to the global feature extraction unit and the deep feature extraction unit to obtain global features and deep features respectively; The fused global features and deep features are then input into the sixth activation layer to obtain the second related weights.

6. The remote sensing image segmentation method based on hierarchical multimodal feature fusion according to claim 5, characterized in that, The global feature extraction unit includes: a global average pooling layer, a fourth depthwise separable convolutional layer, an eighth normalization layer, a seventh activation layer, a fifth depthwise separable convolutional layer, and a ninth normalization layer; The modal interaction features are sequentially input into the global average pooling layer, the fourth depthwise separable convolutional layer, the eighth normalization layer, the seventh activation layer, the fifth depthwise separable convolutional layer, and the ninth normalization layer to obtain the global features; The deep feature extraction unit includes: a sixth depthwise separable convolutional layer, an eighth activation layer, a seventh depthwise separable convolutional layer, an eighth depthwise separable convolutional layer, a ninth depthwise separable convolutional layer, and a tenth normalization layer; The modal interaction features are sequentially input into the sixth depthwise separable convolutional layer and the eighth activation layer to obtain the first extracted features; The first extracted feature is sequentially input into the seventh depth separable convolutional layer and the eighth depth separable convolutional layer to obtain the second extracted feature; The first extracted feature and the second extracted feature are fused and then sequentially input into the ninth depth separable convolutional layer and the tenth normalized layer to obtain the deep features.

7. The remote sensing image segmentation method based on hierarchical multimodal feature fusion according to claim 1, characterized in that, The classification output network includes: a third convolutional layer, an eleventh normalization layer, a ninth activation layer, a fourth convolutional layer, a twelfth normalization layer, a tenth activation layer, a fifth convolutional layer, and an output layer; The multimodal remote sensing features are sequentially input into the third convolutional layer, the eleventh normalized layer, the ninth activation layer, the fourth convolutional layer, the twelfth normalized layer, and the tenth activation layer to obtain segmentation features; The segmentation features are sequentially input into the fifth convolutional layer and the output layer to obtain the classification probability map of the target to be segmented; The segmentation result is obtained based on the classification probability map.

8. A remote sensing image segmentation system based on hierarchical multimodal feature fusion, applied to the remote sensing image segmentation method based on hierarchical multimodal feature fusion as described in any one of claims 1-7, characterized in that, It includes: an image acquisition module, an aggregated feature acquisition module, a first feature acquisition module, a second feature acquisition module, and a result output module; The image acquisition module is used to acquire hyperspectral images and synthetic aperture radar images of the target to be segmented; The aggregated feature acquisition module is used to obtain a first aggregated feature by inputting the hyperspectral image into the HSI feature extraction network; and to obtain a second aggregated feature by inputting the synthetic aperture radar image into the SAR feature extraction network. The first feature acquisition module is used to input the first aggregated feature into the first weight acquisition network to obtain the first relevant weight; and to fuse the second aggregated feature with the first relevant weight to obtain the modal interaction feature. The second feature acquisition module is used to input the modal interaction features into the second weight acquisition network to obtain the second relevant weights; Based on the fusion of the first aggregated features and the second related weights, multimodal remote sensing features are obtained; The result output module is used to input the multimodal remote sensing features into the classification output network to obtain the segmentation result of the target to be segmented.

Citation Information

Patent Citations

  • Multi-source remote sensing image classification method based on multistage feature fusion

    CN117876890A

  • Multi-modal brain tumor image segmentation method based on self-supervised learning

    WO2024108522A1