A multi-modal remote sensing image ground feature classification method based on modal information constraint
By adaptively learning the fusion weights of modality sharing and differential features through a multimodal classification model, the problem of loss of modality-specific information in multimodal remote sensing image land cover classification is solved, achieving higher classification accuracy and better scene adaptability.
Patent Information
- Application Number
- CN202411431849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-10-14
AI Technical Summary
Existing multimodal remote sensing image land cover classification methods struggle to balance the contributions of modality sharing and modality-specific information under complex land cover distribution or noisy conditions, leading to the loss of modality-specific information and affecting classification accuracy.
A multimodal classification model is adopted, including an OPT encoder, a SAR encoder, an MDF decoder, an MSIE-OPT decoder, and an MSIE-SAR decoder. By adaptively learning the fusion weights of modality sharing and modality difference features, modality-specific information is preserved, and modality difference features are used to assist in classification. A loss function is constructed for supervised training.
It enhances the representational ability of fused features, improves the accuracy of ground feature classification, effectively alleviates the recognition performance bottleneck in complex scenarios, and improves the accuracy of multimodal remote sensing image interpretation.
Smart Images

Figure CN119399522B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of remote sensing image processing, and particularly relates to a multi-modal remote sensing image feature classification method based on modal information constraint. BACKGROUND
[0002] Remote sensing image feature classification is a basic task in earth observation, and its main goal is to assign each pixel in a remote sensing image to a predefined feature class. This technology plays an important role in urban planning, environmental protection, precision agriculture, etc. Optical images usually contain rich spectral and texture information, but are susceptible to weather and lighting conditions, making it difficult to accurately classify features. In contrast, synthetic aperture radar (SAR) images can provide structural and scattering characteristics, enhancing the feature diversity of features. Recently, many studies have focused on multi-modal remote sensing image feature classification using optical and SAR data to overcome performance bottlenecks.
[0003] In multi-modal remote sensing image feature classification, it is crucial to integrate the complementary information of each modality. Among them, exploring the relationship between modal shared and modal specific information helps to integrate the complementary information of modalities, thereby achieving comprehensive interpretation. In earth observation, both optical and SAR images contain modal shared and modal specific information. Among them, modal shared information represents the common characteristics of the two modalities, which helps to bridge the gap between them, such as the shape and location of features; modal specific information represents unique attributes in each modality that are essential for accurately describing features, such as spectral and scattering characteristics. Current multi-modal remote sensing image fusion methods are mainly based on deep feature-level fusion. These methods usually use fusion operators such as concatenation, cross-attention, etc. for optical and SAR features, and obtain fused features by learning fusion weights, thereby indirectly regulating the contribution of modal shared and modal specific information, and achieving good classification results.
[0004] However, in challenging scenarios such as complex feature distribution in optical images or speckle noise in SAR images, the performance of these existing multi-modal remote sensing image fusion methods may have performance bottlenecks. Both optical and SAR features contain modal shared and modal specific information, and existing methods directly learn the fusion weights of optical and SAR features, making it difficult to balance the contributions of modal shared and modal specific information, thereby leading to the loss of modal specific information. In order to improve the classification performance of multi-modal remote sensing image feature classification methods in challenging scenarios such as complex feature distribution or noise conditions, the above key problems need to be solved. SUMMARY
[0005] To solve the problem of modal specific information loss in the existing multi-modal remote sensing image ground object classification method, the application provides a multi-modal remote sensing image ground object classification method based on modal information constraint, which enhances the representation ability of fused features and has higher ground object classification accuracy.
[0006] A multi-modal remote sensing image ground object classification method based on modal information constraint, adopts a multi-modal classification model to classify the ground objects of the scene to be measured, wherein the multi-modal classification model comprises an OPT encoder, a SAR encoder, an MDF decoder, an MSIE-OPT decoder and an MSIE-SAR decoder, and the training method of the multi-modal classification model is:
[0007] The OPT encoder is used to obtain an optical image X opt Five different scale optical multi-modal features
[0008]
[0009] The SAR encoder is used to obtain a SAR image X sar Five different scale SAR multi-modal features Wherein, the optical image X opt and the SAR image X sar are taken for the same scene;
[0010] The multi-modal features and are projected to the same dimension to obtain corresponding multi-modal projection features and
[0011] The MDF decoder is used to fuse and to obtain a modal fusion classification result Meanwhile, the MDF decoder is used to obtain the difference between and to obtain a modal difference feature
[0012] The MSIE-OPT decoder is used to fuse the optical multi-modal projection feature and the modal difference feature to obtain an optical fusion classification result
[0013] The MSIE-SAR decoder is used to fuse the SAR multi-modal projection feature and the modal difference feature to obtain a SAR fusion classification result
[0014] According to the modal fusion classification result Optical fusion classification results SAR fusion classification results Construct the loss function L;
[0015] Supervised training of the multimodal classification model is performed based on the loss function L until a multimodal classification model with a loss function L less than a set value is obtained.
[0016] Furthermore, the MDF decoder includes a five-layer network, and except for the first layer which only includes the MDF-VSS module and the VSS module, the remaining layers all include the MDF-VSS module, the VSS module, the sampling module, the summation module, and the residual module.
[0017] The modality fusion classification results The method for obtaining it is as follows:
[0018] Optical multimodal projection features at various scales and SAR multimodal projection features The inputs are given to each layer of the MDF decoder network in ascending order of scale; the MDF-VSS module of each layer is used to acquire the mode-shared features of the optical multimodal projection features and SAR multimodal projection features received by itself. It is also used to obtain the modal difference features of the optical multimodal projection features and SAR multimodal projection features received by itself. Then, the MDF-VSS modules of each network layer are also used to fuse the modality-sharing features they have acquired. Modal difference features Obtain initial fusion features
[0019] The VSS modules of each network layer are used to extract the initial fusion features output by the MDF-VSS modules of the same network layer. Contextual information is used to obtain contextual features.
[0020] The sampling module of the second layer network is used to process the contextual features from the first layer network. Perform double upsampling; the summing module of the second layer network is used to extract the context features output by the VSS module of the second layer network. The summed features are added to the context features from the first layer network after being upsampled by two times to obtain the summed features; the residual module of the second layer network is used to fuse the summed features output by the summed module of this layer network to obtain the secondary fused features;
[0021] The sampling modules of the third layer to the fifth layer network are respectively used for doubling up-sampling the secondary fusion features from the previous layer network; the summation modules of the third layer to the fifth layer network are respectively used for adding the context features output by the VSS modules belonging to the same layer network as the summation modules with the secondary fusion features from the previous layer network and after doubling up-sampling, to obtain summation features; the residual modules of the third layer to the fifth layer network are respectively used for performing feature fusion on the summation features output by the summation modules of the layer network to obtain secondary fusion features; wherein the secondary fusion features output by the residual module of the fifth layer network are the final modal fusion classification results
[0022]
[0023] Further, the modal shared feature of the i-th layer network is The acquisition method is:
[0024]
[0025] wherein, is the optical multi-modal feature of the i-th scale is the optical multi-modal projection feature after projection to the set dimension, is the SAR multi-modal feature of the i-th scale is the SAR multi-modal projection feature after projection to the set dimension, VSS(·) represents the context feature extraction operation by the VSS module with shared parameters, and Concat(·) is a concatenation operator;
[0026] The modal difference feature of the i-th layer network is The acquisition method is:
[0027]
[0028] The initial fusion feature of the i-th layer network is The acquisition method is:
[0029]
[0030] wherein, Conv(·) represents a 3x3 convolution-batch normalization-ReLU convolution module, and Res(·) represents a feature fusion operation by a residual module.
[0031] Further, the MSIE-OPT decoder includes five layer networks, and except that the first layer network only includes a first summation module and a first residual module, the remaining layer networks all include a first summation module, a first residual module, a sampling module, a second summation module and a second residual module;
[0032] The optical fusion classification result is The acquisition method is:
[0033] The optical multi-modal features of each scale are projected to a set dimension to obtain optical multi-modal projection features
[0034] The optical multi-modal projection features of each scale are projected to a set dimension to obtain optical multi-modal projection features and modal difference features The layers of the MDF-OPT decoder are input in order from small scale to large scale; the first summation modules of the layers are respectively used for adding the optical multi-modal projection features and the modal difference features received by the first summation modules to obtain first sum value features; the residual modules of the layers are respectively used for performing feature fusion on the first sum value features output by the first summation modules belonging to the same layer as the residual modules to obtain initial optical fusion features;
[0035] The sampling module of the second layer network is used for doubling the initial optical fusion features from the first layer network; the second summation module of the second layer network is used for adding the initial optical fusion features output by the first residual module of the second layer network to the initial optical fusion features from the first layer network and doubled to obtain second sum value features; and the second residual module is used for performing feature fusion on the second sum value features of the layer to obtain secondary fusion features;
[0036] The sampling modules of the third layer to the fifth layer network are respectively used for doubling the secondary fusion features from the previous layer network; the second summation modules of the third layer to the fifth layer network are respectively used for adding the initial optical fusion features output by the first residual module belonging to the same layer as the second summation modules to the secondary fusion features from the previous layer network and doubled to obtain the second sum value features output by the layer; the second residual module is used for performing feature fusion on the second sum value features of the layer to obtain secondary fusion features; and the secondary fusion features output by the second residual module of the fifth layer network are the final optical fusion classification results
[0037] Further, the MSIE-SAR decoder includes five layers of networks, and in addition to the first layer network including only the first summation module and the first residual module, the remaining layers of networks each include the first summation module, the first residual module, the sampling module, the second summation module, and the second residual module;
[0038] The SAR fusion classification results The acquisition method is:
[0039] The SAR multi-modal features of each scale are projected to a set dimension to obtain SAR multi-modal projection features
[0040] SAR multimodal projection features at various scales Modal difference characteristics The data is input into each layer of the MDF-SAR decoder in order from small scale to large scale. The first summation module of each layer is used to add the received SAR multimodal projection features and modal difference features to obtain the first sum value feature. The first residual module of each layer is used to fuse the first sum value feature output by the first summation module of the same layer as itself to obtain the initial SAR fused features.
[0041] The sampling module of the second layer network is used to double-upsample the initial SAR fusion features from the first layer network; the second summing module of the second layer network is used to add the initial SAR fusion features output by the first residual module of the second layer network to the initial SAR fusion features from the first layer network after double upsampling, to obtain the second sum feature; the second residual module is used to perform feature fusion on the second sum feature of this layer network to obtain the secondary fusion feature;
[0042] The sampling modules of layers 3 through 5 are used to double-upsample the secondary fusion features from the previous layer. The second summing modules of layers 3 through 5 are used to add the initial SAR fusion features output by the first residual module of the same layer to the double-upsampled secondary fusion features from the previous layer, obtaining the second summed feature output by the current layer. The second residual module is used to fuse the second summed feature of the current layer to obtain the secondary fusion feature. The secondary fusion feature output by the second residual module of layer 5 is the final SAR fusion classification result.
[0043] Furthermore, the loss function L is calculated as follows:
[0044] L = L fuse +λ(L opt +L sar )
[0045] Among them, L fuse Let L be the cross-entropy loss function corresponding to the MDF decoder. opt Let L be the cross-entropy loss function corresponding to the MSIE-OPT decoder. sar Let λ be the cross-entropy loss function corresponding to the MSIE-SAR decoder, and λ be the set weight.
[0046] Furthermore, the cross-entropy loss function L corresponding to the MDF decoder fuse The calculation method is as follows:
[0047]
[0048] wherein CE(·) represents a cross-entropy loss function, Y represents an optical image X opt corresponding to the scene. sar corresponding to the scene.
[0049] Further, the cross-entropy loss function L opt of the MSIE-SAR decoder is calculated as follows:
[0050]
[0051] wherein CE(·) represents a cross-entropy loss function, Y represents an optical image X opt corresponding to the scene. sar corresponding to the scene.
[0052] Further, the cross-entropy loss function L sar of the MSIE-SAR decoder is calculated as follows:
[0053]
[0054] wherein CE(·) represents a cross-entropy loss function, Y represents an optical image X opt corresponding to the scene. sar corresponding to the scene.
[0055] Beneficial effects:
[0056] The present application provides a kind of modal information constraint-based multi-modal remote sensing image ground feature classification method, first using MDF decoder adaptive learning modal sharing and modal difference feature Fusion weight, adaptively control modal sharing and modal specific information contribution from multi-scale angle, to retain modal specific information;Then using MSIE decoder guides modal difference feature and multi-modal feature to carry out multi-modal auxiliary classification task, to enhance the optical and SAR specific information conducive to classification;Therefore, compared with other multi-modal remote sensing image ground feature classification method, the present application recognition precision is higher, can effectively alleviate ground feature recognition performance bottleneck under complex scene, provides effective support for multi-modal remote sensing image interpretation. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 the principle block diagram of the multi-modal classification model provided by the present application;
[0058] Figure 2 the principle block diagram of the MDF-VSS module provided by the present application;
[0059] Figure 3 a principle block diagram of the MDF decoder provided by the present application;
[0060] Figure 4 a principle block diagram of the MSIE-OPT decoder provided by the present application;
[0061] Figure 5 a principle block diagram of the MSIE-SAR decoder provided by the present application. DETAILED DESCRIPTION
[0062] In order to enable personnel in the art to better understand the present application, the technical solutions in the present application will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present application.
[0063] A multi-modal remote sensing image feature classification method based on modal information constraint, adopts a multi-modal classification model to classify features of a to-be-measured scene, wherein, as shown in the figure, the multi-modal classification model comprises an OPT encoder, a SAR encoder, an MDF decoder, an MSIE-OPT decoder and an MSIE-SAR decoder, and the training method of the multi-modal classification model is: Figure 1
[0064] S1: acquiring an optical image X opt Five different scale optical multi-modal features
[0065]
[0066] S2: acquiring a SAR image X sar Five different scale SAR multi-modal features Wherein, the optical image X opt and the SAR image X sar are taken for the same scene;
[0067] It should be noted that at this time, if the multi-modal features are projected to the same dimension, the corresponding multi-modal projection features and
[0068] S3: fusing the multi-modal projection features and by using the MDF decoder to obtain a modal fusion classification result At the same time, the MDF decoder is used to obtain the difference between and to obtain a modal difference feature
[0069] Specifically, as shown in the figure, Figure 3 As shown, the MDF decoder includes a five-layer network. Except for the first layer network, which only includes the MDF-VSS module and the VSS module, the other layers all include the MDF-VSS module, the VSS module, the sampling module, the summation module, and the residual module.
[0070] The modality fusion features The method for obtaining it is as follows:
[0071] Optical multimodal projection features at various scales and SAR multimodal projection features The inputs are given to each layer of the MDF decoder network in ascending order of scale; the MDF-VSS module of each layer is used to acquire the mode-shared features of the optical multimodal projection features and SAR multimodal projection features received by itself. It is also used to obtain the modal difference features of the optical multimodal projection features and SAR multimodal projection features received by itself. Then, the MDF-VSS modules of each network layer are also used to fuse the modality-sharing features they have acquired. Modal difference features Obtain initial fusion features
[0072] like Figure 2 As shown, the working principle of the MDF-VSS module in the i-th layer network is as follows:
[0073] First, for each scale, the multimodal projection features are obtained. and The VSS modules with shared parameters are mapped to a common embedding space to obtain optical shared features and SAR shared features. Then, these two projected features are concatenated dimensionally to obtain the modality shared features of the i-th layer network.
[0074]
[0075] in, Optical multimodal features at the i-th scale Optical multimodal projection features after projection onto a set dimension SAR multimodal features at the i-th scale The SAR multimodal projection features projected onto a set dimension, VSS(·) indicates that the context feature extraction operation is performed using shared parameters, that is, VSS modules with completely identical parameters, and Concat(·) is the splicing operator;
[0076] Then, for each layer, the obtained multimodal projection features and The difference feature of the multi-modal feature is calculated using subtraction, and the difference feature is modeled using a VSS module to extract the context information of the complex feature to the difference feature, and finally the modal difference feature of the i-th layer network is obtained
[0077]
[0078] Finally, the modal shared feature obtained above is added to the modal difference feature through a convolution module with a structure of 3x3 convolution-batch normalization-ReLU, and the feature representation capability is improved through a residual module to obtain the initial fusion feature of the i-th layer network
[0079]
[0080] Wherein, Conv(·) represents a convolution module with a structure of 3x3 convolution-batch normalization-ReLU, and Res(·) represents a feature fusion operation using a residual module.
[0081] The VSS module of each layer network is used to extract the context information of the initial fusion feature output by the MDF-VSS module belonging to the same layer network as itself to obtain the context feature
[0082] The sampling module of the second layer network is used to perform two times upsampling in the form of bilinear interpolation on the context feature from the first layer network The summation module of the second layer network is used to add the context feature output by the VSS module of the second layer network and the context feature from the first layer network after two times upsampling to obtain a summation feature; the residual module of the second layer network is used to perform feature fusion on the summation feature output by the summation module of the layer network to obtain a secondary fusion feature;
[0083] The sampling module of the third layer to the fifth layer network is used to perform two times upsampling on the secondary fusion feature from the previous layer network; the summation module of the third layer to the fifth layer network is used to add the context feature output by the VSS module belonging to the same layer network as itself and the secondary fusion feature from the previous layer network after two times upsampling to obtain a summation feature; the residual module of the third layer to the fifth layer network is used to perform feature fusion on the summation feature output by the summation module of the layer network to obtain a secondary fusion feature; wherein the secondary fusion feature output by the residual module of the fifth layer network is the final modal fusion classification result
[0084] Therefore, in the MDF decoder, the fusion features of each layer The context information of the complex ground object is extracted using the VSS module respectively, then the features are added to the adjacent higher resolution features through bilinear interpolation, and finally the residual module is fused to output the fused classification results
[0085] S4: MSIE-OPT decoder is used to fuse optical multi-modal projection features and modal difference features to obtain the optical fusion classification results
[0086] As shown in Figure 4 , the MSIE-OPT decoder includes a five-layer network, and except that the first layer network only includes a first summation module and a first residual module, the remaining layer networks each include a first summation module, a first residual module, a sampling module, a second summation module, and a second residual module.
[0087] The method for obtaining the optical fusion classification results is as follows:
[0088] After projecting the optical multi-modal features of each scale to a set dimension, optical multi-modal projection features
[0089] The optical multi-modal projection features of each scale and the modal difference features are input into the layer networks of the MDF-OPT decoder in the order from small scale to large scale; wherein the first summation module of each layer network is used to add the optical multi-modal projection features and the modal difference features received by itself to obtain a first sum value feature; the first residual module of each layer network is used to fuse the first sum value features output by the first summation modules belonging to the same layer network as itself to obtain initial optical fusion features;
[0090] The sampling module of the second layer network is used to perform two times up-sampling on the initial optical fusion features from the first layer network; the second summation module of the second layer network is used to add the initial optical fusion features output by the first residual module of the second layer network to the initial optical fusion features from the first layer network and subjected to two times up-sampling to obtain a second sum value feature; and the second residual module is used to fuse the second sum value features of the layer network to obtain twice fusion features.
[0091] The sampling modules of the third layer to the fifth layer network are respectively used for doubling up-sampling the secondary fusion features from the previous layer network; the second sum modules of the third layer to the fifth layer network are respectively used for adding the initial optical fusion features output by the first residual module belonging to the same layer network and the secondary fusion features from the previous layer network and after doubling up-sampling, to obtain the second sum value features output by the layer network; the second residual module is used for performing feature fusion on the second sum value features of the layer network to obtain the secondary fusion features; and the secondary fusion features output by the second residual module of the fifth layer network are the final optical fusion classification results
[0092] As can be seen, in the MSIE-OPT decoder, each layer of optical projection features and each layer of modal difference features are input into the MSIE-OPT decoder; first, each layer of modal difference features is added to the corresponding optical projection features ; second, the SAR specific information in the difference features that is beneficial to the classification task is enhanced through the first residual module; then the features are bilinearly interpolated and added to the adjacent higher resolution features, and finally fused through the second residual module; and finally, the fused classification results are output
[0093] S5: fusing SAR multi-modal projection features by using the MSIE-SAR decoder modal difference features to obtain SAR fusion classification results
[0094] As shown in Figure 5 , the MSIE-SAR decoder includes five layers of networks, and except that the first layer network only includes a first sum module and a first residual module, the remaining layer networks all include a first sum module, a first residual module, a sampling module, a second sum module and a second residual module.
[0095] The SAR fusion classification results are obtained by the following method:
[0096] After projecting the SAR multi-modal features of each scale to a set dimension, SAR multi-modal projection features
[0097] The SAR multi-modal projection features of each scale and the modal difference features The networks of each layer of the MDF-SAR decoder are input in order from small scale to large scale; wherein the first summation module of each layer of the network is respectively used for adding the projection SAR multi-modal feature and the modal difference feature received by itself to obtain a first sum value feature; the first residual module of each layer of the network is respectively used for fusing the first sum value feature output by the first summation module belonging to the same layer of network as itself to obtain an initial SAR fusion feature;
[0098] The sampling module of the second layer network is used for doubling up-sampling the initial SAR fusion feature from the first layer network; the second summation module of the second layer network is used for adding the initial SAR fusion feature output by the first residual module of the second layer network and the initial SAR fusion feature from the first layer network and after doubling up-sampling to obtain a second sum value feature; the second residual module is used for fusing the second sum value feature of the network to obtain a secondary fusion feature.
[0099] The sampling module of the third layer to the fifth layer network is respectively used for doubling up-sampling the secondary fusion feature from the last layer network; the second summation module of the third layer to the fifth layer network is respectively used for adding the initial SAR fusion feature output by the first residual module belonging to the same layer of network as itself and the secondary fusion feature from the last layer network and after doubling up-sampling to obtain the second sum value feature output by the network of the layer; the second residual module is used for fusing the second sum value feature of the network to obtain a secondary fusion feature; wherein the secondary fusion feature output by the second residual module of the fifth layer network is the final SAR fusion classification result
[0100] As can be seen, in the MDF-SAR decoder, each layer of SAR projection feature and each layer of modal difference feature is sent into the MSIE-SAR decoder; first, each layer of modal difference feature is added with the corresponding SAR projection feature ; second, the optical specific information in the difference feature which is beneficial to the classification task is enhanced through the first residual module; then these features are bilinearly interpolated and added with the adjacent higher resolution features, and finally fused through the second residual module; finally, the fused classification result is output
[0101] S6: constructing a loss function L according to the modal fusion classification result the optical fusion classification result and the SAR fusion classification result ;
[0102] Further, the calculation method of the loss function L is:
[0103] L = L fuse + λ (L opt + L sar )
[0104] wherein, L fuse is the cross-entropy loss function corresponding to the MDF decoder, L opt is the cross-entropy loss function corresponding to the MSIE-OPT decoder, L sar is the cross-entropy loss function corresponding to the MSIE-SAR decoder, and λ is a hyperparameter balancing the contributions of MSIE decoders, whose empirical value is set to 0.005.
[0105] The calculation method of the cross-entropy loss function L fuse corresponding to the MDF decoder is as follows:
[0106]
[0107] The calculation method of the cross-entropy loss function L opt corresponding to the MSIE-OPT decoder is as follows:
[0108]
[0109] The calculation method of the cross-entropy loss function L sar corresponding to the MSIE-SAR decoder is as follows:
[0110]
[0111] wherein, CE(·) is the cross-entropy loss function, Y represents the ground object label of the scene corresponding to the optical image X opt and the SAR image X sar used when training the multi-modal classification model.
[0112] S7: Supervise the training of the multi-modal classification model based on the loss function L until a multi-modal classification model with a loss function L less than a set value is obtained.
[0113] It should be noted that when the multi-modal classification model is supervised and trained, in each iteration, if the loss function L is not less than the set value, the value of the set weight in the loss function and the weight parameters of each layer network in the OPT encoder, the SAR encoder, the MDF decoder, the MSIE-OPT decoder and the MSIE-SAR decoder in L Figure 1 are changed to update the multi-modal classification model, and then iteration is performed again until the loss function L meets the requirements and the training is ended.
[0114] Taking the practical application scene of multi-modal remote sensing image feature classification as an example, compared with a plurality of existing multi-modal remote sensing image feature classification methods on a plurality of public remote sensing data sets, as shown in Tables 1 and 2, it can be seen that the present application has better accuracy.
[0115] Table 1 Performance comparison of a plurality of existing multi-modal remote sensing image feature classification methods and embodiments on the DFC20 data set
[0116]
[0117]
[0118] Table 2 Performance comparison of a plurality of existing multi-modal remote sensing image feature classification methods and embodiments on the WHU-OPT-SAR data set
[0119]
[0120] In summary, the present application proposes a multi-modal remote sensing image feature classification method based on modal specific information reservation and enhancement, which realizes adaptive control of modal sharing and contribution of modal specific information through a multimodal feature decomposition and fusion (MDF)-visual state space (VSS) module, and introduces a MDF decoder based on the module to realize adaptive control of multi-scale modal sharing and contribution of modal specific information, thereby realizing reservation of modal specific information beneficial to the classification task. In addition, by introducing a multimodal specific information enhancement (MSIE) decoder, an auxiliary classification task guided by modal difference features is carried out to enhance the modal specific information beneficial to classification. As can be seen, the present application enhances the representation ability of the fused features; a large number of experiments show that the present application has stronger performance than other multi-modal remote sensing image feature classification methods; the method proposed in the present application has higher recognition accuracy than other multi-modal remote sensing image feature classification methods, which can effectively alleviate the performance bottleneck of feature recognition in complex scenes and provide effective support for multi-modal remote sensing image interpretation.
[0121] Of course, the present application can also have other various embodiments, and those skilled in the art can certainly make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application, but these corresponding changes and modifications should all belong to the protection scope of the claims attached to the present application.
Claims
1. A method for building feature classification of multi-modal remote sensing images based on modal information constraint, characterized in that, The multi-modal classification model is used for ground object classification of a to-be-tested scene, wherein the multi-modal classification model comprises an OPT encoder, an SAR encoder, a multi-modal feature decomposition and fusion decoder MDF, and a modal specific information enhancement decoder MSIE, wherein the MSIE comprises an MSIE-OPT decoder and an MSIE-SAR decoder; the MDF comprises five layers of networks, and except that the first layer of network only comprises an MDF-VSS module and a VSS module, the remaining layers of network each comprise an MDF-VSS module, a VSS module, a sampling module, a summation module, and a residual module, wherein VSS is a visual state space; the MSIE-OPT decoder comprises five layers of networks, and except that the first layer of network only comprises a first summation module and a first residual module, the remaining layers of network each comprise a first summation module, a first residual module, a sampling module, a second summation module, and a second residual module; The training method of the multi-modal classification model is as follows: Acquiring optical images X with an OPT encoder opt Optical multimodal features of five different scales SAR images X are acquired using a SAR encoder sar SAR multi-modal features of five different scales wherein the optical images X opt with the SAR images X sar are taken of the same scene; projecting the multi-modal features and to the same dimensionality, resulting in corresponding multi-modal projected features and Fusing with MDF decoder and Obtaining modal fusion classification results At the same time, obtaining modal difference features by using the MDF decoder to obtain the difference between and Fusing optical multi-modal projection features with MSIE-OPT decoder and modal difference features obtaining optical fusion classification results Fusing sar multi-modal projection features with MSIE-SAR decoder and modal difference features get sar fusion classification result According to the modal fusion classification result Optical fusion classification result SAR fusion classification result Construct a loss function L; The multi-modal classification model is supervised trained based on a loss function L until the multi-modal classification model with a loss function L less than a set value is obtained. 2.The method of claim 1, wherein, The modal fusion classification result The acquisition method is: Optical multi-modal projection features at various scales and SAR multi-modal projection features The layers of the MDF decoder are inputted in order from small scale to large scale respectively; wherein the MDF-VSS modules of the layers are respectively used to obtain modal shared features of the optical multi-modal projection features and the SAR multi-modal projection features received by the layers themselves and are also used to obtain modal difference features of the optical multi-modal projection features and the SAR multi-modal projection features received by the layers themselves The MDF-VSS modules of the layers are then also used to fuse the modal shared features obtained by the layers themselves and the modal difference features to obtain initial fusion features The VSS modules of the respective layer networks are respectively configured to extract the initial fusion features output by the MDF-VSS modules of the layer networks to which the respective layer networks belong to obtain context features The sampling module of the second layer network is configured to sample the context features from the first layer network perform double up-sampling; the sum module of the second layer network is configured to sum the context features output by the VSS module of the second layer network to obtain summed features; the residual module of the second layer network is configured to perform feature fusion on the summed features output by the sum module of the second layer network to obtain secondary fusion features; The sampling modules of the third layer to the fifth layer network are respectively used for doubling up-sampling the secondary fusion features from the previous layer network; the sum modules of the third layer to the fifth layer network are respectively used for adding the context features output by the VSS modules belonging to the same layer network as the sum modules with the secondary fusion features from the previous layer network and after doubling up-sampling, to obtain sum features; the residual modules of the third layer to the fifth layer network are respectively used for performing feature fusion on the sum features output by the sum modules of the layer network, to obtain secondary fusion features; wherein the secondary fusion features output by the residual module of the fifth layer network are the final modal fusion classification results 3.The method of claim 2, wherein, Modal sharing features of the ith layer network The acquisition method is that: wherein, optical multi-modal feature of the i-th scale projected optical multi-modal projected feature, SAR multi-modal feature of the i-th scale projected SAR multi-modal projected feature, VSS(·) denotes a context feature extraction operation using a VSS module with shared parameters, and Concat(·) is a concatenation operator. Modal difference features of the ith layer network The acquisition method is: Initial fusion features of the i-th layer network The acquisition method is as follows: Wherein, Conv(·) represents a 3*3 convolution-batch normalization-ReLU convolution module, and Res(·) represents a feature fusion operation using a residual module. 4.The method of claim 1, wherein, The optical fusion classification result The acquisition method is: optical multi-modal features of each scale after projection to a set dimension, optical multi-modal projection features are obtained Optical multi-modal projection features of various scales And modal difference features The layers of the MDF-OPT decoder are input in order from small scale to large scale; wherein the first summation modules of the layers are respectively used for adding the optical multi-modal projection features and the modal difference features received by themselves to obtain first sum value features; the residual modules of the layers are respectively used for feature fusion on the first sum value features output by the first summation modules belonging to the same layer as themselves to obtain initial optical fusion features; The sampling module of the second layer of network is used for performing two times of upsampling on the initial optical fusion feature from the first layer of network; the second summation module of the second layer of network is used for adding the initial optical fusion feature output by the first residual module of the second layer of network and the initial optical fusion feature from the first layer of network and subjected to two times of upsampling, to obtain a second summation value feature; and the second residual module is used for performing feature fusion on the second summation value feature of the current layer of network to obtain a secondary fusion feature; The sampling modules of the third layer to the fifth layer network are respectively used for doubling up-sampling the secondary fusion features from the previous layer network; the second sum modules of the third layer to the fifth layer network are respectively used for adding the initial optical fusion features output by the first residual module belonging to the same layer network as itself and the secondary fusion features from the previous layer network and after doubling up-sampling, to obtain the second sum value features output by the layer network; the second residual module is used for performing feature fusion on the second sum value features of the layer network to obtain the secondary fusion features; wherein the secondary fusion features output by the second residual module of the fifth layer network are the final optical fusion classification results 5.The method of claim 1, wherein, The MSIE-SAR decoder comprises five layers of networks, and except that the first layer of network only comprises a first summation module and a first residual module, the remaining layers of network each comprise a first summation module, a first residual module, a sampling module, a second summation module, and a second residual module; The SAR fusion classification result The acquisition method is: SAR multi-modal features of each scale are projected to a set dimension to obtain SAR multi-modal projection features SAR multi-modal projection features SAR multi-modal projection features of various scales and modal difference features The layers of the MDF-SAR decoder are input in order from small scale to large scale; a first summation module of each layer is configured to add the SAR multi-modal projection features and the modal difference features received by the first summation module to obtain a first sum value feature; a first residual module of each layer is configured to fuse the first sum value features output by the first summation modules of the same layer to obtain an initial SAR fusion feature; The sampling module of the second layer of network is used for performing two times of upsampling on the initial SAR fusion feature from the first layer of network; the second summation module of the second layer of network is used for adding the initial SAR fusion feature output by the first residual module of the second layer of network and the initial SAR fusion feature from the first layer of network and subjected to two times of upsampling, to obtain a second summation value feature; and the second residual module is used for performing feature fusion on the second summation value feature of the current layer of network to obtain a secondary fusion feature; The sampling modules of the third layer to the fifth layer network are respectively used for doubling up-sampling the secondary fusion features from the previous layer network; the second sum modules of the third layer to the fifth layer network are respectively used for adding the initial SAR fusion features output by the first residual module belonging to the same layer network as itself and the secondary fusion features from the previous layer network and after doubling up-sampling, to obtain the second sum value features output by the layer network; the second residual module is used for performing feature fusion on the second sum value features of the layer network to obtain the secondary fusion features; wherein the secondary fusion features output by the second residual module of the fifth layer network are the final SAR fusion classification results 6.The method of claim 1, wherein, The calculation method of the loss function L is as follows: L = L fuse + λ(L opt + L sar ) wherein L fuse is the cross-entropy loss function corresponding to the MDF decoder, L opt is the cross-entropy loss function corresponding to the MSIE-OPT decoder, L sar is the cross-entropy loss function corresponding to the MSIE-SAR decoder, and λ is a set weight.
7. The method of claim 6, wherein the method is based on a constraint of modal information. The MDF decoder corresponds to a cross-entropy loss function L fuse The calculation method is: where CE(·) denotes a cross-entropy loss function, Y denotes the optical image X opt corresponding to the SAR image X sar the ground object label of the corresponding scene. 8.The method of claim 6, wherein the method further comprises: The cross-entropy loss function L corresponding to the MSIE-OPT decoder opt The calculation method is: where CE(·) denotes a cross-entropy loss function, Y denotes the optical image X opt corresponding to the scene of the SAR image X sar corresponding to the scene of the SAR image X 9.The method of claim 6, wherein the method further comprises: determining a plurality of feature maps based on the plurality of feature extraction layers; and determining a plurality of feature vectors based on the plurality of feature maps. The cross-entropy loss function L corresponding to the MSIE-SAR decoder sar The calculation method is: where CE(·) denotes a cross-entropy loss function, Y denotes the optical image X opt corresponding to the SAR image X sar the ground object label of the corresponding scene.
Citation Information
Patent Citations
Land cover classification method based on deep fusion of multi-modal remote sensing data
CN113469094A
Ground feature classification method based on multi-source remote sensing image
CN117541873A