Multi-band multi-polarization SAR image fusion classification method for modal missing and non-registration scenes

Through the multi-band multi-polarization SAR image fusion classification method with multi-stream structure and attention mechanism, the problem of decreased classification accuracy caused by image inconsistency and modality loss is solved, and high reliability and high adaptability classification are achieved in complex remote sensing scenarios.

CN120766020APending Publication Date: 2025-10-10HARBIN INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510894335.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing multi-band multi-polarimetric SAR images suffer from the problem of reduced classification accuracy when the image sizes are inconsistent, the frequency bands cannot be aligned, and the modes are missing. In addition, existing methods fail to effectively mine the discriminative and correlation characteristics of each frequency band, limiting the effectiveness of fusion.

Method used

A multi-stream structure is used for data preprocessing, image blocks are sampled independently, polarization features are extracted through a multi-stream sub-network, category conditional alignment loss function and modal Dropout mechanism are introduced, and feature fusion is performed in combination with the attention mechanism. Finally, classification decisions are made through a classifier.

Benefits of technology

Under conditions of incomplete frequency bands or inconsistent resolutions, the semantic consistency and robustness of multi-band polarimetric SAR images are improved, making them suitable for high-reliability classification in complex remote sensing scenarios and enhancing the generalization ability and classification performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766020A_ABST
    Figure CN120766020A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-band multi-polarization SAR (synthetic aperture radar) image fusion classification method for modal missing and non-registration scenes. The method comprises the following steps: step 1, data preprocessing; step 2, feature extraction; step 3, carrying out cross-modal constraint; step 4, carrying out modal Dropout; 5, performing feature fusion; and step 6, performing classification decision. According to the method, spatial alignment or resampling does not need to be carried out on the original image, semantic consistency and modal fusion robustness among different frequency bands can be effectively improved, high-reliability and high-adaptability polarized SAR image intelligent classification in a complex remote sensing scene is realized under the condition that the frequency bands are incomplete or the resolutions are inconsistent, and the method is suitable for popularization and application. The method solves the problem that the classification precision of the existing multi-band multi-polarization SAR image is reduced under the conditions of inconsistent image size, incapability of aligning frequency bands, mode loss and the like, and is suitable for various remote sensing application scenes such as disaster monitoring, land coverage identification, military reconnaissance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of remote sensing image processing and artificial intelligence classification, and relates to a SAR image fusion classification method, in particular to a target classification method suitable for non-registered and multi-modal polarimetric SAR images. BACKGROUND

[0002] Polarimetric synthetic aperture radar (polarimetric SAR) has all-weather and all-day imaging capability, and can provide rich scattering mechanism information, which plays an important role in remote sensing applications such as ground object classification, target detection, and disaster monitoring. With the increasing demand for high-precision ground object identification, the traditional polarimetric SAR image classification method based on a single frequency band often has insufficient discrimination when facing complex scenes, multi-scale structures and similar inter-class ground objects. Therefore, how to effectively utilize multi-source polarimetric information and improve classification accuracy and robustness has become a key research direction in current remote sensing intelligent interpretation.

[0003] In recent years, intelligent classification technology that fuses multi-band and multi-polarization SAR data has gradually attracted attention. Multi-band data has complementary properties in spatial resolution, penetration ability and scattering response, providing more comprehensive discrimination for ground object identification. At the same time, multi-modal feature fusion methods based on deep learning can realize automatic semantic modeling driven by data, and have significant advantages in modal collaborative understanding and structure expression. However, existing methods still have problems such as non-uniform size of frequency band images, non-alignment of modalities, and sharp decline in model performance when some frequency bands are missing. In addition, most methods fail to effectively exploit the discriminative and correlation features of each frequency band in the fusion stage, limiting the performance of fusion. Inspired by this, it is urgent to build a multi-band polarimetric SAR image intelligent classification method with modal robustness, semantic alignment capability and fusion adaptability to better meet the high-reliability interpretation needs in complex remote sensing scenarios. SUMMARY

[0004] To solve the problem of declining classification accuracy of existing multi-band and multi-polarization SAR images under conditions of inconsistent image size, non-aligned frequency bands and missing modalities, the present application provides a multi-band and multi-polarization SAR image fusion classification method for modal missing and non-registered scenarios. This method does not require spatial alignment or resampling of the original image, and can effectively improve the semantic consistency and modal fusion robustness between different frequency bands. It realizes high-reliability and high-adaptability polarimetric SAR image intelligent classification in complex remote sensing scenarios under the condition of incomplete frequency bands or inconsistent resolution, and is suitable for various remote sensing application scenarios such as disaster monitoring, land cover identification and military reconnaissance.

[0005] The purpose of the present application is achieved by the following technical solutions:

[0006] A multi-band multi-polarization SAR image fusion classification method for missing modalities and non-registration scenes, comprising the following steps:

[0007] Step 1: data preprocessing: independent sampling is performed on the polarimetric SAR images of different frequency bands respectively, and fixed-size image blocks are extracted, and the specific steps are as follows:

[0008] Step 1.1: Set a multi-band multi-polarization SAR image set , wherein the first band image is represented by , and are the spatial dimensions of the image, is the number of polarization channels;

[0009] Step 1.2: Extract fixed-size image blocks patch from each image , denoted as , wherein is the width and height of the patch, indicates the sample number, and the patch sampling positions between different frequency bands are not corresponding, and the spatial positions are completely independent;

[0010] Step 1.3: To unify the input structure, the upper triangular elements of the coherence matrix are expanded according to the real part and the imaginary part, and the input form is constructed as:

[0011]

[0012] , wherein is the pixel position in the patch; the vector is composed of the real part and the imaginary part of the upper triangular element of the coherence matrix , indicates the element in the first row and the first column of the matrix, and for the upper triangular matrix, the elements satisfy ;

[0013] Step 2: feature extraction: the image blocks sampled from different frequency bands are respectively input into the corresponding multi-stream sub-network, and the polarization feature representation of each frequency band is extracted, and the specific steps are as follows:

[0014] Step 2.1: An independent feature extraction sub-network is constructed for each frequency band image, and a parallel multi-stream structure with shared parameters is adopted, and each sub-network is responsible for extracting the polarization scattering features in the corresponding frequency band image block, maintaining its modal independence;

[0015] Step 2.2: After the feature extraction sub-network After that, we get the feature map under this mode :

[0016]

[0017] in, are the width and height of the feature map, is the number of channels of the feature map of this frequency band;

[0018] Step 3: Cross-modal constraints: After feature extraction, a category-conditional alignment loss function is introduced at the feature layer. The maximum mean difference (MMD) is used to align the feature distributions of the same category in different frequency bands. The specific steps are as follows:

[0019] Step 3.1: To eliminate the statistical differences between feature representations of different frequency bands and enhance cross-modal semantic consistency, a category-conditional feature alignment loss is introduced during the training phase. :

[0020]

[0021] in, and Respectively represent the feature representations extracted from the C-band and L-band, Indicates that the sample belongs to Class, a total of categories; represents the square of the maximum mean difference;

[0022] Step 3.2: Based on the MMD principle, explicitly align the feature distributions of different frequency bands under the same category;

[0023] Step 3.3: To improve computational efficiency, MMD can be implemented using kernel functions, such as Gaussian kernels or multi-kernel hybrids. Typical definitions are as follows:

[0024]

[0025] in, Represents two distributions and The maximum squared mean difference between and They are respectively and The number of samples in Represents a sample and The kernel function value between and Respectively represent the distribution Sample and distribution and The kernel function value of the interval sample;

[0026] Step 3.4: Feature alignment loss of category condition with the main classification loss Joint optimization, constituting the total loss function :

[0027]

[0028] wherein, is the weight coefficient;

[0029] Step 4: Modal Dropout: During the training process, a modal Dropout mechanism is applied to the input image block of each frequency band, that is, a certain frequency band input is randomly shielded with a set probability. The specific steps are as follows:

[0030] To enhance the robustness of the model in the case of missing frequency bands, a modal Dropout mechanism is introduced in the training stage. By randomly shielding the input image block of a certain frequency band or multiple frequency bands with a preset probability in each training batch, the situation where some frequency bands are unavailable in the actual scenario is simulated, forcing the model to learn the cross-modal redundancy and compensation ability:

[0031]

[0032] wherein, represents the input image block of the th frequency band after modal Dropout processing, the value of which remains the original input with a probability and is set to 0 with a probability ;

[0033] Step 5: Feature fusion: The feature representation of each frequency band is sent to the modal fusion module, and an adaptive weight is given to the feature of each frequency band through the attention mechanism to obtain a fused feature vector. The specific steps are as follows:

[0034] Step 5.1: There are a total of modalities, and the feature map output by each modality is , wherein is the width and height of the feature map, is the unified channel number of each modality feature map, then the spliced feature is:

[0035]

[0036] Step 5.2: Apply channel attention mechanism to the spliced joint feature : First, the feature vector is obtained through adaptive average pooling operation, wherein has a dimension of ​; then the attention weights are calculated by two fully connected networks where and are the weight matrices of the two fully connected networks, respectively, denotes the Sigmoid activation function; finally, the attention weights are applied to the joint features by a channel-wise scaling operation to obtain the weighted features where denotes the channel-wise scaling operation:

[0037]

[0038]

[0039]

[0040] Step 5.3: On the basis of the channel-enhanced features, further introduce a spatial attention mechanism by calculating the maximum and average values at each position, concatenating them, and inputting them into a convolutional layer to generate a spatial attention map that emphasizes discriminative spatial regions: first, calculate the average value and the maximum value in the channel dimension; then, concatenate and along the channel dimension to form a new feature map, and then generate a spatial attention map by a convolutional layer with filters; finally, apply the spatial attention map to the channel-enhanced feature map by element-wise multiplication to obtain the fused feature vector :

[0041]

[0042]

[0043]

[0044]

[0045] Step 6: Classification decision: input the fused feature vector into the classifier to output the corresponding class label, realizing supervised classification of the input image block, with the specific steps as follows:

[0046] Step 6.1: input the fused feature vector into the classifier to complete the prediction of the image block class:

[0047]

[0048] wherein, is a trainable weight, is a bias term, denotes an activation function;

[0049] Step 6.2: Class probability output by the classifier wherein, the dimension represents the confidence of the sample belonging to the class, and the final prediction class is determined by the index corresponding to the maximum probability, that is:

[0050]

[0051] Step 6.3: In the training stage, the prediction result is supervised and optimized with the real label by using the cross-entropy loss function, and the loss function is defined as follows:

[0052]

[0053] wherein, is a one-hot encoding representation of the real label.

[0054] Compared with the prior art, the present application has the following advantages:

[0055] 1. The present application adopts a multi-flow structure to sequentially complete six links of data preprocessing, feature extraction, cross-modal constraint, modal Dropout, feature fusion and classification decision. In the specific implementation process, first, image blocks are independently sampled from different frequency bands, avoiding the destruction of resolution and geometric structure caused by spatial alignment, and preserving the integrity of polarization information; then, multi-flow sub-networks are used to extract polarization SAR image deep features specific to the frequency band, ensuring the expression independence between modalities; further, a class condition alignment constraint is introduced at the feature layer to realize semantic consistent modeling of the same class targets across frequency bands; in the training process, a modal Dropout mechanism is introduced to effectively improve the fault tolerance of the model to missing or abnormal input of the frequency band; finally, the multi-frequency band features are weighted and fused through an attention mechanism, and input into the classifier to complete accurate discrimination. Each link works together to realize the whole process optimization from the original modal heterogeneous data to the semantic consistent output, ensuring the applicability and stability of the system under actual complex remote sensing conditions.

[0056] 2、The application breaks through the dependence on image space alignment in the data preparation stage, avoids texture distortion and polarization information loss caused by image registration, and is suitable for polarimetric SAR images with different resolutions and acquisition angles; the class condition alignment strategy overcomes the problem of large modal distribution difference in traditional fusion methods, significantly improves the semantic consistency of cross-modal, introduces the modal Dropout mechanism, effectively solves the problem of sharp performance decline of existing methods in modal loss, and enhances the generalization ability of the model; the attention fusion mechanism replaces the simple feature splicing method, which can adaptively mine the most discriminative components in each frequency band, thereby realizing more efficient information integration and classification performance improvement. These improvements jointly construct a robust, efficient and flexible multi-frequency polarimetric SAR image intelligent interpretation system, which is obviously superior to existing methods based on single modal or static fusion strategy. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 is a decision flowchart of the method of the application;

[0058] Figure 2 is a non-registration processing sampling method of the method of the application;

[0059] Figure 3 is a network structure schematic diagram of the method of the application;

[0060] Figure 4 is a network training curve of the method of the application;

[0061] Figure 5 is the experimental result of the method of the application. DETAILED DESCRIPTION

[0062] The technical solutions of the application will be further described below in conjunction with the drawings, but are not limited thereto. Any modification or equivalent replacement to the technical solutions of the application without departing from the spirit and scope of the technical solutions of the application shall be covered in the protection scope of the application.

[0063] The application provides a multi-frequency multi-polarization SAR image fusion and classification method for modal loss and non-registration scenarios, as shown in Figure 1 The method comprises the following steps:

[0064] Step 1: data preprocessing: independently sampling from polarimetric SAR images of different frequency bands, extracting fixed-size image blocks, and no spatial alignment is required between different frequency bands. Only the consistency of class labels needs to be ensured during sampling, which is used for input sample construction in the training or test stage. The specific steps are as follows:

[0065] Step 1.1: a multi-frequency multi-polarization SAR image set is provided, wherein the th frequency band image is represented as is the spatial size of the image, is the number of polarization channels (e.g. represents the 9 independent elements of the upper triangular part of the polarimetric coherence matrix, divided by real and imaginary parts. Since the spatial sizes of images in different frequency bands are inconsistent, there is no need for registration between images, only consistent class label information needs to be shared.

[0066] Step 1.2: Extract fixed-size image patches (patches) from each image , denoted as , where is the width and height of the patch, indicates the sample number. The sampling method can use a sliding window or a class-guided random sampling strategy. To achieve class balance, the same number of samples are sampled from each frequency band for each class.

[0067] Step 1.3: The patch sampling positions between different frequency bands do not correspond to each other, and the spatial positions can be completely independent; only the class labels are consistent during training. Considering that the polarization SAR image is dominated by pixel scattering characteristics, the model training can not depend on spatial alignment, so as to fully exploit the cross-band information in the non-registration scenario.

[0068] Step 1.4: To unify the input structure, the upper triangular elements of the coherence matrix can be unfolded according to the real and imaginary parts. The input form can be constructed as:

[0069] (1)

[0070] where is the pixel position in the patch. The vector is composed of the real and imaginary parts of the upper triangular elements of the coherence matrix . The coherence matrix is a matrix containing complex elements, where represents the element in the row and the column of the matrix, and for the upper triangular matrix, its elements satisfy . Each element in the vector is extracted from the upper triangular part of the coherence matrix , the real part is extracted first, and then the imaginary part , these real and imaginary parts are arranged in turn to form the vector. The above method ensures that the input dimensions of different frequency bands are consistent, which is convenient for subsequent multi-stream network processing.

[0071] ​​​Step 2: Feature extraction: The image blocks sampled from different frequency bands are input into the corresponding multi-stream sub-networks. Each sub-network has the same structure but independent parameters, and is used to extract the polarization feature representation of each frequency band. The specific steps are as follows:

[0072] Step 2.1: Construct an independent feature extraction sub-network for each frequency band image, using a parallel multi-stream structure with no parameter sharing. Each sub-network is responsible for extracting the polarization scattering features in the image block of the corresponding frequency band, maintaining its modality independence. The image block input of the frequency band is , after feature extraction sub-network After that, we get the feature map under this mode :

[0073] (2)

[0074] Among them, the parameters Indicates the The image block input of frequency bands has a size of ,in are the width and height of the image block, and Indicates the number of channels of the image block in this frequency band. The size is ,in are the width and height of the feature map, is the number of channels in the feature map of the frequency band, representing the number of features extracted from the original image block. In this way, the image blocks of each frequency band are processed independently to maintain the independence of their modalities, so that the unique information of image blocks of different frequency bands can be better captured and distinguished during the feature extraction process.

[0075] Step 2.2: Each sub-network A lightweight convolutional neural network architecture can be used. A typical configuration includes two or more convolutional layers, a batch normalization layer, and an activation function (such as ReLU), followed by global average pooling to generate fixed-length features. This architecture can extract spatial context while preserving frequency-band specificity.

[0076] Step 2.3: To achieve subsequent feature alignment and fusion operations, the output feature map dimensions of all sub-networks must be consistent, that is, the dimension of each feature map is The sub-network structure parameters can be optimized through cross-validation or transfer learning to adapt to the statistical characteristics of signals in different frequency bands.

[0077] Step 2.4: The multi-stream architecture allows each frequency band to maintain independent distribution during the extraction process, avoiding information interference and enhancing the model's ability to model inter-modal differences. Furthermore, this design facilitates reasoning based on only some of the streams when a modality is missing, providing greater flexibility and generalization.

[0078] Step 3: Cross-modal constraint: After feature extraction, introduce a class-condition alignment loss function at the feature layer to align the feature distributions of the same class under different frequency bands by maximum mean discrepancy (MMD), and improve the multi-modal semantic consistency. The specific steps are as follows:

[0079] Step 3.1: To eliminate the statistical differences between different frequency band feature representations and enhance cross-modal semantic consistency, introduce a class-condition feature alignment loss during the training phase :

[0080] (3)

[0081] where, and represent the feature representations extracted from the C frequency band and the L frequency band, respectively, and represents that the sample belongs to the th class, and there are a total of classes. This loss function measures the distribution distance of the same class samples in two frequency bands through the kernel method, guiding the network to learn a class-consistent feature embedding space.

[0082] Step 3.2: Based on the principle of maximum mean discrepancy (MMD), the feature distributions of different frequency bands under the same class are explicitly aligned.

[0083] Step 3.3: To improve computational efficiency, MMD can be implemented in the form of a kernel function, such as a Gaussian kernel or a multi-kernel hybrid approach, with the typical definition as follows:

[0084] (4)

[0085] where, represents the square maximum mean discrepancy between two distributions and . and are the number of samples in distributions and , respectively. represents the kernel function value between samples and , and similarly and represent the samples in distribution and the distributions and By calculating these kernel function values ​​and substituting them into the formula, the difference between the two distributions can be evaluated, which can be used for feature alignment and distribution matching.

[0086] Step 3.4: Class-Conditional Feature Alignment Loss Can be compared with the main classification loss (such as cross entropy) are jointly optimized to form the total loss function:

[0087] (5)

[0088] in, is the weight coefficient used to balance the optimization objectives of supervised classification and modality alignment.

[0089] This alignment strategy does not rely on the spatial position of the image or external registration information, but only performs alignment based on semantic labels. It is particularly suitable for multi-band scenarios where the original images have inconsistent resolution or spatial misalignment, effectively improving the model's semantic fusion capabilities and cross-modal discrimination performance.

[0090] Step 4: Modal Dropout: During the training process, a modal dropout mechanism is applied to the input image blocks of each frequency band, that is, the input of a certain frequency band is randomly blocked with a set probability, guiding the model to learn feature compensation and redundant information expression capabilities under modality missing conditions. The specific steps are as follows:

[0091] Step 4.1: To enhance the model's robustness to missing frequency bands, a modal dropout mechanism is introduced during training. This mechanism randomly blocks one or more frequency bands of input image patches with a preset probability in each training batch, simulating the unavailability of certain frequency bands in real-world scenarios and forcing the model to learn cross-modal redundancy and compensation capabilities.

[0092] (6)

[0093] in, Represents the first The input image block of frequency bands, whose value is expressed as probability Keep the original input , with probability is set to 0. In this way, the model can adapt to the situation of missing frequency bands during training, improving its robustness and generalization ability in practical applications.

[0094] Step 4.2: Modality Dropout can be performed at the input or feature layer. It is recommended that you apply it before the image patch enters the feature extraction network to maximize modality independence and simplify network processing. During inference, all modalities remain enabled and are no longer randomly masked.

[0095] Step 4.3: The addition of modal dropout can effectively improve the model's adaptability to incomplete or damaged modal samples in the training data, and enhance its fault tolerance performance in practical remote sensing applications. It can be set according to the expected probability of missing modalities, and is typically in the range of 0.1 to 0.5.

[0096] Step 4.4: No additional modifications are required to the Dropout operation in the training loss function. The standard cross entropy loss or other supervised learning loss can be kept unchanged. Joint training with feature alignment or modality fusion modules is also supported.

[0097] Step 5: Feature fusion: The feature representation of each frequency band is sent to the modality fusion module, and the features of each frequency band are given adaptive weights through the attention mechanism to obtain the final fusion feature vector. The specific steps are as follows:

[0098] Step 5.1: To enhance the fusion effect of multi-band multi-polarization SAR features, the feature maps output by each modality sub-network are spliced ​​in the channel dimension to form a joint feature tensor. modalities, and the feature map output by each modal is ,in are the width and height of the feature map, is the uniform number of channels of each modal feature map, then the concatenated features are :

[0099] (7)

[0100] Step 5.2: Combined features Apply channel attention mechanism. First, the feature vector is obtained by adaptive average pooling operation ,in The dimension is Then, the attention weight is calculated through a two-layer fully connected network ,in and are the weight matrices of the two-layer fully connected network, Represents the Sigmoid activation function, which is used to limit the output between 0 and 1. Finally, the attention weight is applied to the joint feature through a channel-by-channel scaling operation , and get the weighted features ,in represents a channel-by-channel scaling operation. In this way, the model can adaptively emphasize important features and suppress unimportant features, thereby improving the discriminative ability of feature representation.

[0101] (8)

[0102] (9)

[0103] (10)

[0104] Step 5.3: Based on the enhanced features in the channel dimension, a spatial attention mechanism is further introduced. By calculating the maximum and average values at each position and concatenating them before inputting into a convolution layer, a spatial attention map is generated to emphasize the discriminative spatial regions. Specifically, first, the maximum and average values of the feature map are calculated The dimensions of these two feature maps are , where denotes the spatial dimensions (width and height) of the feature map. Then, and are concatenated along the channel dimension to form a new feature map, which is then passed through a convolution layer to generate a spatial attention map with a dimension of . Finally, the spatial attention map is applied to the channel-enhanced feature map through element-wise multiplication to obtain the fused feature , which has the same dimension as , i.e., . Here, denotes the Sigmoid activation function, which limits the output of the convolution layer to between 0 and 1, thereby generating attention weights; denotes the element-wise multiplication operation. In this way, the model can adaptively emphasize important spatial regions and suppress unimportant ones, thereby improving the discriminative ability of the feature representation.

[0105] (11)

[0106] (12)

[0107] (13)

[0108] (14)

[0109] Step 5.4: The final fused feature has both channel and spatial significance enhancement capabilities, retains complementary information in multiple frequency band modalities, and suppresses redundant or interfering features, providing more refined and discriminative feature representations for subsequent classification tasks.

[0110] ​Step 6: Classification decision: The fused feature vector is input into the classifier, and the corresponding class label is output to realize the supervised classification of the input image block. The specific steps are as follows:

[0111] Step 6.1: The fused feature map is input into the classifier to complete the prediction of the image block class. The classifier can be composed of three groups of fully connected layers and nonlinear activation functions, and the output is the class probability distribution.

[0112] (15)

[0113] where, is the trainable weight, is the bias term, represents the activation function, usually ReLU.

[0114] Step 6.2: The class probability output by the classifier , where the th dimension represents the confidence that the sample belongs to the th class. The final prediction class is determined by the index corresponding to the maximum probability, that is:

[0115] (16)

[0116] Step 6.3: In the training stage, the cross-entropy loss function is used to supervise and optimize the prediction result and the true label, and the loss function is defined as follows:

[0117] (17)

[0118] where, is the one-hot encoding representation of the true label.

[0119] Step 6.4: The classifier is jointly trained with the aforementioned modal Dropout module, class alignment module, and attention fusion module in an end-to-end manner, and the entire multi-modal feature learning and classification process is optimized. This structure also supports partial modal input testing, that is, any modal combination can be classified in the inference stage.

[0120] Step 6.5: This output mechanism can be used for patch-level classification prediction, and can also be extended to whole-image-level classification map construction through sliding window, realizing the polarization SAR image classification task.

[0121] Embodiment:

[0122] This embodiment verifies the method of the present application based on the published Pol-SF classification dataset. The dataset contains two full-polarimetric SAR images acquired by different satellites, GF-3 (C-band) and ALOS-2 (L-band), both of which contain complete polarization channels: HH, HV, VH, VV. It should be noted that in the present application, the input data is the polarization coherence matrix converted from the initial scattering matrix of the polarimetric SAR image to the second-order statistical form, and only the upper triangular elements of the polarization coherence matrix are selected. These upper triangular elements are 9 independent elements according to the real and imaginary parts, which constitute the input features of the model. The ALOS-2 image has a spatial resolution of 18 meters and a size of 2900x2700 pixels; the GF-3 image has a spatial resolution of about 5 meters and a size of 2200x1600 pixels. During training, 2000 samples are selected per class per image. The test is divided into two cases. First, the modal missing classification is performed on the full image, because the current available frequency band data can be directly copied and then fused with each other. Second, the post-fusion classification is only tested in the labeled area, because two frequency band data are needed for random sampling fusion. Both images cover five types of ground object classes: water, vegetation, high-density city, low-density city, and developed area, with consistent labels and good cross-frequency semantic comparability.

[0123] According to the structural schematic diagram shown in Figure 2 , the method of the present application is implemented according to the following steps:

[0124] Step 1: As shown in Figure 1 , fixed-size image blocks are independently sampled from the three images in a sliding window manner, and the size of each patch is set to 16x16x9, where 9 represents the 9 independent elements extracted from the upper triangular of the polarization coherence matrix, which contains all the polarization information of the target. A number of sample blocks are randomly sampled from each type of ground object in each frequency band to construct a training sample set that is consistent in class but independent in mode. Due to the different spatial sizes and resolutions of the images, no spatial alignment is performed during the sampling process to preserve the original modal characteristics and texture information.

[0125] Step 2: An independent feature extraction subnetwork is constructed for each frequency band using a multi-stream structure that does not share parameters. The structure of each subnetwork is convolution-convolution-pooling-convolution-convolution-pooling, which contains a total of four convolution layers and two maximum pooling layers, and the specific structure is as follows:

[0126] Layer 1: Convolution kernel size 3x3, channel number 8, step 1, ReLU activation;

[0127] Layer 2: Convolution kernel size 3x3, channel number 8, step 1, ReLU activation;

[0128] Layer 3: Maximum pooling 2x2, step 2;

[0129] Layer 4: convolution kernel size 3×3, number of channels 16, stride 1, ReLU activation;

[0130] Layer 5: convolution kernel size 3×3, number of channels 16, stride 1, ReLU activation;

[0131] Layer 6: max pooling 2×2, stride 2;

[0132] Layer 7: convolution kernel size 3×3, number of channels 32, stride 1, ReLU activation;

[0133] Layer 8: convolution kernel size 3×3, number of channels 32, stride 1, ReLU activation;

[0134] Step 3: Introduce the category-conditional MMD alignment loss during training to constrain the feature distribution consistency of samples of the same category in different frequency bands. 、 is the feature from C-band and L-band, then the category condition MMD loss is:

[0135] (19)

[0136] Among them, the MMD distance is estimated using the Gaussian kernel function:

[0137] (20)

[0138] Step 4: Set the probability for each batch during the training phase Perform Dropout on the modal input to simulate the missing modal condition, thereby improving the model's robustness to incomplete modal input. The specific implementation is as follows:

[0139] (18)

[0140] Step 5: Concatenate the feature maps output by each modality sub-network in the channel dimension to form a joint feature tensor. modalities, and the feature map output by each modal is ,in are the width and height of the feature map, is the uniform number of channels of each modal feature map, then the concatenated features are :

[0141] (twenty one)

[0142] Then, the feature vector is obtained by adaptive average pooling operation ,in The dimension is . Then, the attention weights are computed by a two-layer fully connected network where and are the weight matrices of the two-layer fully connected network, denotes the Sigmoid activation function, which limits the output between 0 and 1. Finally, the attention weights are applied to the joint feature by a channel-wise scaling operation, resulting in the weighted feature where denotes the channel-wise scaling operation. In this way, the model can adaptively emphasize important features and suppress unimportant features, thus improving the discriminative ability of the feature representation.

[0143] (22)

[0144] (23)

[0145] (24)

[0146] where denotes the Sigmoid activation function, denotes the channel-wise scaling.

[0147] The average value and the maximum value of the feature map are computed along the channel dimension. The dimensions of these two feature maps are both where denotes the spatial dimension (width and height) of the feature map. Then, and are concatenated along the channel dimension to form a new feature map, which is then passed through a convolutional layer with filters to generate the spatial attention map with a dimension of . Finally, the spatial attention map is applied to the channel-enhanced feature map by element-wise multiplication, resulting in the fused feature with the same dimension as , which is .

[0148] (25)

[0149] (26)

[0150] (27)

[0151] (28)

[0152] wherein, represents a Sigmoid activation function, used to limit the output of the convolutional layer between 0 and 1, thereby generating attention weights; represents an element-wise multiplication operation.

[0153] Step 6: Fuse the features into the classifier for discrimination, and the classifier includes a three-layer fully connected structure, as follows:

[0154] Fully connected layer 1: input 512 dimensions, output 1024 dimensions, ReLU activation;

[0155] Fully connected layer 2: input 1024 dimensions, output 1024 dimensions, ReLU activation;

[0156] Fully connected layer 3: input 1024 dimensions, output 5 dimensions of class number, Softmax output prediction probability;

[0157] The final prediction label is:

[0158] (29)

[0159] The training uses cross-entropy loss and MMD alignment loss for joint optimization:

[0160] (30)

[0161] Step 7: Set the initial learning rate of the network to 0.001, and decay by 0.1 times after every 30 iterations. The parameters are initialized, the parameter is 0.1, and the parameter is 0.3. The Adam optimizer is used to optimize the model and perform 100 iterations of training.

[0162] The experimental environment used is as follows:

[0163] Hardware environment: Intel Xeon CPU, Quadro GV100 GPU;

[0164] Software environment: Python 3.8, pytorch 1.10.0.

[0165] The model obtained through 100 iterations of training on the validation set is verified, and the model with the lowest validation loss function value is used as the optimal network parameter model. Then, the test set is used to output the detection result to verify the actual detection performance of the model. The test set multi-frequency and multi-polarization SAR image is input into the optimal network parameter model to obtain the classification result.

[0166] As Figure 4The network training curve shown indicates that the model shows good convergence characteristics during the training process. The total loss value decreases steadily from the initial higher level (about 1.4) to close to 0.1 after 100 iterations, showing a typical exponential decay trend. It is worth noting that the curve has two obvious downward inflection points near the 30th and 60th iterations, corresponding to the optimization adjustment of the learning rate decay nodes. The whole training process does not appear to be a sharp shock phenomenon, which verifies the stability of the modal Dropout mechanism and the MMD alignment loss. The loss curve finally converges to a lower plateau, indicating that the design of the combined loss function effectively balances the optimization goals of classification accuracy and feature distribution alignment. In addition, the smooth downward trend of the loss value in the later training period reflects the strong generalization ability of the model, and there is no obvious overfitting sign. The experimental result analysis shows that the method of the present application exhibits excellent classification performance under different test scenarios. From the visual classification results of Figure 5 It can be seen from the visual classification results that whether ALOS-2 (L band) or GF-3 (C band) data, the model still has the ability to recognize the complete structure of ground objects under the condition of single modal missing, especially for the clear boundary description of "high-density urban" and "vegetation" categories. The confusion matrix data shows that under the modal missing scene, the recognition accuracy of the model for multiple categories such as high-density urban area, low-density urban area and developed urban area still maintains more than 90%, verifying the effectiveness of the modal Dropout mechanism. After the fusion of features of each band, in the GF-3 data, the inter-class confusion rate of "vegetation" and "developed area" is reduced from 5.57% in single modal to 1.96%. In the ALOS-2 data, the inter-class confusion rate of "vegetation" and "developed area" is reduced from 7.80% in single modal to 3.28%. It shows that the attention mechanism successfully captures the complementary information across bands. The experimental results fully confirm the robustness and generalization ability of the method under the conditions of modal missing and non-registration.

[0167] Finally, the method of the present application is independently tested on GF-3 and ALOS-2 images, respectively, and can achieve high-precision classification results, verifying the high robustness and adaptability of the method under the conditions of non-uniform frequency bands, modal missing and non-registration.

Claims

1. A multi-band multi-polarimetric SAR image fusion classification method for modality missing and non-registration scenarios, characterized by The method comprises the following steps: Step 1: Data preprocessing: Independently sample polarimetric SAR images of different frequency bands and extract image blocks of fixed size; Step 2: Feature extraction: The image blocks sampled from different frequency bands are input into the corresponding multi-stream sub-network to extract the polarization feature representation of each frequency band; Step 3: Cross-modal constraints: After feature extraction, a category-conditional alignment loss function is introduced at the feature layer to align the feature distributions of the same category in different frequency bands using the maximum mean difference (MMD). Step 4: Modal Dropout: During the training process, a modal dropout mechanism is applied to the input image blocks of each frequency band, that is, the input of a certain frequency band is randomly blocked with a set probability; Step 5: Feature fusion: The feature representation of each frequency band is sent to the modality fusion module, and the features of each frequency band are given adaptive weights through the attention mechanism to obtain the fused feature vector; Step 6: Classification decision: Input the fused feature vector into the classifier and output the corresponding category label to achieve supervised classification of the input image block.

2. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 1 is characterized by The specific steps of step 1 are as follows: Step 1.1: Given a multi-band multi-polarimetric SAR image set , among which The frequency band image is represented as , and is the spatial size of the image, is the number of polarization channels; Step 1.2: From each image Extract a fixed-size image patch from ,in is the patch side length, Indicates the sample number. The patch sampling positions between frequency bands do not correspond to each other, and the spatial positions are completely independent. Step 1.3: To unify the input structure, encode the upper triangular elements of the coherence matrix according to the real and imaginary parts. The input form is constructed as follows: in, is the pixel position in the patch; vector By the coherence matrix The real part of the upper triangular elements of and the imaginary part composition, Indicates the first Rank The elements of the column, for an upper triangular matrix, satisfy .

3. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 2 is characterized by The specific steps of step 2 are as follows: Step 2.1: Construct an independent feature extraction sub-network for each frequency band image, using a parallel multi-stream structure with no parameter sharing. Each sub-network is responsible for extracting the polarization scattering features in the image block of the corresponding frequency band, maintaining its modality independence. Step 2.2: After feature extraction sub-network After that, we get the feature map under this mode : in, are the width and height of the feature map, is the number of channels of the feature map of this frequency band.

4. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 3 is characterized by The specific steps of step 3 are as follows: Step 3.1: To eliminate the statistical differences between feature representations of different frequency bands and enhance cross-modal semantic consistency, a category-conditional feature alignment loss is introduced during the training phase. : in, and Respectively represent the feature representations extracted from the C-band and L-band, Indicates that the sample belongs to Class, a total of categories; represents the square of the maximum mean difference; Step 3.2: Based on the MMD principle, explicitly align the feature distributions of different frequency bands under the same category.

5. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 4 is characterized in that The MMD is implemented in the form of a kernel function, which is defined as follows: in, Represents two distributions and The maximum squared mean difference between and They are respectively and The number of samples in Represents a sample and The kernel function value between and Respectively represent the distribution Sample and distribution and The kernel function value of the sample between.

6. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 4 is characterized by The class-conditional feature alignment loss With the main classification loss Joint optimization to form the total loss function : in, is the weight coefficient.

7. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 4 is characterized in that The specific steps of step 4 are as follows: To enhance the model's robustness to missing frequency bands, a modal dropout mechanism is introduced during training. This mechanism randomly blocks input image blocks from one or more frequency bands with a preset probability in each training batch, simulating the unavailability of certain frequency bands in real-world scenarios. This forces the model to learn cross-modal redundancy and compensation capabilities. in, Represents the first The input image block of frequency bands, whose value is expressed as probability Keep the original input , with probability is set to 0.

8. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 7 is characterized in that The specific steps of step 5 are as follows: Step 5.1: Set up a total modalities, and the feature map output by each modal is ,in are the width and height of the feature map, is the uniform number of channels of each modal feature map, then the concatenated features are: Step 5.2: Combined features Apply channel attention mechanism: First, the feature vector is obtained by adaptive average pooling operation ,in The dimension is ; Then, the attention weight is calculated through a two-layer fully connected network ,in and are the weight matrices of the two-layer fully connected network, represents the Sigmoid activation function; finally, the attention weight is applied to the joint feature through a channel-by-channel scaling operation , and get the weighted features ,in Represents a channel-by-channel scaling operation: Step 5.3: Based on the channel-enhanced features, we further introduce the spatial attention mechanism. By calculating the maximum and average values ​​at each position and concatenating them and inputting them into the convolutional layer, we generate a spatial attention map to emphasize the discriminative spatial regions: First, we calculate The average value in the channel dimension and maximum value ; Then, and Splicing along the channel dimension to form a new feature map, and then pass a The convolutional layer generates a spatial attention map ; Finally, the spatial attention map is multiplied element by element Feature map after channel enhancement , get the fused feature vector : 。 9. The multi-band multi-polarization SAR image fusion classification method for modality missing and non-registration scenarios according to claim 8 is characterized in that The specific steps of step 6 are as follows: Step 6.1: The fused feature vector Input to the classifier to complete the prediction of the image block category: in, are trainable weights, is the bias term, represents the activation function; Step 6.2: Classification probability output by the classifier In, The dimension indicates that the sample belongs to The confidence of the class, the final predicted category is determined by the index corresponding to the maximum probability, that is: Step 6.3: During the training phase, the cross entropy loss function is used to supervise and optimize the predicted results and the true labels. The loss function is defined as follows: in, is the one-hot encoding representation of the true label.

Citation Information

Cited By

  • Lightning proximity prediction method based on multi-source heterogeneous image sequence fusion

    CN121962932A