A SAR Interpretability Feature Extraction and Classification Method Based on Decoupled Autoencoders

Through the method based on the decoupled autoencoder, the interpretable features of SAR images are extracted using sliding window sampling and Transformer encoder, and the problems of insufficient feature extraction accuracy and high complexity in the prior art are solved, and efficient feature extraction and classification effects are achieved.

CN116824154BActive Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310722752.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-19
Publication Date
2025-07-25
Estimated Expiration
2043-06-19

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy, high time complexity and difficult to explain features in SAR image feature extraction. The existing methods of decoupling characterization to generate adversarial networks are complex in construction and low training efficiency, making it difficult to be used for downstream classification tasks.

Method used

Using a decoupled autoencoder-based method, a decoupled autoencoder is established by establishing a scattered feature decoupled autoencoder, a sliding window sampling is used to construct a label-free training set, and interpretable features are extracted through the Transformer encoder and the multi-head attention mechanism, and feature fusion is performed by combining grouping convolution and cross attention mechanisms to train the classifier to improve the interpretability and decoupling of features.

Benefits of technology

It realizes efficient interpretable extraction and classification of SAR image features, improves the interpretability and decoupling of features, enhances the classification effect of downstream tasks, and reduces training complexity and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004291649980000051
    Figure BDA0004291649980000051
  • Figure BDA0004291649980000061
    Figure BDA0004291649980000061
  • Figure BDA0004291649980000091
    Figure BDA0004291649980000091
Patent Text Reader

Abstract

The present invention relates to a SAR interpretable feature extraction and classification method based on a decoupled autoencoder. First, a scattering feature decoupled autoencoder model is established, and a sliding window sampling is used to simply and effectively construct an unlabeled training set. The decoupled autoencoder model is trained to extract interpretable features of SAR images with the scattering feature decoupled autoencoder. Since the method of splitting and masking image patches is adopted, the training efficiency of the decoupled network is accelerated. The introduction of the scattering feature decoupling layer maps the abstract semantic features reflecting the scattering characteristics to a specified feature subspace. The features obtained by the scattering feature decoupled autoencoder have strong interpretability and decoupling property, which is beneficial to downstream classification tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the SAR data recognition of autoencoders, and relates to a method for SAR interpretable feature extraction and classification based on a decoupled autoencoder, specifically to a method for SAR interpretable feature extraction and classification based on a scattering feature decoupled autoencoder. Background Art

[0002] SAR (Synthetic Aperture Radar) is a high-resolution imaging technology that can detect targets in the atmosphere while extracting information such as their position, size, and shape. However, there are some problems with traditional SAR image feature extraction methods, such as:

[0003] 1. The accuracy of feature extraction is limited by the algorithm itself and cannot capture all details;

[0004] 2. The time complexity of feature extraction is relatively high, resulting in too long processing time;

[0005] 3. The results of feature extraction are difficult to interpret and it is difficult to explain the meaning of the image to non-technical personnel.

[0006] To solve the above problems, in recent years, Generative Adversarial Networks (GANs) have been widely used in the interpretable feature extraction of SAR images. A GAN consists of two networks, a generator network and a discriminator network. The generator network generates new samples based on the given training data, and the discriminator network classifies real samples according to the features of the samples. By training the adversarial relationship between the two networks, the generator network gradually learns the features of new samples, thus generating more realistic images.

[0007] Decoupled representation is a commonly used method that can separate the representation process of the generative adversarial network from the generation process, thereby eliminating the influence of the learning of the adversarial network on the original data on feature extraction. In the interpretable feature extraction of SAR images, a decoupled representation network can be used for feature extraction of SAR images, while a generative adversarial network can be used for image generation.

[0008] The Decouple representation Network (DCN) consists of two networks, one for feature extraction of SAR images and one for generating SAR images. The DCN first extracts features from the original SAR image, and then uses the generator network to generate a new SAR image. The generator network accepts the output of the decoupled representation network as input and generates a new SAR image. The DCN decouples during the feature extraction and generation processes, thereby eliminating the influence of the learning of the adversarial network on the original data on feature extraction and generation.

[0009] The feature extraction network (Input representation Network, ICN) is used to extract features from the original SAR image. The ICN first uses the DCN to extract the features of the original image, and then uses these features to generate a new SAR image. The output of the ICN is used as the input of the generation network.

[0010] The Generative Adversarial Network (GAN) is used to generate new SAR images. The GAN consists of two networks, a generation network and a discriminant network. The generation network generates new samples based on the given training data, and the discriminant network classifies real samples according to the features of the samples. By training the adversarial relationship between the two networks, the generation network gradually learns the features of new samples, thereby generating more realistic SAR images.

[0011] The Decouple representation Network (DCN) is used to extract features from the original SAR image, and the Generative Adversarial Network (GAN) is used to generate new SAR images. The DCN decouples feature extraction and generation, eliminating the influence of the learning of the adversarial network on the original data on feature extraction and generation, thereby generating more realistic SAR images and being interpretable at the same time.

[0012] As in the patent application 202210616508.9, a method for interpretable feature extraction of SAR images based on a decoupled representation generative adversarial network includes the following steps: Step 1, initialize the model and parameters of the decoupled representation generative adversarial network; Step 2, construct a polarimetric SAR data set and randomly divide it into a training set and a test set; Step 3, train the scattering feature generator of the decoupled representation generative adversarial network to decouple multi-dimensional interpretable scattering component features from the polarimetric coherence matrix; Step 4, train the scattering feature discriminator of the decoupled representation generative adversarial network; Step 5, use the test set to test the trained decoupled representation generative adversarial network to obtain interpretable depth features reflecting the target polarimetric scattering characteristics of the test set SAR images. This method is relatively complex for the construction of the data set and has low training efficiency for the network. Secondly, the interpretable features obtained by this method are relatively specific scattering features and are difficult to be used as effective semantic features for downstream classification tasks. Summary of the Invention

[0013] Technical Problems to be Solved

[0014] In order to avoid the deficiencies of the prior art, the present invention proposes a method for interpretable feature extraction and classification of SAR based on a decoupled autoencoder.

[0015] Technical Solution

[0016] A method for SAR interpretable feature extraction and classification based on a decoupled autoencoder, characterized in that the steps are as follows:

[0017] Step 1: Establish a scattering feature decoupled autoencoder f s , and initialize the network layer parameter weights of the decoupled autoencoder f s , and set the masking rate mr and learning rate lr parameters of the input polarization SAR data T matrix;

[0018] Step 2: Construct an unlabeled sample training set T on the polarization SAR dataset u , and where is the training sample

[0019] Use the unlabeled sample training set T u to train the decoupled autoencoder. The steps are as follows:

[0020] Step 201: Divide the unlabeled training dataset samples into N p image patches:

[0021] N p = s c / s p

[0022] where: s p represents the size of the image patch;

[0023] Mask the divided image patches according to the masking rate mr and record the masking position matrix mask to obtain L um unmasked image patches:

[0024] L um = (1 - mr) * N p

[0025] Step 202: Encode the unmasked image patches with the feature embedding layer, and then add sine position encoding to endow the image patches with position information, thereby obtaining the encoded feature tokens of the unmasked polarization SAR data

[0026] Step 203: Input the encoded L um feature vector tokens into the Transformer encoder f u-enc to capture the context dependencies from the unmasked polarization SAR data and obtain the hidden layer feature tokens T latent :

[0027] T latent = f u-enc (Tum )

[0028] Step 204: Set L m learned feature tokens and insert them into the set of L um hidden layer feature tokens according to the masking positions mask, and then add sine position encoding to endow the layer of feature tokens with global position information, obtaining the updated set of feature tokens T um-m ;

[0029] Step 205: Input the feature token t i ∈T um-m into the Transformer encoder f T-rec , and this layer attempts to reconstruct the T matrix information at the masked positions based on the correlations between the unmasked feature tokens and the spatial context relationship with the masked tokens, and finally obtains the set of feature tokens T containing the global T matrix information global :

[0030] T global = f T-rec (T um-m )

[0031] Step 206: Input the set of feature tokens T global into the scattered feature decoupling layer, and map the input token features to N scatter feature subspaces through the multi-head attention mechanism. Each feature subspace represents a polarization scattering characteristic of the polarimetric SAR data. Through multiple attention heads of the scattered feature decoupling layer f dec , N scatter potential representations of polarization scattering characteristics are obtained The feature dimension is expressed as [N scatter , N p +1, D s , representing the number of types of scattering features, the number of feature tokens, and the feature dimension of the potential representation of polarization scattering characteristics respectively;

[0032] where represents a component of the potential representation of a polarization scattering of each feature token, enhancing the interpretable attribute of the representation;

[0033] Step 207: Perform dimensional transformation on the decoupled potential representation of polarization scattering P scatter to obtain features with dimensions of [C scatter , h p , w p , where N p = h p ·w p , C scatter = N scatter ·Ds ;

[0034] Subsequently, it is input into the scattering feature grouping prediction head to obtain the corresponding polarized SAR decomposition scattering components from the potential representations of each scattering characteristic. After dimensional conversion, the scattering components corresponding to each token are obtained.

[0035] The prediction head f pre is composed of grouped convolutions in the decoupled autoencoder;

[0036] Step 208: Calculate the scattering feature reconstruction loss and update the parameters of the decoupled autoencoder network layer;

[0037] The reconstruction loss is defined as follows:

[0038]

[0039] where t i , respectively represent the true polarized scattering feature token and the scattering feature token generated by the decoupled autoencoder, and t i is transformed from the true polarized scattering feature F scatter by dimensional transformation, that is, the loss function only calculates the mean squared error value of the tokens at the masked positions;

[0040] Step 209: Iteratively train until the maximum number of iteration batches, and save the weight parameter W of the scattering feature decoupled autoencoder. d ;

[0041] Step 3: Fine-tune the backbone of the scattering feature decoupled autoencoder and train the classifier:

[0042] Step 301: Use the scattering feature decoupled autoencoder after removing the grouping prediction head as the interpretable feature extraction backbone f b of the classification model, load the weight parameter W saved in Step 2 d , and initialize the classifier f cls and its parameters and learning rate;

[0043] Step 302: For the labeled polarized SAR data, randomly select the same number of samples from C categories as the training set;

[0044] Step 303: Perform two random data augmentations on each batch of N training samples to obtain 2N augmented data;

[0045] Step 304: Input the randomly augmented SAR data into the interpretable feature extraction backbone f b after image patch feature embedding to obtain the hidden layer feature token T containing context global semantic information. latentand the potential representation P containing polarization scattering characteristics scatter , and this step does not require randomly masking image patches;

[0046] Step 305: The potential representation P containing polarization scattering characteristics scatter is subjected to feature fusion with the hidden layer feature token feature T through the feature embedding layer latent to embed the polarization scattering characteristics into the abstract features with high-level semantics, obtaining the fused feature Z;

[0047] Step 306: Input the fused feature Z into the classifier to obtain the predicted scores for each category, calculate the loss function, and update the weight parameters of the classification model;

[0048] The loss function consists of the cross-entropy loss L ce and the feature contrast loss L con and is defined as follows:

[0049]

[0050] The cross-entropy loss is used to reduce the difference between the true probability distribution of the sample categories and the predicted category probability distribution of the classifier; the feature contrast loss makes the feature distances of the same category in the enhanced data decrease and the features of different categories move away from each other, which is used to train a more robust classifier; represents the number of samples of the i-th category in a batch of training samples, and z i represents the feature of a sample in a batch of samples, that is, the feature vector input to the classifier.

[0051] The decoupled autoencoder includes the Transformer encoder f u-enc , the scattering feature decoupling layer f based on the multi-head attention mechanism T-rec and the prediction head f composed of grouped convolutions pre .

[0052] The construction of the unlabeled sample training set T u is as follows: Sliding window sampling is performed on the polarimetric SAR dataset with a window of size s c to obtain the polarimetric T matrix as the training sample thus constructing the unlabeled sample training set T u , where

[0053] The prediction head f pre is composed of one layer of grouped convolution, and the parameter group is set to N scatter , that is, each polarimetric SAR scattering classification is only inferred from its respective potential representation features, enabling the feature decoupling layer to have the representational ability of interpretable polarization characteristics.

[0054] The data augmentation methods in step 303 include random horizontal flipping, random vertical flipping, random selection, random cropping, and randomly adding Gaussian noise.

[0055] In step 305: The feature embedding layer consists of a cross-attention mechanism, taking the tokens in P scatter as the query vector Q, taking the N latent tokens in the hidden layer feature T p as the key vector K and value vector V, calculating the attention score A, and weighting the hidden layer feature T latent through calculating the attention score A to obtain the fused feature Z.

[0056] Beneficial effects

[0057] A SAR interpretable feature extraction and classification method based on a decoupled autoencoder proposed by the present invention first establishes a scattering feature decoupled autoencoder model. Using sliding window sampling can simply and effectively construct an unlabeled training set. The decoupled autoencoder model is trained to extract interpretable features of SAR images by the scattering feature decoupled autoencoder. Since the way of slicing-masking image patches is adopted, the training efficiency of the decoupled network is accelerated. The introduction of the scattering feature decoupling layer maps the abstract semantic features reflecting the scattering characteristics to the specified feature subspace. The features obtained by the scattering feature decoupled autoencoder have strong interpretability and decoupling, which is beneficial to the downstream classification task. Specific implementation manners

[0058] The present invention will be further described in combination with embodiments as follows:

[0059] Using the data of the present embodiment on the Flevoland polarimetric SAR dataset, the method of the present invention is used for feature extraction and classification, and the steps are as follows:

[0060] Step 1. Establish a scattering feature decoupled autoencoder model and prepare a dataset:

[0061] The decoupled autoencoder includes a Transformer encoder f u-enc , a scattering feature decoupling layer f T-rec based on the multi-head attention mechanism, and a prediction head f pre composed of grouped convolutions.

[0062] Step 101. Initialize the network layer parameter weights of the decoupled autoencoder f s , and set parameters such as the masking rate mr and the learning rate lr of the input polarimetric SAR data T matrix;

[0063] Step 102. Sample the polarimetric SAR data T matrix T c with an image size of s from the polarimetric SAR dataset through a sliding windowu , construct an unlabeled sample training set for the self-decoupling encoder;

[0064] Step 2. Train the scattering feature decoupling autoencoder to extract interpretable features of SAR images:

[0065] Step 201. Divide the samples of the unlabeled training data set into N p image patches, where

[0066] N p = s c / s p

[0067] s p represents the size of the image patch. Subsequently, mask the divided image patches according to the masking rate mr and record the masking position matrix mask to obtain L um unmasked image patches, where:

[0068] L um = (1 - mr) * N p

[0069] Step 202. The unmasked image patches will be encoded by the feature embedding layer, and then sinusoidal position encoding will be added to endow the image patches with position information, thereby obtaining the encoded feature tokens of the unmasked polarimetric SAR data

[0070] Step 203. Input the encoded L um feature vector tokens into the Transformer encoder f u-enc , capture the context dependencies from the unmasked polarimetric SAR data, and obtain the hidden layer feature tokens T latent , where

[0071] T latent = f u-enc (T um )

[0072] Step 204. Set L m learnable feature tokens, and insert them into the set of L um hidden layer feature tokens according to the masking position mask, and then add sinusoidal position encoding to endow the layer of feature tokens with global position information, obtaining the updated set of feature tokens T um-m ;

[0073] Step 205. Input the feature token t i ∈ T um-m into the Transformer encoder f T-rec, this layer attempts to reconstruct the T-matrix information at the masked positions based on the correlations between unmasked feature tokens and the spatial context relationship with the masked tokens, and finally obtains a set of feature tokens T containing global T-matrix information global , where

[0074] T global = f T-rec (T um-m )

[0075] Step 206: Input the set of feature tokens T global into the scattering feature decoupling layer, and map the input token features to N scatter feature subspaces through the multi-head attention mechanism. Each feature subspace represents a polarization scattering characteristic of the polarimetric SAR data, that is, through the multiple attention heads of the scattering feature decoupling layer f dec N scatter potential representations of polarization scattering characteristics can be obtained The feature dimension can be expressed as [N scatter , N p + 1, D s , representing the number of types of scattering features, the number of feature tokens, and the feature dimension of the potential representation of polarization scattering characteristics respectively. Among them represents a component of the potential representation of a polarization scattering of each feature token, enhancing the interpretability attribute of the representation;

[0076] Step 207: Perform dimensionality conversion on the decoupled potential representation P of polarization scattering scatter to obtain features with a dimension of [C scatter , h p , w p , where N p = h p ·w p , C scatter = N scatter ·D s . Subsequently, input it into the scattering feature grouping prediction head to infer the corresponding polarimetric SAR decomposition scattering components from the potential representations of each scattering characteristic After dimensionality conversion, the scattering component corresponding to each token is obtained To ensure that there is no interference between the potential representations of each scattering characteristic and avoid the decoupling of features, the prediction head f pre consists of a layer of grouped convolution, and the parameter group is set to N scatter , that is, each polarimetric SAR scattering classification is inferred only from its own potential representation features, enabling the scattering feature decoupling layer to have the representation ability of interpretable polarization characteristics;

[0077] Step 208: Calculate the scattering feature reconstruction loss and update the parameters of the decoupled autoencoder network layer. The reconstruction loss is defined as follows:

[0078]

[0079] where t i and represent the true polarization scattering feature token and the scattering feature token generated by the decoupled autoencoder respectively, where t i is transformed from the true polarization scattering feature F scatter through dimensional transformation, that is, the loss function only calculates the mean square error value of the tokens at the masked positions.

[0080] Step 209: Iteratively train until the maximum number of iterations, and save the weight parameter W d ;

[0081] Step 3: Fine-tune the backbone of the scattering feature decoupled autoencoder and train the classifier;

[0082] Step 301: Use the scattering feature decoupled autoencoder after removing the grouped prediction head as the interpretable feature extraction backbone f b of the classification model, load the weight parameter W d saved in Step 2, and initialize the classifier f cls and its parameters and learning rate;

[0083] Step 302: For the labeled polarimetric SAR data, randomly select the same number of a small number of samples from C categories as the training set;

[0084] Step 303: Perform two random data augmentations on N training samples in each batch to obtain 2N augmented data. The data augmentation methods include random horizontal flipping, random vertical flipping, random selection, random cropping, and random addition of Gaussian noise;

[0085] Step 304: Input the randomly augmented SAR data into the interpretable feature extraction backbone f b after image patch feature embedding to obtain the hidden layer feature token T latent containing context global semantic information and the latent representation P scatter containing polarimetric scattering characteristics. There is no need to randomly mask the image patches in this step;

[0086] Step 305: Perform feature fusion on the latent representation P scatter containing polarimetric scattering characteristics and the hidden layer feature token feature T latent through the feature embedding layer to embed the polarimetric scattering characteristics into the abstract features with high-level semantics; specifically, the feature embedding layer consists of a cross-attention mechanism. We use P scatterThe token in it is used as the query vector Q, and the hidden layer feature T latent Among the N p tokens are used as the key vector K and the value vector V, so as to calculate the attention score A. By calculating the attention score A, the hidden layer feature T latent is weighted to obtain the fused feature Z;

[0087] Step 306: Input the fused feature Z into the classifier to obtain the prediction scores for each category, calculate the loss function, and update the weight parameters of the classification model. The loss function is defined as follows:

[0088]

[0089] This loss consists of the cross-entropy loss L ce and the feature contrast loss L con The cross-entropy loss is used to reduce the difference between the true probability distribution of the sample categories and the predicted category probability distribution of the classifier; the feature contrast loss makes the feature distances of the same category in the enhanced data decrease, and the features of different categories move away, which is used to train a more robust classifier; represents the number of samples of the i-th category in a batch of training samples, and z i represents the feature of a sample in a batch of samples, that is, the feature vector input to the classifier.

[0090] The proposed method is compared with the supervised learning method based on ViT and the baseline method (self-supervised learning method based on the masked autoencoder MAE), and the overall accuracy (OA) is used as the classification accuracy index. The effectiveness and superiority of the method we proposed are verified.

[0091] In the table, the classification accuracies of the comparison methods and the proposed method under different numbers of samples are listed. It can be seen that the method we proposed is completely superior to other classification methods. Our method uses the scattering feature decoupling autoencoder, so that the hidden layer features of the feature extraction backbone have richer feature expressions, thus being superior to the classification effect of MAE. In addition, after adding the feature embedding module, the polarization characteristics of the feature tokens are enriched, further widening the feature gap between different category samples. And when a small number of samples are given for fine-tuning, due to the introduction of the contrast learning idea, the classification effect of the proposed method has a more prominent performance.

[0092]

Claims

1. A SAR interpretable feature extraction and classification method based on a decoupled autoencoder, characterized in that The steps are as follows: Step 1: Establish the scattering feature decoupled autoencoder f s , and initialize the network layer parameter weights of the decoupled autoencoder f s , and set the masking rate mr and learning rate lr parameters of the input polarization SAR data T matrix; Step 2: Construct an unlabeled sample training set T on the polarimetric SAR dataset u , and where is a training sample Using an unlabeled sample training set T u Training the decoupled autoencoder, the steps are as follows: Step 201: Split the unlabeled training dataset samples into N p image patches: N p = s c / s p where: s p represents the size of the image block; Mask the segmented image patches according to the masking rate mr and record the masking position matrix mask to obtain L um unmasked image patches: L um = (1 - mr) * N p Step 202: Encode the unmasked image patches with the feature embedding layer, and then add sine position encoding thereto to endow the image patches with position information, thereby obtaining the feature tokens after encoding the unmasked polarimetric SAR data Step 203: Input the L um encoded feature vector tokens into the Transformer encoder f u-enc to capture context dependencies from the unmasked polarimetric SAR data and obtain the hidden layer feature tokens T latent : T latent = f u-enc (T um ) Step 204: Set L m learned feature tokens, and insert them into the set of L um hidden layer feature tokens according to the masking positions mask, and then add sine position encoding to endow the layer of feature tokens with global position information, obtaining the updated set of feature tokens T um-m ; Step 205: Input the feature token t i ∈ T um-m into the Transformer encoder f T-rec which attempts to reconstruct the T-matrix information at the masked positions based on the correlations between unmasked feature tokens and the spatial context relationship with masked tokens, and finally obtains the set of feature tokens T global : T global = f T-rec (T um-m ) Step 206: Input the feature token set T global into the scattering feature decoupling layer, and map the input token features to N scatter feature subspaces through the multi-head attention mechanism. Each feature subspace represents a polarization scattering characteristic of the polarimetric SAR data. Through multiple attention heads of the scattering feature decoupling layer f dec obtain N scatter latent representations of polarization scattering characteristics The feature dimension is expressed as [N scatter , N p +1, D s , representing the number of types of scattering features, the number of feature tokens, and the feature dimension of the latent representation of polarization scattering characteristics, respectively; Among them represents a polarization scattering potential characterization component of each feature token, enhancing the interpretable property of the characterization; Step 207: Perform dimensionality conversion on the decoupled polarization scattering potential representation P scatter to obtain features with dimensions of [C scatter , h p , w p , where N p = h p · w p , C scatter = N scatter · D s ; Subsequently, it is input into the scattering feature grouping prediction head to obtain the corresponding polarized SAR decomposition scattering components from the potential representations of each scattering characteristic After dimensional transformation, the scattering components corresponding to each token are obtained The prediction head f pre is composed of grouped convolutions in the decoupled autoencoder; Step 208: Calculate the reconstruction loss of the scattering features and update the parameters of the decoupled autoencoder network layer; The reconstruction loss is defined as follows: Among them, t i , respectively represent the true polarization scattering feature token and the scattering feature token generated by the decoupled autoencoder. t i is transformed from the true polarization scattering feature F scatter through dimensional transformation, that is, the loss function only calculates the mean square error value of the tokens at the masked positions; Step 209: Iteratively train until the maximum number of iterations, and save the weight parameter W of the scattering feature decoupling autoencoder d ; Step 3. Fine-tune the backbone of the scattering feature decoupled autoencoder and train the classifier: Step 301: Use the scattering feature decoupling autoencoder after removing the grouped prediction head as the interpretable feature extraction backbone f of the classification model b , and load the weight parameter W saved in Step 2 d , and initialize the classifier f cls and its parameters and learning rate; Step 302. For the labeled polarimetric SAR data, randomly select the same number of samples from C categories as the training set; Step 303: Perform random data augmentation on N training samples in each batch twice to obtain 2N augmented data; Step 304: Input the randomly augmented SAR data into the interpretable feature extraction backbone f after image patch feature embedding b to obtain the hidden layer feature tokens T containing context global semantic information latent and the latent representation P containing polarimetric scattering characteristics scatter , and there is no need to randomly mask image patches in this step; Step 305: Incorporate the potential representation P of polarization scattering characteristics scatter with the hidden layer feature token feature T through the feature embedding layer latent to perform feature fusion, embed the polarization scattering characteristics into the abstract features with high-level semantics, and obtain the fused feature Z; Step 306: Input the fused feature Z into the classifier to obtain the predicted scores for each category, calculate the loss function, and update the weight parameters of the classification model; The loss function consists of the cross-entropy loss L ce and the feature contrast loss L con and is defined as follows: The cross-entropy loss is used to reduce the difference between the true probability distribution of sample categories and the predicted probability distribution of classifier categories; the feature contrast loss reduces the distance between features of the same category and separates features of different categories in the enhanced data, which is used to train a more robust classifier; represents the number of samples of the i-th category in a batch of training samples, z i represents the feature of a sample in a batch of samples, that is, the feature vector input to the classifier.

2. The SAR interpretable feature extraction and classification method based on the decoupled autoencoder according to claim 1, wherein: The described decoupled autoencoder includes a Transformer encoder f u-enc , a scattering feature decoupling layer f based on the multi-head attention mechanism T-rec , and a prediction head f composed of grouped convolutions pre .

3. The SAR interpretable feature extraction and classification method based on a decoupled autoencoder according to claim 1, wherein: Constructing the unlabeled sample training set T u is as follows: sliding window sampling is performed on the polarimetric SAR dataset with a window of size s c to obtain the polarimetric T matrix as training samples so as to construct the unlabeled sample training set T u , where 4. The SAR interpretable feature extraction and classification method based on a decoupled autoencoder according to claim 1, characterized in that: The prediction head f pre is composed of one layer of grouped convolution, and the parameter group is set to N scatter , that is, each polarization SAR scattering classification is only inferred from its respective potential characterization features, so that the feature decoupling layer has the characterization ability of interpretable polarization characteristics.

5. The SAR interpretable feature extraction and classification method based on a decoupled autoencoder according to claim 1, wherein: The data augmentation methods in Step 303 include random horizontal flipping, random vertical flipping, random selection, random cropping, and randomly adding Gaussian noise.

6. The SAR interpretable feature extraction and classification method based on a decoupled autoencoder according to claim 1, characterized in that: In step 305: The feature embedding layer consists of a cross-attention mechanism, taking the tokens in P scatter as the query vector Q, taking the N latent tokens in the hidden layer feature T p as the key vector K and the value vector V, calculating the attention score A, and weighting the hidden layer feature T latent by calculating the attention score A to obtain the fused feature Z.

Citation Information

Patent Citations

  • Interpretable feature extraction method for SAR images based on decoupled representation generative adversarial network

    CN114898159B

  • Scattering energy and stack self-code-based polarimetric SAR image classification method

    CN107563420A

  • SAR image interpretability feature extraction method based on decoupling representation generative adversarial network

    CN114898159A