A Cross-Domain Remote Sensing Scene Classification Method Based on Intermediate Layer Feature Extraction and Nuclear Norm Maximization

By adopting the methods of intermediate layer feature extraction and kernel norm maximization in remote sensing cross-domain scenario classification, the problems of large image differences and inter-class imbalance are solved, and a more efficient cross-domain classification effect is achieved.

CN117079156BActive Publication Date: 2025-06-13BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311078425.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2025-06-13
Estimated Expiration
2043-08-25

AI Technical Summary

Technical Problem

There are problems in the existing remote sensing cross-domain scene classification methods, which are not effective when testing in the target domain.

Method used

The method based on intermediate layer feature extraction and kernel norm maximization is adopted to extract domain invariant features through intermediate feature layers, and the prediction diversity and resolution of unlabeled samples are constrained by kernel norm maximization.

Benefits of technology

It effectively reduces the distribution differences between the source domain and the target domain, improves the accuracy and diversity of cross-domain classification, and overcomes the negative impact of inter-class imbalance problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004413407940000043
    Figure BDA0004413407940000043
  • Figure BDA0004413407940000046
    Figure BDA0004413407940000046
  • Figure BDA0004413407940000047
    Figure BDA0004413407940000047
Patent Text Reader

Abstract

The present invention relates to a cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization, belonging to the field of remote sensing image processing. The method of the present invention adopts a method of randomly extracting intermediate layer features to enhance the discrimination ability of the model in the feature extraction module. By making full use of the intermediate layer features, the key parts in the remote sensing images can be captured, and the distribution difference between the source domain and the target domain can be reduced. A scaling factor is introduced in the method of the present invention to make the prediction output closer to the ideal state in the hypothesis, and nuclear norm maximization is used to solve the problem of reduced prediction diversity of unlabeled samples caused by the imbalance of the number of samples between categories in the remote sensing dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross - domain remote sensing scene classification method based on intermediate - layer feature extraction and nuclear norm maximization, belonging to the field of remote sensing image processing. Background Art

[0002] With the progress of satellite and remote sensing technologies, the number of remote sensing images is increasing rapidly. Scene classification is a method that can effectively process remote sensing images, aiming to classify remote sensing images into different semantic categories. Remote sensing scene classification plays an important role in fields such as urban planning and geological disaster detection. However, due to differences in geographical distribution, imaging conditions, sensors, etc., the data distributions of different remote sensing datasets vary greatly. To adapt to the distribution differences between different datasets, domain - adaptation methods have been proposed. Domain - adaptation methods include a source domain and a target domain. Due to the data - distribution differences between different domains, it is usually difficult to obtain satisfactory results when testing a model trained on the source domain in the target domain. In domain adaptation, the data of the source domain and the target domain are mapped to the same feature space to minimize the distribution difference between the two domains, enabling the target domain to make full use of the rich information in the source domain.

[0003] Currently, deep learning plays an important role in the field of DA. Convolutional neural network (CNN) is one of the most representative models in deep - learning methods. CNN usually requires a large amount of labeled data during the training process. Although the number of available remote sensing images has increased significantly, labeling these images not only relies on professional knowledge but also consumes a large amount of human resources, which is usually uneconomical. To reduce the dependence on labeled data, unsupervised learning has been introduced. Compared with supervised or semi - supervised methods, the target domain in unsupervised domain - adaptation (UDA) technology does not contain labeled samples but learns relevant information from the labeled source domain.

[0004] Unsupervised domain - adaptation methods can generally be divided into two categories. The first category is the statistics - based method, which uses the mean value or high - order moments to measure the distribution difference between domains and minimizes the statistical metric to align different domains. The second type is the method based on generative adversarial networks. However, due to the complex features of remote sensing images, it is difficult for artificially designed statistical metrics to characterize complex feature - distribution information. Therefore, researchers mainly focus on adversarial - learning - based methods. DANN first introduced generative adversarial networks into the field of transfer learning. In DANN, a domain discriminator is used to distinguish whether a sample comes from the source domain or the target domain, while a feature extractor extracts features that are difficult for the domain discriminator to distinguish. When the domain discriminator and the feature extractor reach a dynamic balance, the feature distributions of the source domain and the target domain are aligned.

[0005] In traditional adversarial training methods, usually only the features of the last layer output by the feature extractor are selected to represent the image features. However, due to differences in factors such as the shooting satellite and location, the key parts of remote sensing images vary greatly. Therefore, it is not enough to only extract the deep features of the images. At the same time, in the same remote sensing dataset, the number of images in different categories may be unbalanced. In this case, the currently used method based on entropy minimization has side effects. It tends to judge the category with fewer samples as the category with more samples, which will reduce the prediction diversity of unlabeled data. Summary of the Invention

[0006] The technical problem to be solved by the present invention is: to overcome the deficiencies of the prior art, that is, the problems of large differences in images and unbalanced numbers of images between categories in the current remote sensing cross-domain scene classification method, and to propose a cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization. In this method, intermediate feature layer extraction can effectively extract domain-invariant features by making full use of intermediate layer features, so as to better achieve cross-domain classification. Nuclear norm maximization can effectively constrain the prediction diversity and discriminability of unlabeled samples by using the nuclear norm of the predicted output of the target domain, so as to better achieve cross-domain classification.

[0007] The technical solution of the present invention is:

[0008] A cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization, the steps of the method include:

[0009] The first step is to obtain two different remote sensing datasets, and select the images in the common classes of the two remote sensing datasets as the objects for scene classification. One of the remote sensing datasets is used as the source domain dataset, and the other remote sensing dataset is used as the target domain dataset. The source domain dataset is fed into the feature extractor F to obtain the source domain feature map f S , and the target domain dataset is fed into the feature extractor F to obtain the target domain feature map f T ; Since the source domain dataset and the target domain dataset are fed into the same feature extractor F, the number of layers of the obtained source domain feature map and target domain feature map is equal;

[0010] The source domain feature map includes N source domain intermediate layer feature maps and the source domain last layer feature map of the feature extractor F;

[0011] The target domain feature map includes N target domain intermediate layer feature maps and the target domain last layer feature map of the feature extractor F;

[0012] The second step is to randomly select n source domain intermediate layer feature maps f from the N source domain intermediate layer feature maps obtained in the first step using the iterative random selection method i S(i ∈ n), select n target domain intermediate feature maps with the same number of layers as the randomly selected source domain intermediate feature maps Send the source domain intermediate feature map f i S and the source domain last layer feature map to the class classifier G, and respectively obtain the corresponding source domain prediction representations and Send the target domain intermediate feature map and the target domain last layer feature map to the class classifier G, and respectively obtain the corresponding target domain prediction representations and

[0013] Thirdly, concatenate the source domain intermediate feature map f i S obtained in the second step with its corresponding source domain prediction representation in a multi-linear mapping manner to obtain the concatenation result of the source domain intermediate layer Send the source domain last layer feature map and its corresponding source domain prediction representation in a multi-linear mapping manner to obtain the concatenation result of the source domain last layer Send the target domain intermediate feature map and its corresponding target domain prediction representation in a multi-linear mapping manner to obtain the concatenation result of the target domain intermediate layer Send the target domain last layer feature map and its corresponding target domain prediction representation in a multi-linear mapping manner to obtain the concatenation result of the target domain last layer

[0014] Concatenate the source domain intermediate layer concatenation result and the target domain intermediate layer concatenation result to obtain h i and send it to the domain classifier D. At the same time, concatenate the source domain last layer concatenation result and the target domain last layer concatenation result to obtain h final and send it to the domain classifier D to obtain the domain classification result;

[0015] Fourthly, obtain the classification loss function according to the source domain prediction representation obtained in the second step, and according to the target domain prediction representation Obtain the nuclear norm maximization loss function, obtain the intermediate layer feature extraction loss function according to the domain classification results obtained in the third step, and use the classification loss function, the nuclear norm maximization loss function, and the intermediate layer feature extraction loss function to optimize the feature extractor, the class classifier, and the domain classifier simultaneously to complete cross-domain remote sensing scene classification.

[0016] In the first step, the source domain dataset is denoted as containing N S labeled samples, and the target domain dataset is denoted as containing N T unlabeled samples;

[0017] In the second step, the iterative random selection method means randomly extracting intermediate layer features in each training cycle;

[0018] In the third step, the source domain intermediate layer cascade result is denoted as:

[0019]

[0020] where represents the i-th source domain cascade result, i ∈ n;

[0021] The source domain last layer cascade result is denoted as:

[0022]

[0023] The target domain intermediate layer cascade result is denoted as:

[0024]

[0025] where represents the j-th target domain cascade result, j ∈ n;

[0026] The target domain last layer cascade result is denoted as:

[0027]

[0028] In the fourth step, the classification loss function is denoted as:

[0029]

[0030] where K represents the number of classes, represents the true value probability of the source domain samples;

[0031] The nuclear norm maximization loss function is denoted as:

[0032]

[0033] Among them, denotes calculating the corresponding nuclear norm for the output matrix of each batch, and P η denotes the output matrix of the improved domain classifier;

[0034] The output matrix of the improved domain classifier is expressed as:

[0035]

[0036]

[0037] η denotes the scaling factor;

[0038] The loss function of the intermediate layer feature extraction is expressed as:

[0039]

[0040] Among them, d l denotes the domain label. When l = 0, it means the corresponding sample comes from the source domain, and when l = 1, it means the corresponding sample comes from the target domain; λ denotes the fixed hyperparameter, and n denotes the number of randomly selected intermediate layer feature maps.

[0041] Beneficial effects

[0042] A cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization according to the present invention. Compared with deep features, the shallow features of an image have a smaller receptive field and higher resolution, enabling the network to capture detailed information in the image. At the same time, the method of selecting intermediate layer features also has a significant impact on performance. Therefore, a method of randomly extracting intermediate layer features is adopted to enhance the discrimination ability of the model in the feature extraction module. By making full use of intermediate layer features, the key parts in remote sensing images can be captured, reducing the distribution difference between the source domain and the target domain.

[0043] Regarding the problem of the imbalance in the number of samples between categories, in an ideal state, the F-norm and rank of the output matrix can be used to measure the prediction discriminability and diversity of unlabeled samples, thereby reducing the impact of the reduction in prediction diversity caused by the imbalance in the number of samples between categories in remote sensing images. However, the actual training results often deviate from the ideal state. Therefore, a scaling factor is introduced to make the prediction output closer to the ideal state in the hypothesis, and nuclear norm maximization is used to solve the problem of the reduction in the prediction diversity of unlabeled samples caused by the imbalance in the number of samples between categories in remote sensing datasets. Specific implementation manners

[0044] The present invention will be further described below in conjunction with embodiments.

[0045] Embodiment

[0046] A cross - domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization. Taking the actual application scenario of cross - domain remote sensing scene classification as an example, the steps of this method are as follows:

[0047] In the first step, obtain two remote sensing datasets AID and NWPU, and select images in 12 common classes (airport, anchorage, beach, dense, farm, overpass, forest, stadium, parking lot, river, sparse, and storage) in the datasets as the objects for scene classification. Take the AID dataset as the source domain dataset and the NWPU dataset as the target domain dataset. Use resent50 as the feature extractor, and send the source domain dataset into the feature extractor F to obtain the source domain feature map f S , and send the target domain dataset into the feature extractor F to obtain the target domain feature map f T ; Since the source domain dataset and the target domain dataset are sent into the same feature extractor F, the number of layers of the obtained source domain feature map and target domain feature map is equal;

[0048] The source domain feature map includes 3 source domain intermediate layer feature maps and the source domain last layer feature map of the feature extractor F;

[0049] The target domain feature map includes 3 target domain intermediate layer feature maps and the target domain last layer feature map of the feature extractor F;

[0050] In the second step, from the 3 source domain intermediate layer feature maps obtained in the first step, use the iterative random selection method to select 2 source domain intermediate layer feature maps f i S (i ∈ [1, 2, 3]), and select 2 target domain intermediate layer feature maps with the same number of layers as the randomly selected source domain intermediate layer feature maps Send the source domain intermediate layer feature map f i S and the source domain last layer feature map into the class classifier G to obtain the corresponding source domain prediction representations and Send the target domain intermediate layer feature map and the target domain last layer feature map into the class classifier G to obtain the corresponding target domain prediction representations and

[0051] In the third step, concatenate the source domain intermediate layer feature map f i S obtained in the second step with its corresponding source domain prediction representation in a multi - linear mapping way to obtain the concatenation result of the source domain intermediate layer Send the source domain last layer feature map and its corresponding source domain prediction representation They are concatenated in the way of multilinear mapping to obtain the concatenation result of the last layer of the source domain The intermediate layer feature map of the target domain and its corresponding target domain prediction representation They are concatenated in the way of multilinear mapping to obtain the concatenation result of the intermediate layer of the target domain The last layer feature map of the target domain and its corresponding target domain prediction representation They are concatenated in the way of multilinear mapping to obtain the concatenation result of the last layer of the target domain

[0052] The concatenation result of the intermediate layer of the source domain and the concatenation result of the intermediate layer of the target domain are spliced to obtain h i and sent to the domain classifier D. At the same time, the concatenation result of the last layer of the source domain and the concatenation result of the last layer of the target domain are spliced to obtain h final and sent to the domain classifier D to obtain the domain classification result;

[0053] In the fourth step, according to the source domain prediction representation obtained in the second step the classification loss function is obtained. According to the target domain prediction representation the nuclear norm maximization loss function is obtained. According to the domain classification result obtained in the third step, the intermediate layer feature extraction loss function is obtained, and the feature extractor, class classifier and domain classifier are optimized simultaneously using the classification loss function, nuclear norm maximization loss function and intermediate layer feature extraction loss function to complete cross-domain remote sensing scene classification.

[0054] In the first step, the source domain dataset is denoted as containing N S labeled samples, and the target domain dataset is denoted as containing N T unlabeled samples;

[0055] In the second step, the iterative random selection method means randomly extracting the intermediate layer features in each training cycle;

[0056] In the third step, the concatenation result of the intermediate layer of the source domain is denoted as:

[0057]

[0058] where represents the i-th source domain concatenation result, i ∈ [1, 2, 3];

[0059] The cascading result of the last layer of the source domain It is expressed as:

[0060]

[0061] The cascading result of the middle layer of the target domain is expressed as:

[0062]

[0063] Among them, represents the cascading result of the j-th target domain, j ∈ [1, 2, 3];

[0064] The cascading result of the last layer of the target domain It is expressed as:

[0065]

[0066] In the fourth step described above, the classification loss function is expressed as:

[0067]

[0068] Among them, K represents the number of categories, where K = 12, represents the true value probability of the source domain samples;

[0069] The nuclear norm maximization loss function is expressed as:

[0070]

[0071] Among them, represents calculating the corresponding nuclear norm of the output matrix for each batch, P η represents the output matrix of the improved domain classifier;

[0072] The output matrix of the improved domain classifier is expressed as:

[0073]

[0074]

[0075] η represents the scaling factor, taking η = 0.75;

[0076] The intermediate layer feature extraction loss function is expressed as:

[0077]

[0078] Among them, d l represents the domain label. When l = 0, it means the corresponding sample comes from the source domain, and when l = 1, it means the corresponding sample comes from the target domain; λ represents the fixed hyperparameter, taking λ = 1, and n represents the number of randomly selected intermediate layer feature maps, where n = 2 here.

[0079] Taking the actual application scenario of cross-domain remote sensing scene classification as an example, on multiple public remote sensing data sets, comparing various existing unsupervised remote sensing cross-domain classification methods, as shown in Table 1, it can be seen that the method proposed in the embodiment has better classification accuracy.

[0080] Table 1 Classification accuracies of various existing cross-domain classification methods and the proposed method on six cross-domain classification tasks formed by three data sets of AID (abbreviated as A), UCM (abbreviated as U), and NWPU (abbreviated as N)

[0081]

[0082] In summary, the above are the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization, characterized in that the steps of this method include: First, obtain two different remote sensing data sets, and select the images in the common classes of the two remote sensing data sets as the objects for scene classification. Use one remote sensing data set as the source domain data set and the other remote sensing data set as the target domain data set. Feed the source domain data set into the feature extractor F to obtain the source domain feature map f S , and feed the target domain data set into the feature extractor F to obtain the target domain feature map f T ; The source domain feature map includes N source domain intermediate layer feature maps and the source domain last layer feature map of the feature extractor F; The target domain feature map includes N target domain intermediate layer feature maps and the target domain last layer feature map of the feature extractor F; Step 2: Randomly select n source domain intermediate layer feature maps f from the N source domain intermediate layer feature maps obtained in Step 1 i S , where i ∈ n, and select n target domain intermediate layer feature maps where j ∈ n; Input the source domain intermediate layer feature map f i S and the source domain last layer feature map into the class classifier G to obtain the corresponding source domain prediction representations and Input the target domain intermediate layer feature map and the target domain last layer feature map into the class classifier G to obtain the corresponding target domain prediction representations and Step 3: The source domain intermediate layer feature map obtained in Step 2 and its corresponding source domain prediction representation are concatenated in a multilinear mapping manner to obtain the concatenation result of the source domain intermediate layer The source domain last layer feature map and its corresponding source domain prediction representation are concatenated in a multilinear mapping manner to obtain the concatenation result of the source domain last layer The target domain intermediate layer feature map and its corresponding target domain prediction representation are concatenated in a multilinear mapping manner to obtain the concatenation result of the target domain intermediate layer The target domain last layer feature map and its corresponding target domain prediction representation are concatenated in a multilinear mapping manner to obtain the concatenation result of the target domain last layer The concatenation result of the intermediate layer in the source domain and the concatenation result of the intermediate layer in the target domain are concatenated to obtain h i and sent to the domain classifier D. At the same time, the concatenation result of the last layer in the source domain and the concatenation result of the last layer in the target domain are concatenated to obtain h final and sent to the domain classifier D to obtain the domain classification result; Step 4: According to the source domain prediction representation obtained in Step 2 obtain the classification loss function, and according to the target domain prediction representation obtain the nuclear norm maximization loss function, obtain the intermediate layer feature extraction loss function according to the domain classification result obtained in Step 3, and simultaneously optimize the feature extractor, class classifier, and domain classifier using the classification loss function, nuclear norm maximization loss function, and intermediate layer feature extraction loss function to complete cross-domain remote sensing scene classification.

2. The cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization according to claim 1, characterized in that: In the first step, the source domain dataset is denoted as containing N S labeled samples, and the target domain dataset is denoted as containing N T unlabeled samples.

3. The cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization according to claim 1 or 2, characterized in that: In the second step, an iterative random selection method is used to select the source domain intermediate layer feature maps.

4. The cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization according to claim 3, characterized in that: In the second step, the iterative random selection method means randomly extracting the intermediate layer features in each training cycle.

5. The cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization according to claim 2, characterized in that: In the third step, the source domain intermediate layer cascade result is expressed as: Among them, represents the cascading result of the i-th source domain, where i ∈ n; The cascade result of the last layer in the source domain It is expressed as: The target domain intermediate layer cascade result is expressed as: Among them, represents the j-th target domain cascade result, where j ∈ n; The cascading result of the last layer of the target domain It is expressed as:

6. The cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization according to claim 5, characterized in that: In the fourth step, the classification loss function is expressed as: where K represents the number of categories, represents the true value probability of the source domain samples.

7. The cross-domain remote sensing scene classification method based on intermediate layer feature extraction and nuclear norm maximization according to claim 6, characterized in that: In the fourth step, the nuclear norm maximization loss function is expressed as: Among them, denotes calculating the corresponding nuclear norm for the output matrix of each batch, and P η denotes the output matrix of the improved domain classifier; The output matrix of the improved domain classifier is expressed as: η represents a scaling factor.