Feature Classification Method for Category-Asymmetric Cross-Domain Hyperspectral Images

Through two-stage network training and the definition of feature alignment loss function, the problem of unsatisfactory classification effect in the cross-domain spectral image classification task with category asymmetry is solved, and higher classification accuracy and robustness are achieved.

CN118781394BActive Publication Date: 2025-06-24HARBIN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410772151.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-16
Publication Date
2025-06-24
Estimated Expiration
2044-06-16

AI Technical Summary

Technical Problem

The prior art is difficult to achieve ideal classification effects in class asymmetric cross-domain spectral image classification tasks, because the target training set may not cover the complete category set.

Method used

Through two stages of network training, effective alignment of features of different domains is achieved. The first stage is pre-trained a neural network, defining the category cohesive loss function and the predicted loss function, and extracting the inter-class relationship between source domain samples. The second stage is to train the main neural network, define the sample alignment loss function, category alignment loss function and feature alignment loss function, and update the main neural network parameters to adapt to the target domain data.

Benefits of technology

The classification accuracy of the target domain dataset is improved, especially in the absence of the target domain category, showing higher label prediction performance and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118781394B_ABST
    Figure CN118781394B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for classifying ground objects for category-asymmetric cross-domain spectral images, belonging to the technical field of ground object classification of spectral images. The method of the present invention realizes effective alignment of features in different domains through network training in two stages. In the first stage, a neural network is pre-trained, and an intra-class cohesion loss function and a prediction loss function are constructed to extract the inter-class relationships in the source domain, and then the class centers in the source domain are calculated to provide a basis for subsequent steps. In the second stage, a main neural network is trained. For the symmetric categories shared by the source domain and the target domain, a sample alignment loss function and a class alignment loss function are constructed to ensure the alignment of the sample distributions and class centers between the two domains; for the asymmetric categories existing in the source domain but missing in the target domain, by constructing a feature alignment loss function, the features extracted by the main neural network are aligned with the features extracted by the pre-trained neural network. Finally, the target domain test set is fed into the trained main neural network to obtain the classification labels of the target domain dataset. The implementation results on the publicly available cross-regional spectral dataset show that, compared with existing methods, this method has higher classification accuracy and more robust performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of spectral image land cover classification, and particularly relates to a land cover classification method for class-asymmetric cross-domain spectral images. Background Art

[0002] In the field of spectral image land cover classification, the task of using the labeled land cover class samples in the source domain to classify the unlabeled but similar land covers in the target domain is called the cross-domain land cover classification task. In recent years, although numerous classification methods have been applied to this field, they generally based on an assumption that the source domain and the target domain have a completely symmetric set of classes. However, this assumption imposes great limitations on practical applications because the target training set may not cover the complete set of classes due to land cover or manual labeling errors. In this context, traditional classification methods often fail to achieve ideal classification results, and this challenge is called class-asymmetric cross-domain spectral image classification. To address this challenge, the present invention proposes a land cover classification method aimed at solving the class-asymmetric cross-domain spectral image land cover classification task. Summary of the Invention

[0003] To solve the above challenges, the present invention provides a land cover classification method for class-asymmetric cross-domain spectral images. The method of the present invention realizes effective alignment of features in different domains through two-stage network training. The method includes the steps:

[0004] Load the source domain dataset and the target domain dataset, and divide the target domain dataset into a training set and a test set.

[0005] In the first stage of the method, a separate neural network is pre-trained, and an intra-class cohesion loss function and a prediction loss function are defined to make the samples of the same class in the source domain close to each other, so as to effectively extract the inter-class relationship between the source domain samples.

[0006] Calculate the class centers of the source domain using the pre-trained neural network.

[0007] In the second stage of the method, the main neural network is trained. For the symmetric classes of the source domain and the target domain, a sample alignment loss function and a class alignment loss function are defined.

[0008] For the asymmetric classes of the source domain and the target domain, a feature alignment loss function is defined to update the parameters of the main neural network.

[0009] Send the target domain test set into the trained main neural network to obtain the classification labels of the target domain dataset.

[0010] Further, the loaded source domain dataset and target domain dataset are such that the source domain data is a spectral data sample with ground object class labels collected at a certain geographical location, and the target domain data is an unclassified spectral data sample without ground object class labels collected at another different geographical location. The target domain dataset is further divided into a training set and a test set. The training set is set to have incomplete ground object classes, thus simulating the situation where the target domain training set and the source domain dataset have asymmetric classes, while the test set contains complete ground object classes and is used to evaluate the accuracy of the trained network model in the classification task.

[0011] Further, in the first stage of the method, a neural network is pre-trained, and a within-class cohesion loss function is defined. The formulated loss function is as follows:

[0012]

[0013] Where, and are respectively the i-th input sample in the source domain and its corresponding label, are the parameters of the pre-trained neural network, D dis is the Euclidean distance, represents the predicted output of the pre-trained neural network for the input sample, C represents the total number of classes in the source domain data, and Λ(a, b) represents 1 if a equals b and 0 otherwise. The within-class cohesion loss function makes samples of the same class close to each other and samples of different classes far from each other.

[0014] Further, the prediction loss function is defined. The formulated loss function is as follows:

[0015]

[0016] The prediction loss function is used to calculate the gap between the predicted output of the pre-trained neural network on the source domain and the true label, and by minimizing this gap, the label prediction performance of the network for the source domain data is improved.

[0017] The within-class cohesion loss function and the prediction loss function are combined to obtain the total loss function in the first stage. The formulated loss function is as follows:

[0018]

[0019] Further, the class center of the source domain is calculated using the pre-trained neural network, providing a basis for subsequent steps. The class center Cen S of the source domain is expressed as:

[0020]

[0021] Where, n srepresents the number of samples of the source domain data, C bal represents the number of symmetric classes between the source domain and the target domain.

[0022] Furthermore, in the second stage of the method, a main neural network is trained. For the symmetric classes of the source domain and the target domain, a sample alignment loss function and a class alignment loss function are constructed. The formulated sample alignment loss function is as follows:

[0023]

[0024] where is the j-th input sample in the target domain, are the parameters of the main neural network, represents the predicted output of the main neural network for the input sample, n t represents the number of samples of the target domain data. The sample alignment loss function ensures that the sample distributions between the two domains are aligned with each other.

[0025] The class alignment loss function ensures that the class centers between the two domains are aligned with each other. The formulated class alignment loss function is as follows:

[0026]

[0027] where D cos (a, b) represents the cosine similarity between a and b, Cen T is the class center of the target domain, expressed as:

[0028]

[0029] Furthermore, for the asymmetric classes that exist in the source domain but are missing in the target domain, first, a pre-trained neural network is used to extract features for all classes in the source domain, and then the main neural network is used to extract features for these classes again. By constructing a feature alignment loss function, the features extracted by the main neural network are aligned with the features extracted by the pre-trained network. The expression of the formulated feature alignment loss function is as follows:

[0030]

[0031] If the feature alignment is successful, then the target features should also be aligned with the source features. Combining the sample alignment loss function, the class alignment loss function, and the feature alignment loss function can obtain the total loss function in the second stage. The formulated loss function is as follows:

[0032]

[0033] where λ is the hyperparameter of the feature alignment loss.

[0034] Finally, the target domain test set containing all categories is fed into the main neural network to obtain the classification labels of the target domain data set. Compared with the existing methods, the method of the present invention has higher classification accuracy and more robust performance.

[0035] The specific advantages are as follows:

[0036] The embodiments of the present invention adopt a new neural network model. After two-stage training with careful design, this neural network not only shows excellent label prediction ability for symmetric categories in the target domain, but also shows excellent label prediction performance for missing categories in the target domain that are difficult to handle by traditional network models. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the exemplary embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the exemplary embodiments of the invention or the prior art.

[0038] By referring to the detailed description below and reading in conjunction with the drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention can be better understood. Several embodiments of the present invention are shown in the drawings, which are for illustration only and not for limitation, where:

[0039] Figure 1 is a flowchart of a method for land cover classification of category-asymmetric cross-domain spectral images provided by the present invention;

[0040] Figure 2 is a schematic diagram of the cross-domain hyperspectral image used in the exemplary method; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the specific embodiments and with reference to the drawings. It should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0042] Exemplary Method

[0043] As Figure 1 shown, after the processing flow starts, step S110 is first executed.

[0044] Step S110: Load the data sets, including the source domain data set with manually labeled land cover category labels and the target domain data set without land cover category labels to be classified, where the target domain data set is divided into a training set and a test set.

[0045] Step S120: Pretrain a separate neural network and define the intra-class cohesion loss function and the prediction loss function.

[0046] As an example, the expression of the intra-class cohesion loss function defined in step S120 is as follows:

[0047]

[0048] Where, and are the i-th input sample and its corresponding label in the source domain respectively, are the parameters of the pre-trained neural network, D dis is the Euclidean distance, represents the predicted output of the pre-trained neural network for the input sample, C represents the total number of classes in the source domain data, and Λ(a, b) represents 1 if a equals b and 0 otherwise. The intra-class cohesion loss function makes samples of the same class close to each other and samples of different classes far from each other.

[0049] As an example, the expression of the prediction loss function defined in step S120 is as follows:

[0050]

[0051] The prediction loss function is used to calculate the gap between the predicted output of the pre-trained neural network on the source domain and the true label, and by minimizing this gap, the label prediction performance of the network for the source domain data is improved.

[0052] By combining the intra-class cohesion loss function and the prediction loss function, the total loss function in the first stage can be obtained. As an example, the expression of the total loss function in the first stage is as follows:

[0053]

[0054] By the loss function, it can be ensured that the pre-trained neural network can effectively extract the inter-class correlation in the source domain.

[0055] Step S130: Use the neural network pre-trained in S120 to calculate the class centers of the source domain, providing a basis for subsequent steps. As an example, the class center Cen S of the source domain is expressed as:

[0056]

[0057] Step S140: Train the main neural network, and define the sample alignment loss function and the class alignment loss function for the symmetric classes of the source domain and the target domain.

[0058] As an example, the expression of the defined sample alignment loss function is as follows:

[0059]

[0060] Wherein, is the j-th input sample in the target domain, are the parameters of the main neural network, represents the predicted output of the main neural network for the input sample, and n t represents the number of samples in the target domain data. The sample alignment loss function ensures that the sample distributions between the two domains are aligned with each other.

[0061] As an example, the expression of the defined class alignment loss function is as follows:

[0062]

[0063] Wherein, D cos (a, b) represents the cosine similarity between a and b, and Cen T is the class center of the target domain, which is expressed as:

[0064]

[0065] The class alignment loss function ensures that the class centers between the two domains are aligned with each other.

[0066] Step S150: For the asymmetric classes in the source domain and the target domain, first use the pre-trained neural network in the first stage to extract features for all classes in the source domain, and then use the main neural network trained in the second stage to extract features for these classes again. Define a feature alignment loss function to align the features extracted by the main neural network with the features extracted by the pre-trained network.

[0067] As an example, the feature alignment loss function defined in step S150 is expressed as follows:

[0068]

[0069] If the source domain feature alignment is successful, then the target features should also be aligned with the source features. Combining the sample alignment, class alignment, and feature alignment loss functions can obtain the total loss function in the second stage.

[0070] As an example, the expression of the total loss function in the second stage is as follows:

[0071]

[0072] Wherein, λ is the hyperparameter of the feature alignment loss.

[0073] Step S160: Finally, the target domain test set containing all categories is sent into the main neural network trained in S140 to obtain the classification labels of the target domain data set.

[0074] Results of the specific implementation

[0075] This implementation uses the Pavia data set. The details of the data set are described as follows:

[0076] The Pavia data set contains two hyperspectral images: Pavia U and Pavia C, both of which are acquired using a ROSIS sensor. The data set of Pavia U includes 103 spectral bands with a size of 610x340 pixels, and the Pavia C data set includes 102 spectral bands with an image size of 1096x715 pixels. Both images contain 7 different land cover categories. Table 1 shows the numbers and names of the experimental classes and the number of class samples in two different regional scenarios. To verify the superiority of this implementation (Ours), this implementation is compared with several existing methods, including methods such as DANN and TST. The classification accuracies of these methods for the above-mentioned publicly available data sets will be compared. The specific data comparisons are shown in Table 2.

[0077] Table 1 Numbers of source samples and target samples in the Pavia data set

[0078]

[0079] Table 2 Classification results of the Pavia data set (evaluation metric: Overall accuracy)

[0080]

[0081] From the data comparison in the above table, it can be clearly seen that as the number of missing categories increases, this method can still achieve good classification results, and significantly improves the classification performance compared with other methods.

[0082] This embodiment proposes a method for land cover classification of category-asymmetric cross-domain spectral images. After two stages of training, this method not only demonstrates excellent label prediction ability for symmetric categories in the target domain, but also shows excellent label prediction performance for missing categories in the target domain that are difficult to handle by traditional network models. The results on the publicly available test data set demonstrate the superiority of this implementation.

[0083] It should be understood that the above specific implementation manners of the present invention are only used for exemplary illustration or explanation of the principles of the present invention, and do not constitute a limitation to the present invention. Therefore, the present invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A method for classifying objects in asymmetric cross-domain spectral images, characterized in that: The method comprises the steps of: S110. Divide the spectral image dataset into a source domain dataset and a target domain dataset, wherein the target domain dataset is divided into a training set and a test set; S120. Pre-train a neural network, construct a category cohesion loss function and a prediction loss function, and update the parameters of the pre-trained neural network by minimizing the two loss functions; S130. Calculate the category center of the source domain using the pre-trained neural network in S120; S140. Train a main neural network, construct a sample alignment loss function and a category alignment loss function for the symmetric categories shared by the source domain and the target domain, and update the parameters of the main neural network by minimizing the above two loss functions; S150. For asymmetric categories that are present in the source domain but missing in the target domain, use the pre-trained neural network in S120 and the main neural network in S140 to extract features from all categories in the source domain, and construct a feature loss function, and update the parameters of the main neural network in S140 by minimizing the feature loss function; S160. Send the target domain test set containing all categories to the main neural network in S140, and finally obtain the classification label of the target domain data set; In the S110, the spectral image dataset is divided into a source domain dataset and a target domain dataset, the source domain data is a spectral data sample with a ground object category label collected at a certain geographical location, and the target domain data is a spectral data sample to be classified without a ground object category label collected at another different geographical location. In order to more realistically simulate the scenario of missing category data due to ground object coverage or labeling errors in actual applications, the target domain dataset is further divided into a training set and a test set, wherein the training set is set to have incomplete ground object categories, thereby simulating the situation where the target domain training set and the source domain dataset have asymmetric categories, and the test set contains complete ground object categories, which is used to evaluate the accuracy of the trained network model in the classification task; In S120, a neural network is pre-trained to define a category cohesion loss function and a prediction loss function. The expression of the category cohesion loss function is as follows: in, and are the i-th input sample in the source domain and its corresponding label, are the parameters of the pre-trained neural network, D dis is the Euclidean distance, represents the predicted output of the pre-trained neural network for the input sample, C represents the total number of categories of the source domain data, and Λ(a,b) means that if a is equal to b, the output is 1, otherwise the output is 0; The expression defining the prediction loss function is as follows: By combining the category cohesion loss function and the prediction loss function, the total loss function of the first stage is expressed as: In S130, the source domain category center is calculated using a pre-trained neural network. S It is expressed as: Among them, n s Represents the number of samples of source domain data; In S140, for the balanced categories shared by the source domain and the target domain in the network training phase, a sample alignment loss function and a category alignment loss function are respectively constructed, and the expression of the sample alignment loss function is as follows: in, is the jth input sample in the target domain, are the parameters of the main neural network, represents the predicted output of the main neural network for the input sample, C bal Indicates the number of categories that are symmetric between the source domain and the target domain, n t Represents the number of samples of target domain data; The expression of the category alignment loss function is as follows: Among them, D cos (a,b) represents the cosine similarity between a and b, Cen T The category center of the target domain is expressed as: In S150, for the imbalanced categories that the source domain has but the target domain lacks during the network training phase, the pre-trained neural network of S120 is first used to extract features from all categories in the source domain, and then the main neural network trained in S140 is used to extract features from these categories again, and a feature alignment loss function is constructed. The expression of the feature alignment loss function is as follows: By combining the sample alignment loss function, the category alignment loss function, and the feature alignment loss function, the total loss function of the second stage is obtained: Among them, λ is the hyperparameter of feature alignment loss; In S160, the target domain test set containing all categories is sent to the main neural network trained in S140, and finally the classification label of the target domain data set is obtained.

Citation Information

Patent Citations

  • Hyperspectral image cross-domain classification method based on multi-level feature alignment

    CN116188830A

  • Dual-drive feature learning method for cross-region spectral image terrain classification

    CN116935121A