Feature Divide-and-Conquer Remote Sensing Image Target Recognition Method

Through the method of feature division and conquer, the feature features of most and minority classes in the remote sensing image target recognition task are spatially separated, and the class boundaries are learned in their respective spaces, which solves the problem of insufficient learning of minority classes in the long-tail problem, and improves the robustness and generalization ability of the model.

CN116385890BActive Publication Date: 2025-06-17DALIAN UNIV OF TECH +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210905706.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-06-17
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

There is a long tail problem in the remote sensing image target recognition task, which leads to the model's insufficient learning of a few classes of boundaries and the problems of overfitting and underfitting.

Method used

A method of dividing and conquering features is proposed, separating the feature space of majority and minority classes, and learning specific class boundaries in their respective feature spaces, and gradually separating and refining features through four training stages.

Benefits of technology

It effectively alleviates the model's overfitting of most classes and underfitting of minority classes, improves the model's performance in all categories identification, and enhances the ability to generalize to minority classes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385890B_ABST
    Figure CN116385890B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image information processing, and discloses a remote sensing image target recognition method based on divide-and-conquer of features. Based on decoupling feature learning and classification learning, this method first separates the feature spaces of the majority class and the minority class, and then learns the specific class boundaries in their respective feature spaces, reducing the mutual influence between the majority class and the minority class in a "divide-and-conquer" manner. This method can also be directly embedded into other classification-related visual tasks affected by the long-tail problem and can achieve good classification results. At the same time, this method has good generalization for other classification-related visual tasks and can be used as an effective solution to the long-tail problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image information processing, and in particular proposes a new solution to the long-tail problem commonly existing in remote sensing image target recognition tasks. Background Art

[0002] The long-tail problem mainly refers to the imbalance of samples caused by the long-tail distribution of the training data of the neural network, that is, a small number of categories occupy the vast majority of samples, while a large number of categories have only a small number of samples. At present, in dealing with the long-tail problem of vision-related tasks, the solutions related to this patent mainly cut into the following three levels: data sampling, loss function and learning strategy.

[0003] Common solutions to the long-tail problem from the data sampling level are mainly divided into two categories: oversampling and undersampling. The purpose of oversampling is to increase the number of training samples sampled from the minority class as much as possible, but if only the trivial strategy of copying samples is used to expand the minority class, it is very easy to cause the overfitting problem of the model. Therefore, most oversampling methods focus on the effective expansion of minority class samples. Typical methods include the sample synthesis strategy Borderline-SMOTE proposed by HuiHan et al. in the paper "Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning". It selects minority class samples close to the boundary between the majority and minority classes, and synthesizes samples based on their k-nearest neighbor sample information to improve the learning value of the expanded samples. The purpose of undersampling is to reduce the number of training samples sampled from the majority class as much as possible, but reducing the number of samples will inevitably lead to information loss and may cause the model to fail to learn important features about the majority class. Therefore, undersampling methods focus more on selecting the most comprehensive and representative samples from the majority class samples. Typical methods include the majority class sample selection strategy NearMiss proposed by ManiI et al. in the paper "kNNApproach to Unbalanced Data Distributions: A Case Study involving Information Extraction". The sample with the closest local or global majority class is selected for training based on the Euclidean distance between the two classes of samples. Whether it is oversampling or undersampling, the core idea is to artificially intervene in the data distribution in the data set, which will more or less change the original form of the data. When the number of majority class and minority class samples is very different, sample balancing will lead to a sample distribution that deviates greatly from the actual one, which will cause the model to learn wrong feature information.

[0004] To avoid modifying the original data distribution, another feasible entry point for solving the long-tail problem is to appropriately guide the learning process of the model. The most direct solution is to set different loss weights for different category samples in the design of the network loss function. For example, the loss function EqualizationLoss proposed by Jingru Tan et al. in the literature "Equalization Loss for Long-Tailed Object Recognition" ignores the gradients from the majority class samples for the minority class by introducing a weight term on the basis of the cross-entropy loss, preventing the predictions of the minority class by the model from being overly suppressed by a large number of majority class samples during the learning process. The biggest advantage of the improvement method based on the loss function is its flexibility and convenience, which can be directly applied to complex tasks and various datasets.

[0005] The improvement space of the loss function is very limited by the classification task. In contrast, the model learning strategy is more customizable, and its operation space for the model learning process is also larger. Therefore, more current mainstream research tends to design new model learning strategies to specifically solve the long-tail problem. For example, the strategy based on transfer learning proposed by Xi Yin et al. in the literature "Feature Transfer Learning for Face Recognition with Under-Represented Data" models the majority class samples and the minority class samples separately and transfers the learned knowledge of the majority class to the minority class for use. Another example is the idea of decoupling feature learning and classification learning proposed by Bingyi Kang et al. in the literature "Decoupling Representation and Classifier for Long-Tailed Recognition", which samples normally during the feature learning stage and balances sampling during the classification learning stage, resulting in better long-tail learning results. The implementation of this type of method is relatively complex, but due to its underlying reconstruction of the model learning process, the actual effect improvement is often more significant. Summary of the Invention

[0006] Aiming at the long-tail problem in remote sensing image target recognition tasks, to improve the robustness of remote sensing image target recognition models, alleviate their overfitting to head classes and underfitting to tail classes, and make them perform better in the recognition of all classes, the present invention proposes a remote sensing image target recognition method based on feature divide-and-conquer. Based on decoupling feature learning and classification learning, this method first separates the feature spaces of majority classes and minority classes, and then learns the specific class boundaries in their respective feature spaces to reduce the mutual influence between majority classes and minority classes in a "divide-and-conquer" manner. This method can also be directly embedded into other classification-related vision tasks affected by the long-tail problem and can achieve good classification results.

[0007] The technical solution of the present invention:

[0008] A remote sensing image target recognition method based on feature divide-and-conquer, the steps are as follows:

[0009] Given a training data set with N picture samples where y i is the true label of one of the K classes corresponding to the picture sample x i ; First, define the long-tail distribution of the training data set D; count the number of samples N j contained in each label class in the training data set D, and then calculate the frequency f j =N j / N; Set a threshold λ∈(0,1), for each class j, if f j ≥λ, then consider this class to belong to the majority class; otherwise, consider this class to be a minority class; expand a binary label i for each sample x in the training data set D to represent that the class to which this sample belongs is the majority class, that is or the minority class, that is to obtain the data set where:

[0010]

[0011] The model is based on the ResNet residual network and consists of two main parts: a feature extraction network and a classification head C; among them, the feature extraction network corresponds to the structure of conv1~conv4 in the ResNet network, and downsamples the input RGB image by 16 times to obtain the corresponding feature map representation; the classification head C corresponds to conv5 in the ResNet network and the subsequent average pooling layer and fully connected layer, and converts the feature representation into the probability distribution prediction in the specific classification task; the training of the model is divided into four stages, specifically as follows:

[0012] In the first stage of model training, after the feature extraction network parallel classification heads C1 and C2 are set up, and the dataset D′ is used to perform overall training on C1 and C2;

[0013] C1 is a binary classifier, activated using the sigmoid function and trained with the label as the target ground truth to separate the majority class and the minority class in the feature space; the loss function is binary cross-entropy loss:

[0014]

[0015] where, is the probability of sample x predicted by C1 i for the minority class;

[0016] C2 is a multi-class classifier, activated using the softmax function and constraining the feature extraction with the label y i as the target ground truth; the loss function is multi-class cross-entropy loss:

[0017]

[0018] where, is the true class probability distribution of sample x predicted by C2 i , and each element represents the probability that sample x i belongs to class c; if the true class of sample x i is c, then y ic = 1, otherwise y ic = 0;

[0019] The overall loss L1 of the first-stage network is defined as the weighted combination of L C1 and L C2 :

[0020] L1 = αL C1 + (1 - α)L C2 (4)

[0021] where, α ∈ (0, 1) is the weight coefficient;

[0022] During the training of the first stage, the feature extraction network is encouraged to separate the majority class and the minority class samples in the feature space. However, due to the constraints of the original classification task, such feature separation is likely to be imperfect. When the loss tends to be stable, there are still some difficult majority class and minority class samples that will be mispredicted by the model. These hard-to-separate samples are obstacles in the feature separation process, but at the same time contain important boundary information and need to be further carefully distinguished.

[0023] In the training of the second stage, to further determine the boundary between the majority class and the minority class to better separate them, according to the prediction of the C1 classifier on all training datasets D at the end of the first stage of training, the dataset D′ is divided into two subsets D′ H and D′ T ; where D′ H is the set of all samples in D′ predicted by C1 as the majority class; D′ T is the set of all samples in D′ predicted by C1 as the minority class, and there is D′ H ∪D′ T = D′; in the second stage, the parameters of the feature extraction network will be fixed to the state at the end of the previous stage of training and not updated to fix the feature space; C1 and C2 will be replaced by two independent classification heads C H and C T to finely divide the boundaries of the majority class and the minority class respectively;

[0024] The classifier C H needs to distinguish all the minority class samples misclassified as the majority class in the first stage from the true majority class samples in D′ H , and its loss function is still expressed as binary cross-entropy loss:

[0025]

[0026] where, is the probability prediction of the classifier C H that the sample is a minority class;

[0027] Similarly, the classifier C T needs to distinguish the misclassified majority class and the true minority class samples in D′ T , and its loss function is:

[0028]

[0029] where, is the probability prediction of the classifier C T that the sample is a minority class;

[0030] Since the features of the samples are fixed and the learning of the boundary is decomposed into two independent tasks, after the learning of the second stage, the classifiers C H and C T can fully explore the key information for distinguishing the majority class and the minority class from difficult samples. Therefore, in the training of the third stage, they will reverse-guide the feature extraction network to separate the majority class and the minority class in the feature space.

[0031] The classification head C during the third - stage training H and C T will be fixed to guide the feature extraction network to update the parameters again; all the data in the dataset D′ is used for training, and both classifiers have the same binary classification task; the overall loss in the third stage is:

[0032]

[0033] After the above three - stage training, the feature extraction network will have sufficient ability to globally separate the majority - class and minority - class samples in the feature space, and then learn the specific boundaries of each class in their respective feature spaces. Thus, the problem that when the majority - class and minority - class samples are mixed together, a large number of majority - class samples will cause the classification boundary to deviate towards the majority - class, and the model will ignore learning the boundaries in the minority - class is solved.

[0034] In the final - stage training, the network parameters of the feature extraction network will be fixed again and connected to a multi - classification head C, and C H and C T will be removed; the original dataset D is used to train C to divide the boundaries of the true classes in the separated feature space, and the objective function is the ordinary multi - classification cross - entropy loss:

[0035]

[0036] To obtain better final classification performance, other oversampling and undersampling strategies can be used in the final - stage training. Since the majority - class and minority - class have been separated at the feature level, the noise introduced by oversampling will be significantly reduced; at the same time, the fixed sample features also help to retain as much key information as possible during undersampling.

[0037] Advantages of the present invention: The present invention proposes a method for remote - sensing image target recognition based on feature divide - and - conquer. First, the majority - class and minority - class are separated in the feature space, and then the specific class divisions are learned in their respective spaces, effectively solving the problem that the long - tail effect of the training dataset causes the model to insufficiently learn the boundaries of the minority - class. At the same time, this method has good generalization for other vision tasks related to classification and can be used as an effective solution to the long - tail problem. Brief Description of the Drawings

[0038] Figure 1 It is a schematic diagram of the long - tail distribution of the remote - sensing image dataset.

[0039] Figure 2 It is a flowchart of the method for remote - sensing image target recognition based on feature divide - and - conquer.

[0040] Figure 3 Schematic diagram of the training strategy for the target recognition method of remote sensing images with feature divide-and-conquer.

[0041] Figure 4 Schematic diagram of the effect of the target recognition method of remote sensing images with feature divide-and-conquer. Detailed implementation manners

[0042] The following further elaborates on the detailed implementation manners of the present invention in combination with the technical solutions and the drawings.

[0043] Problem description: The long-tail distribution is widespread in the real world, which stems from the imbalance of the original data distribution. Taking the remote sensing images targeted by the present invention as an example, since the remote sensing images are collected from the real situation on the ground, the number of samples of each category reflects the distribution in the real world: the number of "vehicles" is often much larger than that of "bridges"; the number of "ports" is obviously less than that of "ships". The class imbalance will cause the model to focus on the majority class during the learning process and ignore the learning of the information of the minority class, resulting in the generalization effect of the model on the minority class being much worse than that on the majority class.

[0044] This embodiment is based on a remote sensing image dataset with a long-tail distribution phenomenon as shown in Figure 1 to implement the training of an identification model on it, and finally make the model have good generalization performance in both the majority class and the minority class. The overall process of the method proposed by the present invention is as shown in Figure 2 as follows.

[0045] Model training: First, define the majority class and the minority class according to the frequencies of the samples of each category in the original dataset D, and add additional labels to all samples according to Equation (1). In the dataset used in the embodiment, the majority classes statistically obtained are Vehicle, Ship, Tennis court, etc. Add additional majority class labels to the labels corresponding to the samples of these categories, such as: (index: 114, category: Vehicle, minority class: no); the minority classes are Dam, Windmill, Airplane, Golfcourse, Harbor, etc. Add minority class labels to the samples of these categories, such as: (index: 514, category: Dam, minority class: yes). The updated dataset is denoted as D'. The model training strategy proposed by the present invention is as shown in Figure 3 as follows.

[0046] In the first stage, set two parallel binary classification heads C1 and K classification heads C2 (K = 20 in the embodiment) to connect with the feature extraction network . The training process uses Equation (2) to guide the classification results of C1, aiming to make the feature extraction network Generate majority and minority features that can be binary-classified at the feature level; Equation (3) is used to constrain the classification results of C2, aiming to make the generated features still have discriminability at the true class level. The overall model is trained using Equation (4) as the loss function. At the end of the training in this stage, the prediction results given by C1 are statistically analyzed, and accordingly the dataset D′ is divided into two subsets D′ H and D′ T , where D′ H is the set of all samples in D′ that are predicted by C1 as the majority class; D′ T is the set of all samples in D′ that are predicted by C1 as the minority class.

[0047] In the second stage, remove the classifiers C1 and C2 and freeze the feature extraction network to make the feature space fixed. Train two independent binary classification heads C H and C T on the datasets D′ H and D′ T respectively using Equations (5) and (6), and learn more detailed class boundaries for those difficult samples misclassified by C1 to achieve a highly refined division of the majority and minority classes. At the end of the training in this stage, the classifiers C H and C T have strong discriminative ability for the majority and minority classes, and can inversely guide the distribution of the feature space to achieve the feature separation of the majority and minority class samples.

[0048] In the third stage of training, freeze the classifiers C H and C T learned in the previous stage, and use Equation (7) to train the feature extraction network to fully separate the majority and minority class samples in the feature space.

[0049] In the fourth stage, remove the classifiers C H and C T and freeze again to protect the separated feature space, add a task with a K-classification head to train a multi-classification head C, and use Equation (8) to learn the detailed true class division within the two major categories of the majority and minority classes, and assist in using class-balanced sampling to further improve the model performance.

[0050] Model Inference: After all the training plans in the above four stages are completed, the obtained feature extraction network and the classifier C are combined to form the final model. For the test image x, input it into the model for inference to obtain a K-dimensional output, use the softmax function to convert it into the probability predictions of K classes, and take the class corresponding to the maximum value as the class to which the model predicts the image x belongs.

[0051] The method proposed by the present invention first separates the features of the majority class and the minority class, and then learns the true class boundary in their respective feature spaces, alleviating the bias towards the majority class and the neglect of the minority class when directly classifying in the mixed feature space. The classification effect of the method proposed by the present invention is as Figure 4 shown. This method forces the classifier to learn the class division in the minority class sample space, avoiding the mutual influence between the majority class and the minority class to a certain extent, so as to obtain stronger discrimination ability in the minority class. Under the condition of the same model structure, adopting the training strategy proposed by the present invention can obtain higher recognition accuracy of the minority class and overall recognition accuracy on the test data set compared with the general training strategy.

Claims

1. A method for remote sensing image target recognition with feature divide-and-conquer, characterized in that, The steps are as follows: Given a training dataset with N image samples where y i is the ground truth label of one of the K classes corresponding to the image sample x i ; First, define the long-tailed distribution of the training dataset D; Count the number of samples N contained in each label class in the training dataset D j , and then calculate the frequency f j = N j / N; Set a threshold λ ∈ (0, 1). For each class j, if f j ≥ λ, then this class is considered to belong to the majority class; otherwise, this class is considered to be a minority class. For each sample x in the training dataset D i augment a binary label indicating that the class to which this sample belongs is the majority class, that is or the minority class, that is to obtain the dataset where: The model is based on the ResNet residual network and includes a feature extraction network ε and a classification head C. Among them, the feature extraction network ε corresponds to the structure of conv1 to conv4 in the ResNet network, and downsamples the input RGB image by 16 times to obtain the corresponding feature map representation. The classification head C corresponds to conv5 in the ResNet network and the subsequent average pooling layer and fully connected layer, and converts the feature representation into the probability distribution prediction in a specific classification task. The training of the model is divided into four stages, which are specifically as follows: In the first stage of model training, parallel classification heads C1 and C2 are set after the feature extraction network ε, and the dataset D' is used to perform overall training on ε, C1, and C2. C1 is a binary classifier, activated by the sigmoid function and trained with the label as the ground truth to separate the majority class from the minority class in the feature space; the loss function is binary cross-entropy loss: Among them, is the sample x predicted by C1 i is the probability of the minority class; C2 is a multi-classifier, activated by the softmax function and constrained by the label y i as the target true value for feature extraction; the loss function is multi-class cross-entropy loss: Among them, is the true class probability distribution of the sample x i , and each element represents the probability that the sample x i belongs to the class c; if the true class of the sample x i is c, then y ic = 1, otherwise y ic = 0; The overall loss L1 of the first-stage network is defined as the weighted combination of L C1 and L C2 : L1 = αL C1 +(1 - α)L C2 (4) Among them, α ∈ (0, 1) is the weight coefficient. In the training of the second stage, to further determine the boundary between the majority class and the minority class to better separate them, according to the predictions of the C1 classifier on all training datasets D at the end of the first stage of training, the dataset D' is divided into two subsets D' H and D' T ; where D' H is the set of all samples in D' predicted by C1 as the majority class; D' T is the set of all samples in D' predicted by C1 as the minority class, and there is D' H ∪D' T = D'; in the second stage, the parameters of the feature extraction network ε will be fixed to the state at the end of the previous stage of training without being updated to achieve the fixation of the feature space; C1 and C2 will be replaced by two independent classification heads C H and C T to finely divide the boundaries of the majority class and the minority class respectively; Classifier C H It is necessary to distinguish all the minority class samples misclassified as the majority class in the first stage from the true majority class samples in D' H The loss function is still expressed as binary cross-entropy loss: Among them, is the probability prediction of classifier C H for samples of the minority class; Similarly, classifier C T needs to distinguish the misclassified majority class and the true minority class samples in D' T , and its loss function is as follows: Among them, is the probability prediction of classifier C T for samples belonging to the minority class; The classification head C in the third stage of training H and C T will be fixed to guide the feature extraction network ε to update its parameters again; all the data in the dataset D' are used for training, and both classifiers have the same binary classification task; the overall loss in the third stage is as follows: After the training of the above three stages, the feature extraction network ε will have sufficient ability to globally separate the majority-class and minority-class samples in the feature space. On this basis, it will then learn the specific boundaries of each class in their respective feature spaces. In the final stage of training, the network parameters of the feature extraction network ε will be fixed again and connected to a multi-classification head C, where C H With C T will be removed; the original dataset D is used to train C to divide the boundaries of the true classes in the separated feature space, and the objective function is the ordinary multi-classification cross-entropy loss:

Citation Information

Patent Citations

  • Image salient detecting method based on layering sparse modeling

    CN104240256A

  • Landsat8 and MODIS fusion-construction high space-time resolution data identification autumn grain crop method

    CN104915674A