Remote sensing scene classification method based on sort learning
By constructing a ranking classification network model, utilizing ranking learning and deep learning techniques, introducing virtual neutral samples, and optimizing the ranking relationship, the problem of weak classification ability of rare objects in remote sensing data is solved, and unified processing of single-label and multi-label classification is achieved, thereby improving the model accuracy and efficiency.
Patent Information
- Application Number
- CN202510815550.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
AI Technical Summary
In remote sensing datasets, the scarcity of rare land object samples leads to weak model classification capabilities. Existing technologies cannot effectively utilize image features, and single-label and multi-label classification tasks need to be designed and trained separately, which increases resource waste and management complexity.
Construct a ranking classification network model, introduce virtual neutral samples through ranking learning and deep learning technology, calculate label ranking loss, sample ranking loss and feature ranking loss, optimize the ranking relationship, perform end-to-end training and evaluation, and uniformly handle single-label and multi-label classification tasks.
It improves the overall classification performance and accuracy of the model, reduces the workload of switching between different classification tasks, improves the classification ability of rare landforms, and reduces resource waste and management complexity.
Smart Images

Figure CN120673258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing scene classification, and more specifically, to a remote sensing scene classification method based on ranking learning and deep learning. Background Art
[0002] Remote sensing datasets often have abundant samples of common feature categories (such as roads and buildings), while rare features (such as specific mineral distribution areas and endangered species habitats) are scarce, resulting in a long-tail distribution. This data imbalance causes the model to tend to learn features of high-frequency categories during training, making it less capable of extracting and classifying features from minority categories. This leads to an imbalance in overall classification performance, making it difficult to meet the requirements for comprehensive and accurate monitoring of various features.
[0003] Single-label and multi-label classification tasks typically require the design, training, and maintenance of two separate models. From data collection and preprocessing to model tuning, each step requires significant human, material, and time investment. This duplication of effort not only results in a significant waste of resources but also increases the complexity of model management. Furthermore, single-label classification focuses on extracting key features that distinguish between different categories, while multi-label classification requires capturing correlations between labels and the features of multiple objects in an image. Existing technologies treat these two tasks separately, resulting in an inability to fully utilize common image features during feature extraction and learning. Summary of the Invention
[0004] The purpose of the present invention is to overcome the problems of the existing technology and provide a remote sensing scene classification method based on ranking learning. With the help of ranking learning and deep learning technology, a ranking classification network model is constructed, which can simultaneously solve the two tasks of single-label classification and multi-label classification, and uniformly perform end-to-end training and evaluation, reducing the workload of switching different classification tasks.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A remote sensing scene classification method based on ranking learning includes the following steps:
[0007] S1: Construct a sorting classification network model to simultaneously process single-label classification tasks and multi-label classification tasks of remote sensing images;
[0008] S2: Introducing virtual neutral samples, the predicted values of virtual neutral samples are constrained to the interval between the positive label prediction value and the negative label prediction value of the image;
[0009] S3: Combined with virtual neutral samples, calculate the label ranking loss, sample ranking loss and feature ranking loss, and optimize the label ranking relationship, sample ranking relationship and feature ranking relationship;
[0010] S4: Calculate the approximate normalized discounted cumulative gain loss and optimize the ranking loss of the ranking classification network model;
[0011] S5: Combine multiple ranking losses, train through multi-task learning, evaluate the ranking classification network model, and apply the trained ranking classification network model to remote sensing scene classification.
[0012] Furthermore, a sorting and classification network model is constructed, specifically:
[0013] The ranking classification network model framework consists of three parts: label ranking, sample ranking, and feature ranking. It uses multiple network structures as feature extractors, introduces virtual neutral samples, and performs training and evaluation in an end-to-end architecture.
[0014] Let the label space be The label of image x is called the positive label of x, and the remaining labels are called the negative labels of x; assuming that each small batch contains M training images, the mth image is x m , and record x m The positive label set is The negative label set is
[0015] Input x m , the sorting classification network model generates N prediction values z m1 ,z m2 ,...,z mN ∈[0,1], where z mn Represents the positive label set of the mth image Contains the label value y n probability; expectation When z mn >0.5, and When z mn <0.5;
[0016] join in A virtual neutral label is created, and the value of θ is set to limit the probability range of the neutral label. For a given θ∈(0,0.5), the "virtual prediction value" of the virtual neutral label is set to fall within the interval [0.5-θ,0.5+θ], and an "interval interval" is established between the positive and negative labels. When z mn >0.5+θ; When z mn <0.5-θ.
[0017] Furthermore, combined with the virtual neutral samples, the label ranking loss is calculated as follows:
[0018] The set of neutral labels is For any given and θ∈(0,0.5), let the vector in
[0019]
[0020] Let vector in
[0021]
[0022] The calculation formula for label ranking loss is
[0023]
[0024] Where Φ is the sorting optimization function, is the label prediction vector, is the virtual neutral label vector.
[0025] Furthermore, combined with the virtual neutral samples, the sample ranking loss is calculated as follows:
[0026] If y is the positive label of image x, then x is called a positive sample of y, otherwise x is called a negative sample of y; in a small batch, if Then x m y n Positive samples of
[0027] M training samples generate M predicted values: z 1n ,z 2n ,...,z Mn ; Sort the corresponding samples in descending order of predicted values, and insert between positive and negative samples Virtual neutral samples;
[0028] For any given Let vector in
[0029]
[0030] Furthermore, let the vector in
[0031]
[0032] The sample ranking loss calculation formula is:
[0033]
[0034] Where, is the sample prediction vector; is the virtual neutral sample vector.
[0035] Furthermore, combined with the virtual neutral samples, the feature ranking loss is calculated as follows:
[0036] Let v m =[v m1 ,v m2 ,...,v m(m-1) ,v m(m+1) ,...,v mM ], where v mm′ For collection and The similarity coefficient, v mm′ Measured the mth image x m and the m′th image x m′ The label similarity of represents the positive label set of the m′th image;
[0037] The calculation formula of similarity coefficient is:
[0038]
[0039] Let r m =[r m1 ,r m2 ,...,r m(m-1) ,r m(m+1) ,...,r mM ], where r mm′ For image x m and x m′ The feature similarity of
[0040] The calculation formula for feature ranking loss is:
[0041]
[0042] Where, v m is the image label similarity vector; r m is the image feature similarity vector.
[0043] Furthermore, the feature similarity of the images is calculated using cosine distance.
[0044] Furthermore, the approximate normalized discounted cumulative gain loss is calculated to optimize the ranking loss of the ranking classification network model, specifically:
[0045] Let s=(s1,s2,...,s K ), s k is the network's predicted value of the similarity of the k-th key image; let r=(r1,r2,...,r K), r K Indicates the correlation level of the k-th image;
[0046] Let π be any ranking list of K key images, then the cumulative loss gain of π is
[0047]
[0048] Where π(k) represents the position of the kth key image in π;
[0049] Let s and r correspond to the ranking list respectively π s and π r , then π s The normalized discounted cumulative gain is
[0050]
[0051] Among them, Ψ(π s ,r)∈(0,1];
[0052] When θ→+∞, the Sigmoid function Infinitely approach the unit step function h(Δ);
[0053] The unit step function h(Δ) is expressed as:
[0054]
[0055] According to π s (k)=1+∑ k′≠k h(s k′ -s k ),have
[0056]
[0057] Let μ be the right side of the equation s (k), using μ s (k) replaces π s (k), we get π s The approximate cumulative loss gain is
[0058]
[0059] Then we get π s The approximate normalized discounted cumulative gain is
[0060]
[0061] Where, is a derivable optimization objective.
[0062] Furthermore, training is performed through multi-task learning, specifically:
[0063] The sorting classification network model includes three tasks: label sorting, sample sorting, and feature similarity sorting. Combining the label sorting loss, sample sorting loss, and feature sorting loss, the total loss function for sorting classification network model training is:
[0064] l total =l 标 +l 样 +λl 特
[0065] Among them, λ>0 is an adjustable hyperparameter; l 标 is the label ranking loss; l 样 is the sample sorting loss; l 特 is the feature ranking loss.
[0066] Furthermore, the ranking classification network model is evaluated, and the evaluation indicators are as follows:
[0067] For single-label classification, the overall accuracy is used as the evaluation metric; for multi-label classification, the mean F1 score, mean F2 score, and mean p e 、mean r e 、mean p l and mean r l Multiple indicators are evaluated;
[0068] The calculation formulas for the F1 score and F2 score are as follows:
[0069]
[0070] Where p e Indicates the accuracy calculated based on samples; r e Represents the recall rate calculated based on samples;
[0071] The calculation formulas for precision and recall are as follows:
[0072]
[0073] Where TP represents the number of samples correctly predicted as positive examples; FP represents the number of samples incorrectly predicted as positive examples; FN represents the number of samples incorrectly predicted as negative examples; when β = e, it means calculation based on samples, and when β = l, it means calculation based on labels.
[0074] Furthermore, the ranking classification network model is used to simultaneously process the single-label classification task and the multi-label classification task of remote sensing images, specifically:
[0075] In the single-label classification task, the sorting classification network model uses ResNet50 and Swin Transformer, and uses the Softmax function to obtain the probability output value of single-label classification;
[0076] In the multi-label classification task, the sorting classification network model uses ResNet50 and VGG-16, and the Sigmoid function is used to obtain the probability output value of multi-label classification.
[0077] Compared with the existing technology, the present invention uses ranking learning and deep learning technologies to construct a ranking classification network model, and uniformly performs end-to-end training and evaluation. It can simultaneously solve single-label classification and multi-label classification tasks. By calculating the label ranking loss, sample ranking loss and feature ranking loss, it optimizes the label, sample and feature ranking relationship. By calculating the approximate normalized discounted cumulative gain loss, it optimizes the ranking loss of the ranking classification network model, improves the overall classification performance and accuracy of the model, and reduces the workload of switching between different classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 Flowchart of the remote sensing scene classification method based on ranking learning.
[0079] Figure 2 This is a structural diagram of the sorting classification network model.
[0080] Figure 3 Sorting diagram for labels.
[0081] Figure 4 Schematic diagram of sample sorting.
[0082] Figure 5 Schematic diagram of feature sorting.
[0083] Figure 6 Schematic diagram of four data set examples. DETAILED DESCRIPTION
[0084] The remote sensing scene classification method based on ranking learning of the present invention is further described below with reference to the accompanying drawings and specific embodiments.
[0085] See also Figure 1 The present invention discloses a remote sensing scene classification method based on ranking learning, comprising the following steps:
[0086] S1: Construct a sorting classification network model to simultaneously process single-label classification tasks and multi-label classification tasks of remote sensing images;
[0087] S2: Introducing virtual neutral samples, the predicted values of virtual neutral samples are constrained to the interval between the positive label prediction value and the negative label prediction value of the image;
[0088] S3: Combined with virtual neutral samples, calculate the label ranking loss, sample ranking loss and feature ranking loss, and optimize the label ranking relationship, sample ranking relationship and feature ranking relationship;
[0089] S4: Calculate the approximate normalized discounted cumulative gain loss and optimize the ranking loss of the ranking classification network model;
[0090] S5: Combine multiple ranking losses, train through multi-task learning, evaluate the ranking classification network model, and apply the trained ranking classification network model to remote sensing scene classification.
[0091] In step S1, a sorting classification network model is constructed.
[0092] Specifically, see Figure 2 This paper integrates the concept of ranking learning into the remote sensing scene classification task, and constructs a unified and efficient ranking classification network model framework, aiming to simultaneously address the two major challenges of multi-label classification and single-label classification. The ranking classification network model framework includes three major parts: label ranking, sample ranking, and feature ranking, and can use a variety of network structures as feature extractors, such as CNN, Transformer, etc. In order to overcome the difficulties faced by multi-label classification in the label decision stage, the concept of virtual neutral samples is introduced. Through this unique design, different backbone networks can be trained and evaluated under an end-to-end architecture.
[0093] In order to avoid introducing too many notations, the label space is The label of image x (which can be a single-label or multi-label image) is called the positive label of x, and the remaining labels are called the negative labels of x; assuming that each small batch contains M training images, the mth image is x m , and record x m The positive label set is The negative label set is
[0094] Input x m , the sorting classification network model generates N prediction values z m1 ,z m2 ,...,z mN ∈[0,1], where z mn Represents the positive label set of the mth image Contains the label value y n probability; expectation When z mn >0.5, and When z mn <0.5.
[0095] In order to separate the predicted values of positive and negative labels, add A virtual neutral label is created, and a value θ is set to limit the probability range of the neutral label. For a given θ∈(0,0.5), the "virtual prediction value" of the virtual neutral label is set to fall within the interval [0.5-θ,0.5+θ], thereby establishing an "interval interval" between the positive and negative labels. When z mn >0.5+θ; When z mn <0.5-θ.
[0096] In step S3, the label ranking loss is calculated in combination with the virtual neutral samples.
[0097] Specifically, the label ranking loss enables the model to learn the relationship between different scenes in a single image sample. The set of neutral labels is denoted as For any given and θ∈(0,0.5), let the vector in
[0098]
[0099] Let vector in
[0100]
[0101] The calculation formula for label ranking loss is
[0102]
[0103] Where Φ is the sorting optimization function, is the label prediction vector, is the virtual neutral label vector
[0104] See also Figure 3 , using label ranking loss l 标 Training a ranking classification network model is essentially sorting labels, and expecting positive labels to be ranked before virtual neutral labels, which in turn are ranked before negative labels. Figure 3 Explain the meaning of label sorting. For the predicted values in a batch, add neutral labels to the label dimension and force the loss function to and In other words, we expect positive labels to be larger than neutral labels and negative labels to be smaller than neutral labels.
[0105] In step S3, the sample ranking loss is calculated in combination with the virtual neutral samples.
[0106] Specifically, the sample ranking loss enables the model to learn the inter-class relationship between different samples. If y is the positive label of image x, then x is called a positive sample of y, otherwise x is called a negative sample of y; in a small batch, if Then x m y n Positive samples.
[0107] M training samples generate M predicted values: z 1n ,z 2n ,...,z Mn ; Sort the corresponding samples in descending order of predicted values, hoping that positive samples are ranked before negative samples. Similar to label sorting, insert between positive and negative samples Virtual neutral samples.
[0108] For any given Let vector in
[0109]
[0110] Furthermore, let the vector in
[0111]
[0112] The sample ranking loss calculation formula is:
[0113]
[0114] Where, is the sample prediction vector; is the virtual neutral sample vector.
[0115] See also Figure 4 , Figure 4 Explains the meaning of sample sorting, which is similar to label sorting, but is different from the sample dimension. and Calculate the losses.
[0116] In step S3, the feature ranking loss is calculated in combination with the virtual neutral samples.
[0117] Specifically, feature ranking loss can enable the model to learn deeper inter-class relationships. Let v m =[v m1 ,v m2 ,...,v m(m-1) ,v m(m+1) ,...,v mM ], where v mm′ For collection and The similarity coefficient, v mm′ Measured the mth image x m and the m′th image x m′ The label similarity of represents the positive label set of the m′th image.
[0118] The calculation formula of similarity coefficient is:
[0119]
[0120] Let r m =[r m1 ,r m2 ,...,r m(m-1) ,r m(m+1) ,...,r mM ], where r mm′ For image x m and x m′ There are many ways to define the feature similarity of two images, for example, it can be defined as the cosine distance of the global feature vectors of the two images.
[0121] The calculation formula for feature ranking loss is:
[0122]
[0123] Where, v m is the image label similarity vector; r m is the image feature similarity vector.
[0124] See also Figure 5 , Figure 5 Give a schematic diagram of the above process, by forcing v m and r m The order of sizes is the same to enable the feature network to learn feature knowledge.
[0125] In step S4, the approximate normalized discounted cumulative gain loss is calculated to optimize the ranking loss of the ranking classification network model.
[0126] Specifically, the normalized discounted cumulative gain is used to measure the quality of the ranking, and the approximate normalized discounted cumulative gain loss function achieves the purpose of constructing an ideal ranking by optimizing the normalized discounted cumulative gain retrieval index.
[0127] Given a query image, the network needs to sort K images (called "key images") based on the similarity between the key image and the query image (calculated using the feature embedding of the two images). The higher the similarity, the higher the ranking. Let s = (s1, s2, ..., s K ), where s kis the network's predicted value of the similarity of the kth key image, which is a derivative function of the network parameter ω. Let r=(r1,r2,...,r K ), r K It represents the relevance level of the k-th image, which is a predetermined value that measures the relevance between the k-th image and the query image.
[0128] Let π be any ranking list of K key images, then the cumulative loss gain of π is
[0129]
[0130] Wherein, π(k) represents the position (ie, serial number) of the k-th key image in π.
[0131] Let s and r correspond to the ranking list respectively π s and π r , then π s The normalized discounted cumulative gain is
[0132]
[0133] Among them, Ψ(π s ,r)∈(0,1], the closer to 1, the π s The more perfect.
[0134] When π=π s When π in formula (9) s (k) is not differentiable, so Φ(π s ,r) is not differentiable, and thus Ψ(π s ,r) is also not differentiable. Therefore, if the gradient descent method is used to update the network parameters, Ψ(π s ,r) cannot be used as the optimization target. Therefore, this embodiment turns to find a differentiable approximation to replace Ψ(π s ,r).
[0135] Easy to see:
[0136] π s (k)=1+∑ k′≠k h(s k′ -s k ) (11)
[0137] where h(Δ) is the unit step function:
[0138]
[0139] When θ→+∞, the Sigmoid function
[0140]
[0141] Infinitely approximate the unit step function h(Δ). Therefore, for sufficiently large θ, we have
[0142]
[0143] Let μ be the right side of the equation s (k), using μ s (k) replaces π s (k), we get π s Approximate discounted cumulative gain
[0144]
[0145] Then we can get π s The approximate normalized discounted cumulative gain of
[0146]
[0147] in, It is a derivable optimization objective that is only targeted at one query image.
[0148] In step S5, multiple ranking losses are combined and trained through multi-task learning.
[0149] Specifically, the sorting classification network model includes three tasks: label sorting, sample sorting, and feature similarity sorting. Combining the label sorting loss, sample sorting loss, and feature sorting loss, the total loss function for sorting classification network model training is:
[0150] l total =l 标 +l 样 +λl 特 (17)
[0151] Among them, λ>0 is an adjustable hyperparameter and its optimal value is analyzed later.
[0152] The present invention combines the above three sorting losses and innovatively proposes the following Figure 1 The three-rank loss network framework shown in the figure can be introduced into various classification backbone networks. It can be used as a leading function to guide model training or as an auxiliary loss to optimize model performance. At the same time, it can be trained and evaluated in an end-to-end manner.
[0153] The following is a specific example, in which the present invention is applied to the problem of remote sensing image scene classification, and experiments are conducted on four remote sensing datasets, such as Figure 6 As shown, the technical solution of the present invention is compared with the implementation effect of the prior art.
[0154] AID: The dataset is a single-label dataset released by Wuhan University, containing 30 categories of scene images, each category contains 200-400 images, totaling 10,000 images, where the pixel size of each image is 600x600, and the spatial resolution ranges from 0.5m to 8m.
[0155] NWPU-RESISC dataset: A single-label dataset collected and annotated by Northwestern Polytechnical University, containing 31,500 images with a pixel size of 256x256, spatial resolution ranging from 0.2m to 30m, covering 45 different scene categories, with 700 samples in each category.
[0156] UCM Multilabel Dataset: A multi-label dataset based on aerial images in the UCM dataset. It has 17 labels and contains 2,100 images. Each image is 256x256 pixels in size and has a spatial resolution of one foot.
[0157] AID Multilabel Dataset: A multi-label dataset obtained by re-annotating the AID dataset. It uses the same label definition as UCM Multilabel and has a total of 3,000 images.
[0158] The hardware and software environment for implementing the present invention is as follows: All experiments in the present invention are completed on an AMAX workstation equipped with two GPUs (NVIDIA Titan X) and 128G of memory. At the same time, the proposed sorting classification network model is implemented in the PyTorch framework and trained in an end-to-end manner.
[0159] In step S5, the ranking classification network model is evaluated.
[0160] Specifically, for single-label classification, the overall accuracy is used as the evaluation metric; for multi-label classification, the mean F1 score, mean F2 score, mean p e 、mean r e 、mean p l and mean r l Multiple indicators are evaluated;
[0161] The calculation formulas for the F1 score and F2 score are as follows:
[0162]
[0163] Where p e Indicates the accuracy calculated based on samples; r e Indicates the recall rate calculated based on samples;
[0164] The calculation formulas for precision and recall are as follows:
[0165]
[0166] Where TP represents the number of samples correctly predicted as positive examples; FP represents the number of samples incorrectly predicted as positive examples; FN represents the number of samples incorrectly predicted as negative examples; when β = e, it means calculation based on samples, and when β = l, it means calculation based on labels.
[0167] To comprehensively evaluate the proposed method and provide a fair comparison with other methods, we selected ResNet50 and VGG-16 as the model backbone for multi-label classification tasks, while ResNet50 and Swin Transformer were used for single-label classification tasks. The number of neurons in the last linear layer was adjusted based on the number of labels in the dataset. In the multi-label classification task, the Sigmoid function was used instead of the Softmax function in the single-label classification task to obtain the probabilistic output value of the multi-label classification. At the same time, the threshold for label decision was set to 0.5.
[0168] Regarding the data set partitioning, two split ratios were set for each single-label dataset: AID (20%-80%), AID (50%-50%), NWPU-RESISC (10%-90%), and NWPU-RESISC (20%-80%). The former in parentheses indicates the training set ratio, and the latter indicates the test set ratio. The experiment was repeated ten times for each ratio setting, and the mean and standard deviation are reported in the results.
[0169] For multi-label datasets, the division method defined by their publishers is adopted. 80% of the image samples for each scene category are selected as the training set, and the other 20% of the images are used as the test set. In addition, for each dataset, 10% are randomly selected from the training set to form the validation set. The validation set can be used in subsequent hyperparameter tuning experiments, and it also helps improve model training and prevent model overfitting.
[0170] During model training, the images of the AID dataset and NWPU-RESISC dataset were randomly cropped to 512x512 and 224x224 sizes, respectively, while the images of the two multi-label datasets were directly scaled to 512x512 and 224x224 sizes (random cropping of multi-label images may lose label information). At the same time, horizontal flipping, vertical flipping and 90-degree rotation were applied for data augmentation. The Adam optimizer was used, and the initial learning rate was set to 1*10 -5The learning rate is decayed in stages according to the epoch. For AID and AID Multilabel, the batch_size of the training is set to 16, and for the other two datasets it is set to 64.
[0171] To explore the impact of hyperparameters on model performance, we selected AID Multilabel and NWPU-RESISC (10%-90%) and determined the optimal hyperparameter values for both classification tasks on their validation sets. We used ResNet50 as the model backbone for the experiments. For multi-label classification, we used the mean F1 score as the evaluation metric, while for single-label classification, we used overall accuracy.
[0172] To address the label decision problem, virtual neutral samples are introduced. However, the number of virtual neutral samples can be a key factor affecting model performance. To determine the optimal number of virtual neutral samples, an exponential search method is used. Specifically, a search space of {0, 2, 4, 8, 16, 32, 64} is defined to find the optimal number of virtual neutral samples.
[0173] To eliminate the impact of other hyperparameters on the experimental results, the similarity of the virtual neutral samples was fixed at 0.5, and the correlation was set to 1. This means that the values corresponding to each label category of the virtual neutral samples were uniformly equal to 1. Furthermore, for more accurate evaluation, only the label ranking loss was introduced for model training. Six different models were trained based on this setting. Table 1 shows the results of these models on the validation set.
[0174] Table 1 Validation set results for different numbers of virtual neutral samples
[0175]
[0176]
[0177] Based on the experimental results in the table above, we determined that the optimal number of virtual neutral samples for the AID Multilabel dataset and the NWPU-RESISC dataset is 16 and 64, respectively. These two values coincide with the batch size (batch_size) set during the experiments. Based on this important finding, the number of virtual neutral samples was uniformly set to the batch_size in subsequent experiments.
[0178] In order to comprehensively evaluate the impact of the similarity range of virtual neutral samples on model performance, in the following experiments, the number of virtual neutral samples was fixed to the optimal value obtained previously, and the similarity of virtual samples was set to increase uniformly within the preset range. The experimental results are shown in Table 2.
[0179] Table 2 Validation set results for different similarity ranges
[0180]
[0181] The similarity range in subsequent experiments was set to 0.4-0.6.
[0182] In order to ensure the rationality and effectiveness of the selection of weight λ, the present invention sets a transformation range containing multiple candidate values, specifically λ∈{0.5, 1, 2, 4, 8, 16, 32}. The experimental results are shown in Table 3 below.
[0183] Table 3 Validation set results with different λ
[0184]
[0185]
[0186] In subsequent experiments, we set λ = 4 and obtained the best performing model based on the three tuned hyperparameters.
[0187] To verify the effectiveness of the three-ranking loss framework of the ranking classification network model of the present invention, a series of ablation experiments were conducted on two benchmark datasets, and finally six models with different configurations were trained to systematically compare their performance.
[0188] In Table 4, Models 1, 2, and 3 represent models that use only label ranking loss, sample ranking loss, and feature ranking loss as training losses, respectively. Models 4 and 5 explore the effects of combining label ranking loss, sample ranking loss, and feature ranking loss, respectively. Model 6 is trained using a combination of all three losses. Model 3 uses only feature ranking loss for training; therefore, during the testing phase, we used the KNN method for performance evaluation. The other models employed an end-to-end training and evaluation strategy.
[0189] Table 4 Ablation experiment results
[0190]
[0191] Table 5 shows the comparison results of the proposed method and other SOTA methods on the UCM Multilabe dataset. It can be observed that the proposed method surpasses all competitors regardless of the model backbone used. In particular, when using VGGNet16 as the model backbone, the proposed method is significantly superior to other SOTA methods, which also shows that the proposed method has good robustness to different model backbones. In addition, compared with other methods, the proposed method also basically achieves the best performance in various precision and recall rate indicators, which fully demonstrates the superiority of the proposed method.
[0192] Table 5 Comparison results of the three ranking loss frameworks and existing methods on the UCM Multilabel Dataset
[0193]
[0194]
[0195] To further evaluate the three-rank loss framework, we also compare our method with other SOTA methods on the AID Multilabel Dataset. The specific results are shown in Table 6. Consistent with the observations on the UCM Multilabel dataset, our method still significantly outperforms other SOTA methods.
[0196] Table 6 Comparison results of the three ranking loss frameworks and existing methods on AID Multilabel Dataset
[0197] method <![CDATA[mean F1]]> <![CDATA[mean F2]]> <![CDATA[mean p e ]]> <![CDATA[mean r e ]]> <![CDATA[mean p l ]]> <![CDATA[mean r l ]]> ResNet-50 (2016) 86.23 85.57 89.31 85.65 72.39 52.82 ResNet-RBFNN (2017) 83.77 85.87 82.84 88.32 60.85 70.45 CA-ResNet-BiLSTM (2019) 87.63 88.03 89.03 88.99 79.50 65.60 AL-RN-ResNet (2020) 88.72 88.54 91.00 88.95 80.81 71.12 GC-MLFNet (2021) 92.03 91.31 93.25 90.84 83.87 76.31 ResNet50-SR-Net (2021) 89.97 - 89.42 90.52 87.24 82.25 ResNet50 (Ours) 92.18 93.33 90.37 94.14 87.51 87.55 VGGNet (2014) 85.52 85.60 87.41 86.32 70.60 58.89 VGG-RBFNN (2017) 84.58 85.99 84.56 87.85 62.90 69.15 CA-VGG-BiLSTM (2019) 86.68 86.88 88.68 87.83 72.04 60.00 AL-RN-VGGNet (2020) 88.09 88.31 89.96 89.27 76.94 68.31 CNN-GNN (2020) 88.26 88.68 89.61 89.55 - - MLRSSC-CNN-GNN (2020) 88.64 89.18 89.83 90.20 - - VGG16-SR-Net (2021) 87.15 - 86.84 87.46 75.79 65.38 VGGNet16(Ours) 90.88 90.86 90.93 90.85 85.58 74.44
[0198] Table 7 shows the classification results of the three-ranking loss framework using different model backbones and other SOTA methods. The present invention uses the NWPU-RESISC (10%-90%) dataset to conduct parameter experiments and determine the optimal hyperparameters. Then, the hyperparameter configuration is directly used for training on AID (20%-80%), AID (50%-50%), and NWPU-RESISC (20%-80%).
[0199] When using ResNet50 as the model's backbone network, the accuracy gap between the three-rank loss framework and the current best method is less than 2% under four dataset configurations. These best results are achieved by different methods. In particular, on the AID (20%-80%) dataset, the accuracy gap between the experimental results of our invention and the best result is only 0.29%.
[0200] When using the Transformer as the backbone network, the proposed method outperforms all Transformer-based methods on the AID (20%-80%) dataset. Although slightly inferior to those combining CNN and Transformer, the proposed method maintains similar accuracy to the best results in the other three configurations.
[0201] Table 7 Comparison results on single-label remote sensing datasets
[0202]
[0203] In summary, the present invention constructs a ranking classification network model and uniformly performs end-to-end training and evaluation, which can simultaneously solve single-label classification and multi-label classification tasks. By calculating the label ranking loss, sample ranking loss and feature ranking loss, the label, sample and feature ranking relationship is optimized. By calculating the approximate normalized discounted cumulative gain loss, the ranking loss of the ranking classification network model is optimized, the overall classification performance and accuracy of the model are improved, and the workload of switching between different classification tasks is reduced.
[0204] The above description is a detailed description of the preferred embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention should fall within the patent scope covered by the present invention.
Claims
1. A remote sensing scene classification method based on ranking learning, characterized in that: The following steps are involved: S1: Construct a sorting classification network model to simultaneously process single-label classification tasks and multi-label classification tasks of remote sensing images; S2: Introducing virtual neutral samples, the predicted values of virtual neutral samples are constrained to the interval between the positive label prediction value and the negative label prediction value of the image; S3: Combined with virtual neutral samples, calculate the label ranking loss, sample ranking loss and feature ranking loss, and optimize the label ranking relationship, sample ranking relationship and feature ranking relationship; S4: Calculate the approximate normalized discounted cumulative gain loss and optimize the ranking loss of the ranking classification network model; S5: Combine multiple ranking losses, train through multi-task learning, evaluate the ranking classification network model, and apply the trained ranking classification network model to remote sensing scene classification.
2. The remote sensing scene classification method based on ranking learning according to claim 1 is characterized in that Construct a sorting classification network model, specifically: The ranking classification network model framework consists of three parts: label ranking, sample ranking, and feature ranking. It uses multiple network structures as feature extractors, introduces virtual neutral samples, and performs training and evaluation in an end-to-end architecture. Let the label space be The label of image x is called the positive label of x, and the remaining labels are called the negative labels of x; assuming that each small batch contains M training images, the mth image is x m , and record x m The positive label set is The negative label set is Input x m , the sorting classification network model generates N prediction values z m1 ,z m2 ,...,z mN ∈[0,1], where z mn Represents the positive label set of the mth image Contains the label value y n probability; expectation When z mn >0.5, and When z mn <0.5; join in A virtual neutral label is created, and the value of θ is set to limit the probability range of the neutral label. For a given θ∈(0,0.5), the "virtual prediction value" of the virtual neutral label is set to fall within the interval [0.5-θ,0.5+θ], and an "interval interval" is established between the positive and negative labels. When z mn >0.5+θ; When z mn <0.5-θ.
3. The remote sensing scene classification method based on ranking learning according to claim 2 is characterized in that: Combined with the virtual neutral samples, the label ranking loss is calculated as follows: The set of neutral labels is For any given and θ∈(0,0.5), let the vector in Let vector in The calculation formula for label ranking loss is Where Φ is the sorting optimization function, is the label prediction vector, is the virtual neutral label vector.
4. The remote sensing scene classification method based on ranking learning according to claim 2 is characterized in that: Combined with the virtual neutral samples, the sample ranking loss is calculated as follows: If y is the positive label of image x, then x is called a positive sample of y, otherwise x is called a negative sample of y; in a small batch, if Then x m y n Positive samples of M training samples generate M predicted values: z 1n ,z 2n ,...,z Mn ; Sort the corresponding samples in descending order of predicted values, and insert between positive and negative samples Virtual neutral samples; For any given Let vector in Furthermore, let the vector in The sample ranking loss calculation formula is: Where, is the sample prediction vector; is the virtual neutral sample vector.
5. The remote sensing scene classification method based on ranking learning according to claim 2 is characterized in that: Combined with the virtual neutral sample, the feature ranking loss is calculated as follows: Let v m =[v m1 ,v m2 ,...,v m(m-1) ,v m(m+1) ,...,v mM ], where v mm′ For collection and The similarity coefficient, v mm′ Measure the mth image x m and the m′th image x m′ The label similarity of represents the positive label set of the m′th image; The calculation formula of similarity coefficient is: Let r m =[r m1 ,r m2 ,...,r m(m-1) ,r m(m+1) ,...,r mM ], where r mm′ For image x m and x m′ The feature similarity of The calculation formula for feature ranking loss is: Where, v m is the image label similarity vector; r m is the image feature similarity vector.
6. The remote sensing scene classification method based on ranking learning according to claim 5 is characterized in that: The feature similarity of images is calculated using cosine distance.
7. The remote sensing scene classification method based on ranking learning according to claim 1 is characterized in that: Calculate the approximate normalized discounted cumulative gain loss and optimize the sorting loss of the sorting classification network model, specifically: Let s=(s1,s2,...,s K ), s k is the network's predicted value of the similarity of the k-th key image; let r=(r1,r2,...,r K ), r K Indicates the correlation level of the k-th image; Let π be any ranking list of K key images, then the cumulative loss gain of π is Where π(k) represents the position of the kth key image in π; Let s and r correspond to the ranking list respectively π s and π r , then π s The normalized discounted cumulative gain is Among them, Ψ(π s ,r)∈(0,1]; When θ→+∞, the Sigmoid function Infinitely approach the unit step function h(Δ); The unit step function h(Δ) is expressed as: According to π s (k)=1+∑ k′≠k h(s k′ -s k ),have Let μ be the right side of the equation s (k), using μ s (k) replaces π s (k), we get π s The approximate cumulative loss gain is Then we get π s The approximate normalized discounted cumulative gain is Where, is a derivable optimization objective.
8. The remote sensing scene classification method based on ranking learning according to claim 1 is characterized in that: Training is performed through multi-task learning, specifically: The sorting classification network model includes three tasks: label sorting, sample sorting, and feature similarity sorting. Combining the label sorting loss, sample sorting loss, and feature sorting loss, the total loss function for sorting classification network model training is: Among them, λ>0 is an adjustable hyperparameter; loss for label ranking; is the sample ranking loss; is the feature ranking loss.
9. The remote sensing scene classification method based on ranking learning according to claim 1, characterized in that: Evaluate the sorting classification network model. The evaluation indicators are as follows: For single-label classification, the overall accuracy is used as the evaluation metric; for multi-label classification, the meanF1 score, meanF2 score, and meanp e 、meanr e 、meanp l and meanr l Multiple indicators are evaluated; The calculation formulas for the F1 score and F2 score are as follows: Where p e Indicates the accuracy calculated based on samples; r e Indicates the recall rate calculated based on samples; The calculation formulas for precision and recall are as follows: Where TP represents the number of samples correctly predicted as positive examples; FP represents the number of samples incorrectly predicted as positive examples; FN represents the number of samples incorrectly predicted as negative examples; when β = e, it means calculation based on samples, and when β = l, it means calculation based on labels.
10. The remote sensing scene classification method based on ranking learning according to claim 1, characterized in that: The sorting classification network model is used to simultaneously process single-label classification tasks and multi-label classification tasks of remote sensing images, specifically: In the single-label classification task, the sorting classification network model uses ResNet50 and Swin Transformer, and uses the Softmax function to obtain the probability output value of single-label classification; In the multi-label classification task, the sorting classification network model uses ResNet50 and VGG-16, and the Sigmoid function is used to obtain the probability output value of multi-label classification.