Image Scene Classification Method for Progressive Training of Multi-Scale Information Retrieval Network
By building a multi-scale information retrieval network and adopting an incremental learning algorithm, the problems of insufficient feature extraction and noise samples in remote sensing scene classification are solved, and a higher accuracy of remote sensing image classification is achieved.
Patent Information
- Application Number
- CN202311203883.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-09-18
AI Technical Summary
The existing remote sensing scene classification method lacks feature extraction capabilities when facing noisy scene images, and is greatly affected by noise samples, resulting in low classification accuracy.
A multi-scale information retrieval network of dual-twin branches is built, and the characteristics of different scales are fused through the Transformer module, and a progressive learning algorithm is used to train the network in stages, including reverse learning, sample selection and re-labeling strategies to reduce the impact of noise samples.
The classification accuracy of remote sensing images is improved, the characteristics in complex remote sensing scenarios can be accurately understood, the impact of noise samples on the model is reduced, and the classification performance is achieved.
Smart Images

Figure CN117079056B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to an image scene classification method for progressively training a multi-scale information retrieval network in the field of image classification technology. The present invention can be used for classifying remote sensing scene images. Background Art
[0002] Among various remote sensing technologies, remote sensing scene classification is a basic remote sensing discrimination technology, whose purpose is to define scene labels that conform to the content of remote sensing scenes. Accurate scene classification results are beneficial to different remote sensing tasks and applications, such as image retrieval, land cover classification, hazard and environmental monitoring, and resource exploration. However, due to the diverse geographical environments and different land use types contained in remote sensing images, such as mountains, rivers, cities, and wetlands, etc., remote sensing scenes have significantly complex features compared with natural images. Understanding the complex features of remote sensing scenes has practical significance for remote sensing image scene classification methods. At the same time, in engineering practice, due to the characteristics of remote sensing scene images such as a large number of land cover categories and large volumes, it is very difficult to label newly obtained remote sensing scenes. Manually labeling large-scale remote sensing data is challenging, time-consuming and laborious. In addition, it requires practitioners to have professional knowledge. The accuracy of machine-labeling remote sensing data still needs to be improved. Due to the powerful learning ability of neural networks, mislabeled remote sensing scenes will directly affect the final classification performance.
[0003] Kang Jian et al. proposed a remote sensing image classification method for noise scenarios based on metric learning in their published paper "Noise-tolerant deep neighborhood embedding for remotely sensed images with label noise" (IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021). This method accurately encodes the semantic relationships between remote sensing scenarios by improving the original metric learning loss function, maximizes the scores of leave-one-out K-nearest neighbors to explain the inherent neighborhood structure between images in the feature space, and reduces the contribution of potential noisy images by learning local structures and pruning images with low K-nearest neighbor scores. Finally, features are extracted by a neural network and classified using the K-nearest neighbor method. Although this method reduces the impact of noisy samples on the classification results by selecting samples, the drawback of this method is that the network structure in this method uses the original neural network, which has limited feature extraction ability, and the extracted features are not sufficient to accurately describe the content of complex remote sensing scenarios. Therefore, in the classification task for noise scenarios, using a basic neural network cannot exhibit good classification performance due to insufficient feature extraction ability.
[0004] Qilu University of Technology proposed a remote sensing scene classification method based on a multi-head self-attention convolutional neural network in its patent document "A Remote Sensing Scene Classification Method Based on a Multi-Head Attention Convolutional Neural Network" (Patent Application No.: CN 202210381142.1, Authorization Publication No.: CN 114463646 B). This method further encodes the convolutional feature maps learned by the convolutional neural network using the multi-head self-attention layer, thereby fully utilizing context information to capture features. However, the drawback of this method is that it only classifies correctly labeled remote sensing scenarios. Therefore, only the cross-entropy loss function is used for training during the training of the network. Since remote sensing scenarios may be mislabeled in actual applications, in scenarios where the dataset contains noisy samples, the classification performance of the model constructed by this method cannot meet the expected index requirements due to the interference of noisy samples. Summary of the Invention
[0005] The purpose of the present invention is to propose an image scene classification method for progressively training a multi-scale information retrieval network to solve the technical problems of insufficient feature extraction ability and low classification accuracy caused by being greatly affected by noisy samples when existing remote sensing scene classification methods are used for classifying noise-scenario images.
[0006] To achieve the above object, the idea of the present invention is to fuse the multi-scale information extracted by the neural network from the remote sensing scene through a Transformer. The fused global context features contain both information at different scales and the relationships between information at different scales / within the same scale. Compared with the global convolutional features extracted by the neural network, the global context features can supplement the missing long-range dependencies, thereby solving the problem of insufficient model feature extraction ability when classifying images in a noisy scene. At the same time, a dual twin-branch structure is adopted to reduce the impact of noisy samples on a single model. The present invention uses a progressive learning algorithm to divide the training process of the model into three stages. First, a reverse learning strategy is adopted to initially train the multi-scale information retrieval network, and a reverse learning loss function is used to simply learn the relationship between samples and labels. Second, a sample selection strategy is adopted to select some samples and further train the multi-scale information retrieval network through a cross-entropy loss function. Finally, in order to avoid information loss, a relabeling strategy is adopted to train and use the cross-entropy loss function to train the network to complete the training of the multi-scale information retrieval network, thereby solving the problems of low classification accuracy and the model being greatly affected by noisy samples.
[0007] The specific steps to achieve the object of the present invention are as follows:
[0008] Step 1, generate a training set:
[0009] Select at least P remote sensing images to form a training set. The training set includes at least C remote sensing scene categories. Each remote sensing scene category contains at least N images, and among them, there are M noisy remote sensing images in this category. Wherein, N is greater than or equal to 1, M ≤ N, C is greater than or equal to 2, and P = C × N;
[0010] Step 2, build a multi-scale information retrieval network with a dual twin-branch structure composed of a first sub-network and a second sub-network with the same structure in parallel; each sub-network includes two branches and four scale reduction layers; the first branch is composed of a downsampling module, a first convolutional module, a second convolutional module, a third convolutional module, and a fourth convolutional module connected in series in sequence; the second branch is composed of a splicing layer, a Transformer module, and a classifier connected in series in sequence; the first to fourth convolutional modules in the first branch are respectively connected to the first to fourth scale reduction layers and then connected to the splicing layer of the second branch; the fourth convolutional module in the first branch is connected to the classifier of the second branch;
[0011] Step 3, adopt a reverse learning strategy to initially train the multi-scale information retrieval network:
[0012] Input the training set into the multi-scale information retrieval network. Using the gradient descent method, iteratively update the weight values of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the reverse learning loss function converges, obtaining a preliminarily trained multi-scale information retrieval network;
[0013] Step 4, adopt a sample selection strategy to further train the multi-scale information retrieval network:
[0014] Input the training set into the preliminarily trained multi-scale information retrieval network, and use the sample selection strategy to select some samples to participate in the training; use the gradient descent method to iteratively update the weights of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the cross-entropy loss function converges, obtaining a further trained multi-scale information retrieval network;
[0015] Step 5, adopt a relabeling strategy to complete the training of the multi-scale information retrieval network:
[0016] Input the training set into the further trained multi-scale information retrieval network, and use the relabeling strategy to reassign labels to the training samples in the training set; use the gradient descent method to iteratively update the weights of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the cross-entropy loss function converges, obtaining a trained multi-scale information retrieval network;
[0017] Step 6, classify the remote sensing image:
[0018] Input the remote sensing image to be classified into the trained multi-scale information retrieval network, and output a classification result vector. This vector contains probability values corresponding to each remote sensing scene category in the training set. Take the category corresponding to the maximum probability value as the classification result of this remote sensing image.
[0019] The present invention has the following advantages compared with the existing technologies:
[0020] First, since the present invention constructs a multi-scale information retrieval network containing double twin branches, fuses the features of different scales of the remote sensing scene through the Transformer module, and fuses the obtained global context features and global convolution features to complete the classification task, it overcomes the defect that the existing technology simply designs a robustness learning algorithm starting from noise samples and the features extracted by the original neural network are insufficient, resulting in low classification accuracy. The present invention can fully understand the complex features contained in the complex remote sensing scene, thereby improving the classification accuracy of the remote sensing image.
[0021] Second, since the present invention trains the network through a progressive learning algorithm, the training process of the network is divided into three stages, namely, initially training the multi-scale information retrieval network using a reverse learning strategy, further training the multi-scale information retrieval network using a sample selection strategy, and completing the training of the multi-scale information retrieval network using a relabeling strategy. These three training methods reduce the impact of noise samples in the dataset on the network, overcome the problem that the method of training the network using the cross-entropy loss function commonly used in existing remote sensing scene classification technologies is greatly affected by noise samples, and enable the present invention to accurately classify image scenes even in the face of noise samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the overall flowchart of the implementation of the present invention;
[0023] Figure 2 is the structural schematic diagram of the multi-scale information retrieval network constructed by the present invention;
[0024] Figure 3 is the structural schematic diagram of the sub-network in the multi-scale information retrieval network of the present invention;
[0025] Figure 4 is the structural schematic diagram of the residual block in the multi-scale information retrieval network of the present invention;
[0026] Figure 5 is the structural schematic diagram of the Transformer module in the multi-scale information retrieval network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0028] Refer to Figure 1 to further describe in detail the implementation steps of the embodiments of the present invention.
[0029] Step 1, generate a training set.
[0030] In the embodiments of the present invention, the dataset used is the UCM21 remote sensing scene classification dataset, which was created by researchers at the University of California, Merced. It contains 2,100 remote sensing scene images, which are evenly divided into 21 different scene categories. According to the different categories, 80 pictures are selected from each scene category, and among them, 32 remote sensing pictures containing noise are included. The training labels of the remote sensing pictures without noise in the training set are their corresponding scene categories, and the training labels of the remote sensing pictures containing noise are randomly selected from one of the 21 scene categories that is different from their own scene category.
[0031] Step 2: Build a multi-scale information retrieval network with a dual-twin branch structure composed of a first sub-network and a second sub-network with the same structure in parallel. Its structure is as shown in Figure 2 . The multi-scale information retrieval network includes two branches, and the two branches share an input end. The first branch structure is composed of a first sub-network and a first output end connected in series in sequence. The second branch structure is composed of a second sub-network and a second output end connected in series in sequence.
[0032] Refer to Figure 3 for a further description of the structures of the first sub-network and the second sub-network.
[0033] The structures of the first sub-network and the second sub-network are the same. Each sub-network includes two branches and four scale reduction layers. The first branch is composed of a downsampling module, a first convolution module, a second convolution module, a third convolution module, and a fourth convolution module connected in series in sequence. The second branch is composed of a splicing layer, a Transformer module, and a classifier connected in series in sequence. The first to fourth convolution modules in the first branch are respectively connected to the first scale reduction layer, the second scale reduction layer, the third scale reduction layer, and the fourth scale reduction layer, and then connected to the splicing layer of the second branch. The fourth convolution module in the first branch is connected to the classifier of the second branch.
[0034] The downsampling module is composed of a convolution layer, a normalization layer, an activation layer, and a max pooling layer connected in series in sequence. Set the number of input channels of the convolution layer to N r , and the value of N r is the same as the number of channels of the input remote sensing image. In the embodiment of the present invention, the input image includes 3 channels. Therefore, N r is 3, the number of output channels is 64, the convolution kernel size is set to 7×7, the convolution stride is set to 1, and the boundary padding value is set to 1. Set the number of channels of the normalization layer to 64; the activation layer is implemented using the ReLU activation function; the stride of the max pooling layer is set to 2.
[0035] The structures of the first convolution module, the second convolution module, the third convolution module, and the fourth convolution module are as follows: The first convolution module is composed of a first residual block, a second residual block, and a max pooling layer connected in series in sequence. The second convolution module is composed of a third residual block, a fourth residual block, a fifth residual block, and a max pooling layer connected in series in sequence. The third convolution module is composed of a sixth residual block, a seventh residual block, an eighth residual block, a ninth residual block, and a max pooling layer connected in series in sequence. The fourth convolution module is composed of a tenth residual block, an eleventh residual block, and a max pooling layer connected in series in sequence. Set the stride of the max pooling layer in the first to third convolution modules to 2; set the stride of the max pooling layer in the fourth convolution module to 7.
[0036] The structures of the first to eleventh residual blocks in the first convolution module, the second convolution module, the third convolution module, and the fourth convolution module are all the same.
[0037] Refer to Figure 4 , and a further description of the structure of the residual block is given.
[0038] Each residual block is composed of a first convolution layer, a first batch normalization layer, an activation layer, a second convolution layer, and a second batch normalization layer connected in series in sequence. The output of the first convolution layer is added to the output of the second convolution layer. The convolution kernel sizes of the first and second convolution layers are both set to 3×3, the convolution strides are both set to 1, and the boundary padding values are both set to 1; the number of channels of the first and second batch normalization layers is equal to the number of output channels of the corresponding residual block; the activation layer is implemented using the ReLU activation function.
[0039] The input channel numbers of the first to eleventh residual blocks are set to 64, 64, 64, 128, 128, 128, 256, 256, 256, 256, 512 respectively; the output channel numbers are set to 64, 64, 128, 128, 128, 256, 256, 256, 256, 512, 512 respectively. The input channel number of the first convolution layer corresponds to and is equal to the input channel number of the corresponding residual block, the output channel number corresponds to and is equal to the output channel number of the corresponding residual block, and the input channel number and output channel number of the second convolution layer are both equal to the output channel number of the corresponding residual block.
[0040] The structures of the four scale reduction layers in each sub-network are all the same. Each scale reduction layer is composed of a convolution layer, a regularization layer, and a pooling layer connected in series in sequence. The convolution kernel size of the convolution layer is set to 1×1, the convolution stride is set to 1, and the boundary padding value is set to 1. The input channel numbers of the convolution layers in the first to fourth scale reduction layers are set to 64, 128, 256, 512 respectively; the output channel numbers are all set to 512. The regularization layer is implemented using the L2 regularization function; the strides of the pooling layers in the first to fourth scale reduction layers are set to 8, 4, 2, 1 respectively.
[0041] Refer to Figure 5 , and a further description of the structure of the Transformer module in the multi-scale information retrieval network of the present invention is given.
[0042] The Transformer modules in the first and second sub-networks each include a first regularization layer, a multi-head attention layer, a second regularization layer, and a multi-layer perceptron unit. Among them, the first regularization layer, the multi-head attention layer, the second regularization layer, and the multi-layer perceptron unit are connected in series in sequence. The input of the first regularization layer is also connected by addition to the output of the multi-head attention layer; the input of the second regularization layer is also connected by addition to the output of the multi-layer perceptron unit. The number of input channels of the first and second regularization layers is set to 512, and the number of layers of the multi-head attention layer is 4. The multi-layer perceptron unit consists of two hidden layers connected in series, and the number of input channels and output channels of the two hidden layers are both set to 512.
[0043] The classifier in the sub-network uses a perceptron, and the number of input channels of the perceptron is set to 1024, and the number of output channels is equal to the number of remote sensing scene categories in the training set.
[0044] Step 3, adopt a reverse learning strategy to preliminarily train the multi-scale information retrieval network.
[0045] Input the training set into the multi-scale information retrieval network, use the gradient descent method to iteratively update the weight values of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the reverse learning loss function converges, and obtain the preliminarily trained multi-scale information retrieval network.
[0046] The parameter transfer method mentioned above is completed by the following formula:
[0047]
[0048] Among them, represents the parameters of the second sub-network after the current iteration update when the multi-scale information retrieval network is preliminarily trained using the reverse learning strategy, represents the parameters of the second sub-network after the update in the previous iteration when the multi-scale information retrieval network is preliminarily trained using the reverse learning strategy, represents the parameters of the first sub-network after the current iteration update when the multi-scale information retrieval network is preliminarily trained using the reverse learning strategy.
[0049] The reverse learning loss function mentioned above is as follows:
[0050]
[0051] Among them, L NL represents the reverse learning loss value between the sample prediction probability value and the actual probability value output when the multi-scale information retrieval network is preliminarily trained using the reverse learning strategy, N represents the total number of samples in the training set, C represents the total number of scene categories included in the training set, It represents the sign function, k represents the serial number of the category of the samples in the training set, k ≤ C. When the complementary category of the i-th sample is equal to the category of the k-th sample in the training set Otherwise The complementary category is another category that belongs to the categories of the samples in the training set and is inconsistent with the true category, p ik It represents the probability that the prediction result of the i-th sample in the training set belongs to category k, and log(·) represents the logarithmic operation with base 10. Since the number of pictures in the training set in the embodiments of the present invention is 1680 and the total number of scene categories is 21, so N in the present invention takes the value of 1680 and C takes the value of 21.
[0052] Step 4: Adopt a sample selection strategy to further train the multi-scale information retrieval network.
[0053] Input the training set into the preliminarily trained multi-scale information retrieval network, use the sample selection strategy to select some samples to participate in the calculation of the cross-entropy loss function in the training process; use the gradient descent method to iteratively update the weights of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the cross-entropy loss function converges, and obtain the further trained multi-scale information retrieval network.
[0054] The sample selection strategy mentioned above refers to selecting samples whose own labels are consistent with the predicted labels of the first sub-network and the second sub-network to participate in the training.
[0055] The parameter transfer method is completed by the following formula:
[0056]
[0057] Where It represents the parameters of the second sub-network after the current iterative update when further training the multi-scale information retrieval network using the sample selection strategy It represents the parameters of the second sub-network after the update in the previous iteration when further training the multi-scale information retrieval network using the sample selection strategy It represents the parameters of the first sub-network after the current iterative update when further training the multi-scale information retrieval network using the sample selection strategy.
[0058] The cross-entropy loss function is as follows:
[0059]
[0060] Where L t It represents the cross-entropy loss value between the predicted probability value and the actual probability value of the sample during the t-th training of the multi-scale information retrieval network, t = 2, 3, y ik It represents the sign function. When the true category of the i-th sample is equal to the category of the k-th sample in the training set yik = 1, otherwise, y ik = 0.
[0061] Step 5, adopt the relabeling strategy to complete the training of the multi-scale information retrieval network.
[0062] Input the training set into the further trained multi-scale information retrieval network, and use the relabeling strategy to reassign labels to the training samples in the training set; use the gradient descent method to iteratively update the weights of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the cross-entropy loss function converges, obtaining the trained multi-scale information retrieval network.
[0063] The relabeling strategy refers to reassigning labels to the samples in the training set using the following formula:
[0064]
[0065] where y i represents the relabeled label of the i-th sample in the training set, and p clean represents the probability that the prediction result of the i-th sample in the training set belongs to its true class. represents the sign function. When the predicted label of the second sub-network of the i-th sample is equal to the class of the k-th sample in the training set, otherwise,
[0066] The parameter transfer method is completed by the following formula:
[0067]
[0068] where represents the parameters of the second sub-network after the current iteration update when further training the multi-scale information retrieval network using the relabeling strategy, represents the parameters of the second sub-network after the update in the previous iteration when further training the multi-scale information retrieval network using the relabeling strategy, represents the parameters of the first sub-network after the current iteration update when further training the multi-scale information retrieval network using the relabeling strategy.
[0069] The cross-entropy loss function is as follows:
[0070]
[0071] L2 represents the cross-entropy loss value between the predicted probability value and the actual probability value of the sample when training the multi-scale information retrieval network is completed using the relabeling strategy.
[0072] Step 6, classify the remote sensing images.
[0073] Input the remotely sensed image to be classified into the trained multi-scale information retrieval network, and output a classification result vector, which contains probability values corresponding to each remotely sensed scene category in the training set. Take the category corresponding to the maximum probability value as the classification result of the remotely sensed image.
[0074] In the embodiment of the present invention, select pictures that do not belong to the test set from the UCM21 dataset. These pictures belong to a certain category among the 21 scene categories of UCM21. After inputting these pictures into the trained multi-scale information retrieval network, add the output of the first sub-network and the output of the second sub-network to obtain a classification result vector. The classification result vector contains the probability values of the pictures belonging to each scene category, and the scene category corresponding to the maximum probability value is the classification result of the sample point.
[0075] The following further describes the effect of the present invention in combination with simulation experiments.
[0076] 1. Simulation experiment conditions.
[0077] The hardware platform for the simulation experiment of the present invention: The processor is an Intel(R) Xeon(R) E5-2650 v4 CPU, the main frequency is 2.20 GHz, the memory is 125 GB, and the graphics card is a GeForce GTX TiTan XP.
[0078] The software platform for the simulation experiment of the present invention is: Ubuntu operating system and python 3.7.
[0079] The remotely sensed scene dataset used in the simulation experiment of the present invention is UCM21, which was created by researchers at the University of California, Merced. It contains 21 scene categories, each category contains 100 pictures, the size of the pictures is 224×224×3, the image format is RGB, and the spatial resolution is 0.3 m.
[0080] 2. Simulation content and its result analysis:
[0081] The simulation experiment of the present invention uses the present invention and four existing technologies (denoising CMR-NLD classification method based on covariance matrix, robust metric learning NTDNE classification method, RS-COCL classification method based on reverse learning and forward learning, NCE classification method based on regularization cross entropy) to classify the remotely sensed pictures in the input UCM21 respectively, and obtain classification results.
[0082] In the simulation experiment, the four existing technologies used refer to:
[0083] The existing covariance matrix-based denoising CMR-NLD classification method refers to the remote sensing image classification method for noisy scenes proposed by B. Tu et al. in "Robust learning of mislabeled training samples for remote sensing image scene classification[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, PP(13):5623-5639.", abbreviated as the CMR-NLD classification method.
[0084] The existing robust metric learning NTDNE classification method refers to the remote sensing image classification method for noisy scenes proposed by J. Kang et al. in "Noise-tolerant deep neighborhood embedding for remotely sensed images with label noise[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14:2551-2562.", abbreviated as the NTDNE classification method.
[0085] The existing reverse learning and forward learning RS-COCL classification method refers to the remote sensing image classification method for noisy scenes proposed by Q. Li et al. in "Complementary learning-based scene classification of remote sensing images with noisy labels[J]. IEEE Geoscience&Remote Sensing Letters, 2021, 19:1-5.", abbreviated as the RS-COCL classification method.
[0086] The existing regularized cross-entropy NCE classification method refers to the image classification method proposed by X. Ma et al. in "Normalized loss functions for deep learning with noisy labels[C]. in International conference on machine learning, PMLR, 2020, pp.6543–6553", abbreviated as the NCE classification method.
[0087] Using the overall classification accuracy OA as the evaluation index, the classification results of the present invention and four existing classification methods are evaluated respectively. The formula of OA is as follows:
[0088]
[0089] The performance evaluation index comparison of the classification results of the above-mentioned present invention and four existing remote sensing image classification methods for the UCM21 dataset is shown in Table 1:
[0090] Table 1 List of comparison results of evaluation indexes
[0091] method OA CMR-NLD 88.57 NTDNE 91.04 RS-COCL 95.90 NCE 75.80 the present invention 96.21
[0092] As can be seen from Table 1, compared with the other four existing classification methods, the present invention shows better classification performance, and its index value of the overall classification accuracy OA is better than the other four algorithms, proving that the present invention can obtain higher classification accuracy of remote sensing scene images.
[0093] The above simulation experiments show that the method of the present invention uses a multi-scale information retrieval network to mine information from remote sensing images, which can effectively mine the complex information contained in remote sensing images and ensure the accuracy and integrity of image feature information; by adopting a three-stage progressive learning algorithm to train the network, it can effectively complete the remote sensing image classification task, reduce the influence of noisy label samples on the model during training, and thus improve the classification accuracy; it solves the problems of insufficient network feature extraction ability and large influence of using the cross-entropy loss function to train the network by noisy samples in the prior art, resulting in low classification accuracy, and is a very practical remote sensing image classification method.
Claims
1. An image scene classification method for progressively training a multi-scale information retrieval network, characterized in that, Construct a multi-scale information retrieval network with double twin branches and train the network using a progressive learning algorithm; the steps of this classification method are as follows: Step 1, generate a training set: Select at least P remote sensing images to form a training set. The training set includes at least C remote sensing scene categories. Each remote sensing scene category contains at least N images, and there are M noisy remote sensing images in this category, where N is greater than or equal to 1, M ≤ N, C is greater than or equal to 2, and P = C × N; Step 2, build a multi-scale information retrieval network with a double twin branch structure composed of two sub-networks with the same structure in parallel; each sub-network includes two branches and four scale reduction layers; the first branch is composed of a downsampling module, a first convolution module, a second convolution module, a third convolution module, and a fourth convolution module connected in series in sequence; the second branch is composed of a splicing layer, a Transformer module, and a classifier connected in series in sequence; the first to fourth convolution modules in the first branch are respectively connected to the first to fourth scale reduction layers and then connected to the splicing layer of the second branch; the fourth convolution module in the first branch is connected to the classifier of the second branch; Step 3, adopt a reverse learning strategy to preliminarily train the multi-scale information retrieval network: Input the training set into the multi-scale information retrieval network, use the gradient descent method to iteratively update the weight values of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the reverse learning loss function of the network converges, obtaining a preliminarily trained multi-scale information retrieval network; Step 4, adopt a sample selection strategy to further train the multi-scale information retrieval network: Input the training set into the preliminarily trained multi-scale information retrieval network, use the sample selection strategy to select some samples to participate in the calculation of the cross-entropy loss function during the training process; use the gradient descent method to iteratively update the weights of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the cross-entropy loss function of the network converges, obtaining a further trained multi-scale information retrieval network; Step 5, adopt a relabeling strategy to complete the training of the multi-scale information retrieval network: Input the training set into the further trained multi-scale information retrieval network, use the relabeling strategy to reassign labels to the training samples in the training set; use the gradient descent method to iteratively update the weights of the first sub-network, and use the parameter transfer method to update the weights of the second sub-network until the cross-entropy loss function of the network converges, obtaining a trained multi-scale information retrieval network; Step 6, classify remote sensing images: Input the remote sensing image to be classified into the trained multi-scale information retrieval network, and output a classification result vector. This vector contains probability values corresponding to each remote sensing scene category in the training set. Take the category corresponding to the maximum probability value as the classification result of the remote sensing image to be classified.
2. The method for image scene classification of the progressive training multi-scale information retrieval network according to claim 1, characterized in that, The downsampling module described in step 2 consists of a convolutional layer, a normalization layer, an activation layer, and a max pooling layer connected in series in sequence; the number of input channels of the convolutional layer is set to N r , N r is equal to the number of channels of the input remote sensing scene image, the number of output channels is set to 64, the size of the convolutional kernel is set to 7×7, the convolutional stride is set to 1, and the boundary padding value is set to 1; the number of channels of the normalization layer is set to 64; the activation layer is implemented using the ReLU activation function; the stride of the max pooling layer is set to 2.
3. The image scene classification method for progressively training a multi-scale information retrieval network according to claim 1, wherein The structures of the first convolution module, the second convolution module, the third convolution module, and the fourth convolution module described in step 2 are as follows: The first convolution module is composed of a first residual block, a second residual block, and a max pooling layer connected in series in sequence; the second convolution module is composed of a third residual block, a fourth residual block, a fifth residual block, and a max pooling layer connected in series in sequence; the third convolution module is composed of a sixth residual block, a seventh residual block, an eighth residual block, a ninth residual block, and a max pooling layer connected in series in sequence; the fourth convolution module is composed of a tenth residual block, an eleventh residual block, and a max pooling layer connected in series in sequence; the strides of the max pooling layers in the first to third convolution modules are all set to 2; the stride of the max pooling layer in the fourth convolution module is set to 7.
4. The method for image scene classification of the progressive training multi-scale information retrieval network according to claim 3, wherein The structures of the first to eleventh residual blocks in the first convolution module, the second convolution module, the third convolution module, and the fourth convolution module described in step 2 are the same; each residual block is composed of a first convolutional layer, a first batch normalization layer, an activation layer, a second convolutional layer, and a second batch normalization layer connected in series in sequence; the output of the first convolutional layer is added to the output of the second convolutional layer; the kernel sizes of the first and second convolutional layers are both set to 3×3, the convolutional strides are both set to 1, and the boundary padding values are both set to 1; the number of channels of the first and second batch normalization layers is equal to the number of output channels of the corresponding residual block; the activation layer is implemented using the ReLU activation function; the input channel numbers of the first to eleventh residual blocks are set to 64, 64, 64, 128, 128, 128, 256, 256, 256, 256, 512 respectively; the output channel numbers are set to 64, 64, 128, 128, 128, 256, 256, 256, 256, 512, 512 respectively; the input channel number of the first convolutional layer corresponds to the input channel number of the corresponding residual block, the output channel number corresponds to the output channel number of the corresponding residual block, and the input channel number and output channel number of the second convolutional layer are both equal to the output channel number of the corresponding residual block.
5. The image scene classification method for progressively training a multi-scale information retrieval network according to claim 1, wherein The structures of the four scale reduction layers described in step 2 are the same, and each is composed of a convolutional layer, a regularization layer, and a pooling layer connected in series in sequence; the kernel size of the convolutional layer is set to 1×1, the convolutional stride is set to 1, and the boundary padding value is set to 1; the input channel numbers of the convolutional layers in the first to fourth scale reduction layers are set to 64, 128, 256, 512 respectively; the output channel numbers are all set to 512; the regularization layer is implemented using the L2 regularization function; the strides of the pooling layers in the first to fourth scale reduction layers are set to 8, 4, 2, 1 respectively.
6. The image scene classification method for progressively training a multi-scale information retrieval network according to claim 1, characterized in that The Transformer module described in step 2 includes a first regularization layer, a multi-head attention layer, a second regularization layer, and a multi-layer perceptron unit; among them, the first regularization layer, the multi-head attention layer, the second regularization layer, and the multi-layer perceptron unit are connected in series in sequence; the input of the first regularization layer is also connected by adding to the output of the multi-head attention layer; the input of the second regularization layer is also connected by adding to the output of the multi-layer perceptron unit; the input channel numbers of the first and second regularization layers are both set to 512, and the number of layers of the multi-head attention layer is 4; the multi-layer perceptron unit consists of two hidden layers connected in series, and the input channel numbers and output channel numbers of the two hidden layers are both set to 512.
7. The image scene classification method for progressively training a multi-scale information retrieval network according to claim 1, wherein The classifier described in step 2 uses a perceptron, and the input channel number of the perceptron is set to 1024, and the output channel number is equal to the number of remote sensing scene categories in the training set.
8. The method for image scene classification of a progressive training multi-scale information retrieval network according to claim 1, wherein The parameter transfer method described in step 3, step 4, and step 5 is completed by the following formula: Among them, represents the parameters of the second sub-network after the current iteration update during the t-th training of the multi-scale information retrieval network, where t = 1, 2, 3. t = 1 represents the stage of initially training the multi-scale information retrieval network using the backward learning strategy described in step 3, t = 2 represents the stage of training the multi-scale information retrieval network using the sample selection strategy described in step 4, and t = 3 represents the stage of training the multi-scale information retrieval network using the relabeling strategy described in step 5; represents the parameters of the second sub-network after the update in the previous iteration during the t-th training of the multi-scale information retrieval network, represents the parameters of the first sub-network after the current iteration update during the t-th training of the multi-scale information retrieval network.
9. The method for image scene classification of the progressive training multi-scale information retrieval network according to claim 8, characterized in that, The reverse learning loss function described in step 3 is as follows: Among them, L NL represents the reverse learning loss value between the sample prediction probability value and the actual probability value output when initially training the multi-scale information retrieval network using the reverse learning strategy. represents the sign function, k represents the serial number of the category of the sample in the training set, k ≤ C. When the complementary category of the i-th sample is equal to the category of the k-th sample in the training set Otherwise The complementary category is another category that belongs to the categories of the samples in the training set and is inconsistent with the true category. p ik represents the probability that the prediction result of the i-th sample in the training set belongs to category k, and log(·) represents the logarithmic operation with base 10.
10. The image scene classification method for progressively training a multi-scale information retrieval network according to claim 1, wherein, The sample selection strategy described in step 4 refers to selecting samples whose own labels are consistent with the predicted labels of the first sub-network and the second sub-network to participate in training.
11. The method for image scene classification of progressive training of multi-scale information retrieval network according to claim 9, wherein, The cross-entropy loss function described in step 4 and step 5 is as follows: Among them, L t represents the cross-entropy loss value between the predicted probability value and the actual probability value of the sample during the t-th multi-scale information retrieval network training, where t = 2, 3, y ik represents the sign function. When the true category of the i-th sample is equal to the category of the k-th sample in the training set, y ik = 1; otherwise, y ik = 0.
12. The method for image scene classification of progressive training multi-scale information retrieval network according to claim 11, characterized in that The relabeling strategy described in step 5 refers to reassigning labels to the samples in the training set using the following formula: where y i represents the re-assigned label of the i-th sample in the training set, and p clean represents the probability that the prediction result of the i-th sample in the training set belongs to its true class, represents the sign function. When the predicted label of the second sub-network of the i-th sample is equal to the class of the k-th sample in the training set, otherwise,
Citation Information
Patent Citations
A Remote Sensing Scene Classification Method Based on Multi-Head Self-Attention Convolutional Neural Network
CN114463646B
High-resolution remote sensing image scene classification method for small data set
CN107220657A
Hyperspectral remote sensing image classification method
CN113705526A