Optical Remote Sensing Image Ground Object Classification Method Based on Multi-Level Pseudo-Relationship Learning

By adopting multi-level pseudo-relationship learning method in optical remote sensing image classification, and using the multi-level pseudo-relationship network of teacher model and student model, the problem of relying on a large amount of manual annotation information in the existing technology is solved, the classification accuracy and efficiency are improved, and the generalization ability of the model is enhanced.

CN116434075BActive Publication Date: 2025-06-20XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310227880.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-09
Publication Date
2025-06-20
Estimated Expiration
2043-03-09

AI Technical Summary

Technical Problem

The prior art relies on a large amount of manual labeling information in the classification of optical remote sensing images, resulting in high data labeling costs and insufficient generalization capabilities of deep learning models in the absence of labeling of target domain data.

Method used

Using a multi-level pseudo-relationship learning method, a multi-level pseudo-relationship network of teacher model and student model is constructed, and a teacher model is used to guide the student model to learn the relationship between pixel points and local blocks in the target domain, and a variety of pseudo-relationship learning loss functions are constructed to optimize the model.

Benefits of technology

It improves the accuracy and efficiency of optical remote sensing images, reduces the dependence on manual annotation information, and enhances the generalization ability of the model in the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434075B_ABST
    Figure CN116434075B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning, which is used to solve the problem that the existing self-training method ignores the potential connections between pixel points in ground object classification. The implementation steps are as follows: obtaining the optical remote sensing image to be classified from a remote device and constructing a multi-level pseudo-relationship network model, obtaining the optimization objective function of the optimized model by constructing a source domain relationship learning loss, a target domain pixel-level pseudo-relationship loss, a source domain local patch-level pseudo-relationship loss, and a target domain local patch-level pseudo-relationship loss, then training the model using a training set, and finally classifying the optical remote sensing image. The present invention realizes the classification of ground objects in optical remote sensing images through multi-level pseudo-relationship learning, improves the classification effect, makes up for the deficiencies of the self-training method, and can be used in application fields such as urban planning, land use, and environmental detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to a semi-supervised ground object classification method, specifically to an optical remote sensing image ground object classification method based on multi-level pseudo-relationship learning, which can be used in application fields such as urban planning, land use, and environmental detection. Background Technique

[0002] Optical remote sensing images are usually obtained through sensing devices such as airborne and spaceborne satellites. Compared with other geographical information data, optical remote sensing images have a wide perspective and a high richness of information content. However, due to their unique acquisition method, optical remote sensing images are easily affected by the atmospheric environment (such as light, climate, etc.). Before the rise of big data and deep learning, the ground object classification process of optical remote sensing images usually needed to be completed by manual visual interpretation. Ground object classification requires assigning category annotation information to each pixel of the image one by one, which requires a large amount of manpower and material resources for remote sensing images with rich semantic information. The total cost of the process from obtaining optical remote sensing data to its engineering application is relatively high. The emergence of deep learning provides a convenient way to implement the ground object classification task for remote sensing images.

[0003] Using machines to replace humans in classifying optical remote sensing images mainly benefits from the booming development of big data and computer vision. Computer vision simulates the visual function of the human eye through machines. After training, neural networks can also obtain spatial perception capabilities similar to humans through the input planar images, and then participate in the classification and detection of specific targets. The learning of neural networks requires a large number of training samples. The emergence of massive data provides data support for artificial intelligence and guarantees the improvement of subsequent algorithms.

[0004] Even today, although the remote sensing image processing method based on deep learning has made great improvements, the data problem is still an important factor restricting the performance of neural network models. Different factors such as geographical location and acquisition settings will lead to differences in the distribution of ground object categories. When deploying applications, it is impossible to ensure that all ground object categories in the data are visible to the model. The most direct way to solve this problem is to continue to input a large number of training samples into the model, and corresponding sample annotation information is also required. The development of optical remote sensing image ground object classification based on deep learning has experienced supervised learning, semi-supervised learning, and unsupervised learning, and its dependence on manual annotation is decreasing, and the efficiency of image interpretation is gradually improving. Currently, the research on ground object classification based on deep learning has become a trend. While reducing the interpretation cost, it can also ensure a certain interpretation efficiency. The ground object classification algorithms based on deep learning are constantly improving, the classification effect is gradually enhanced, and the application prospects in related fields will also be more and more extensive.

[0005] With the development of computer vision, more and more deep learning-based methods have been used for image processing. Long et al. replaced the last fully connected layer of the convolutional neural network with a convolutional layer to propose the fully convolutional neural network (FCN) for image segmentation. The fully convolutional neural network can upsample the last layer of feature maps through deconvolution operations to ensure that the network can accept images of any size as input. This retains more original image information and enables it to achieve good performance in pixel-level classification tasks. Later scholars have made many improvements on the basis of FCN and achieved remarkable results. Badrinarayanan et al. proposed SegNet, which improved the computational efficiency of FCN by using the indices of max pooling for upsampling. Ronneberger et al. designed U-Net, which adopted skip connections between the encoder and the decoder, and fused all upsampled maps with the original map to obtain more scale information. However, all these works lost some information due to pooling operations, resulting in relatively rough segmentation results. Therefore, the DeepLab series of models adopted dilated convolution to obtain a larger receptive field without losing information and designed the ASPP module to fuse features of different scales.

[0006] The rapid development of semantic segmentation methods based on deep convolutional neural networks has led to many applications in the remote sensing field. Dimitrios Marmanis et al. applied FCN and its various variants to urban high-resolution images and achieved good results. Audebert et al. found that using SegNet could improve the classification accuracy of small targets (vehicles). Martin et al. combined multiple convolutional neural networks and obtained context regions of different sizes to retain high-resolution information. Mou et al. learned long-range dependencies and enhanced feature representation by obtaining global spatial and channel relationships. Domain differences make the generalization effect of models pre-trained on the source domain poor when fine-tuning on the target domain. This requires enhancing the discriminative ability of the model for a specific target domain, which is also the part ignored by domain alignment methods. The self-training method well makes up for this point. Its aim is to generate pixel-level pseudo-labels for target domain images. The introduction of pseudo-labels can, on the one hand, alleviate the problem that target domain data is not labeled, and on the other hand, enable the classifier to obtain a more accurate decision boundary for target domain data.

[0007] Liang Yan et al. regarded the prediction labels of target data with high confidence as pseudo-labels classified by category. Xuedong Yao et al. selected pseudo-labels through threshold setting. Famao Ye et al. generated pseudo-labels through clustering. Yangyang Li et al. used a teacher-student model to update the pseudo-labels of target images on average. And Lefei Zhang et al. used a generator to adaptively update the pseudo-labels. Yue Wu et al. corrected the pseudo-labels by adding spatial constraints. These works only focused on how to generate reliable pseudo-labels, but ignored the spatial relationship between pixels, resulting in not very high classification accuracy of the model. Summary of the Invention

[0008] To solve the above problems existing in the prior art, the present invention provides a method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0009] The present invention provides a method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning, including:

[0010] Step 1, obtain the optical remote sensing image to be classified from a remote device and obtain a training set composed of multiple optical remote sensing images from a database;

[0011] Among them, the training set includes source domain samples and target domain samples. The target domain samples are distinguished from the source domain samples by different geographical locations. The source domain samples carry the target category labels included in the images.

[0012] Step 2, construct a multi-level pseudo-relationship network model, where the multi-level pseudo-relationship network model includes a teacher model and a student model with exactly the same structure and initialization parameters, and the parameters between the teacher model and the student model are not shared;

[0013] Step 3, add perturbations or noises to the training samples and send them into the student model; send the training samples into the teacher model, and perform iterative training on the student model and the teacher model; during the iterative training process, construct a source domain relationship learning loss and a source domain patch-level pseudo-relationship loss for the source domain samples in sequence, and construct a target domain pixel-level pseudo-relationship loss and a target domain patch-level pseudo-relationship loss for the target domain samples, so as to obtain a final optimization objective function for optimizing the multi-level pseudo-relationship network model; extract the primary features and multi-scale features of the training samples through their respective models, and both output a first output result at the pixel level and a second output result at the patch level;

[0014] Step 4, repeat the training process of Step 3 for each batch of training samples, and use the final optimization objective function to optimize the training process during the iteration to obtain a trained multi-level pseudo-relationship network model;

[0015] Step 5: Use the trained multi-level pseudo-relationship network model to extract the features of the optical remote sensing image to be classified, and classify according to the features to obtain the category of the target contained in the optical remote sensing image to be classified.

[0016] The present invention provides an optical remote sensing image ground object classification device based on multi-level pseudo-relationship learning, including:

[0017] An acquisition module, configured to acquire the optical remote sensing image to be classified from a remote device and acquire a training set composed of multiple optical remote sensing images from a database;

[0018] Wherein, the training set includes source domain samples and target domain samples, the target domain samples are distinguished from the source domain samples by different geographical locations, and the source domain samples carry the target category labels contained in the images;

[0019] A construction module, configured to construct a multi-level pseudo-relationship network model, the multi-level pseudo-relationship network model includes a teacher model and a student model with exactly the same structure and initialization parameters, and the parameters between the teacher model and the student model are not shared;

[0020] An optimization training module, configured to add perturbations or noises to the training samples and send them into the student model; send the training samples into the teacher model, and perform iterative training on the student model and the teacher model; during the iterative training process, construct a source domain relationship learning loss and a source domain patch-level pseudo-relationship loss for the source domain samples in sequence, and construct a target domain pixel-level pseudo-relationship loss and a target domain patch-level pseudo-relationship loss for the target domain samples, so as to obtain a final optimization objective function for optimizing the multi-level pseudo-relationship network model; extract the primary features and multi-scale features of the training samples through their respective models, and both output a first output result at the pixel level and a second output result at the patch level;

[0021] An iteration module, configured to repeat the training process of step 3 for each batch of training samples, and use the final optimization objective function to optimize the training process during the iteration to obtain a trained multi-level pseudo-relationship network model;

[0022] A classification module, configured to use the trained multi-level pseudo-relationship network model to extract the features of the optical remote sensing image to be classified, and classify according to the features to obtain the category of the target contained in the optical remote sensing image to be classified.

[0023] The beneficial effects of the present invention:

[0024] 1. The present invention mainly proposes a method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning, which uses a teacher model to guide the student model to learn the relationships between pixel points and local patches in the target domain. The parameters of the teacher model are calculated by means of exponential moving average, and the parameters of the student model are updated by backpropagation of the loss function.

[0025] 2. The present invention constructs four pseudo-relationship learning loss functions, namely source domain relationship learning loss, target domain pixel-level pseudo-relationship loss, source domain local patch-level pseudo-relationship loss, and target domain local patch-level pseudo-relationship loss. By combining these four losses, an optimization objective function for the learning model of multi-level pseudo-relationships is constructed, and the performance of the model is improved compared with the self-training method. Therefore, the present invention can improve the classification accuracy.

[0026] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0027] Figure 1 is a schematic flowchart of a method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning provided by the present invention;

[0028] Figure 2 is the network structure diagram of the present invention;

[0029] Figure 3 is a schematic diagram of the cutting of the smallest subgraph;

[0030] Figure 4 is a visualization result diagram of segmentation of different methods under the ISPRS 2D semantic label challenge dataset. Detailed Embodiments

[0031] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0032] As Figure 1 shown, the present invention provides a method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning, including:

[0033] Step 1, obtaining the optical remote sensing image to be classified from a remote device and obtaining a plurality of historical optical remote sensing images from a database, and using the historical optical image training set;

[0034] Among them, the training set includes source domain samples and target domain samples. The target domain samples are distinguished from the source domain samples by different geographical locations, and the source domain samples carry the target category labels included in the images.

[0035] Step 2, construct a multi-level pseudo-relationship network model. The multi-level pseudo-relationship network model includes a teacher model and a student model with exactly the same structure and initialization parameters, and the parameters between the teacher model and the student model are not shared;

[0036] Reference Figure 2 , both the teacher model and the student model include a feature extractor, two ASPP networks, a patch-level classifier, and a pixel-level classifier;

[0037] The output of the feature extractor is connected to the inputs of the two ASPP networks. The output of one ASPP network is connected to the input of the patch-level classifier, and the output of the other ASPP network is connected to the input of the pixel-level classifier.

[0038] Parameter settings of the feature extractor:

[0039] The convolution kernel size of the first convolutional layer is set to 3*3, the stride is set to 1, and the size of the feature map is set to 64; the convolution kernel size of the second convolutional layer is set to 3*3, the stride is set to 2, and the size of the feature map is set to 128; the convolution kernel size of the third convolutional layer is set to 3*3, the stride is set to 2, and the size of the feature map is set to 256; the convolution kernel size of the fourth convolutional layer is set to 3*3, the stride is set to 1, and the size of the feature map is set to 512; the activation function for each layer is the Relu function, and batch normalization operations are performed on all of them.

[0040] Parameter settings of the ASPP model:

[0041] The convolution kernel size of the first convolutional layer is set to 3*3, the stride is set to 1, the dilation rate is set to 1, and the size of the feature map is set to 256; the convolution kernel size of the second convolutional layer is set to 3*3, the stride is set to 3, the dilation rate is set to 6, and the size of the feature map is set to 256; the convolution kernel size of the third convolutional layer is set to 3*3, the stride is set to 3, the dilation rate is set to 12, and the size of the feature map is set to 256; the convolution kernel size of the fourth convolutional layer is set to 3*3, the stride is set to 3, the dilation rate is set to 18, and the size of the feature map is set to 256; the fifth layer merges the outputs of the first four layers and performs global average pooling operation. The activation function for each layer is the Relu function.

[0042] Parameter settings of the classifier model:

[0043] The convolution kernel size of the first convolutional layer is set to 3*3, the stride is set to 1, and the feature map size is set to 256; the convolution kernel size of the second convolutional layer is set to 3*3, the stride is set to 1, and the feature map size is set to 256; the output size of the pixel-level classifier is 256*256, and the output size of the patch-level classifier is 64*64; the activation function for each layer is the Relu function.

[0044] Step 3: Add perturbations or noise to the training samples and send them into the student model; send the training samples into the teacher model, and perform iterative training on the student model and the teacher model; during the iterative training process, construct a source domain relationship learning loss and a source domain patch-level pseudo-relationship loss for the source domain samples in sequence, and construct a target domain pixel-level pseudo-relationship loss and a target domain patch-level pseudo-relationship loss for the target domain samples, so as to obtain a final optimization objective function for optimizing the multi-level pseudo-relationship network model; extract the primary features and multi-scale features of the training samples through their respective models, and both output a first output result at the pixel level and a second output result at the patch level.

[0045] In the present invention, the training samples are sent into the teacher model and then sent into the student model after adding perturbations or noise. The primary features of the original image are extracted through the feature extraction network, the multi-scale features of the image are obtained through the ASPP network, and finally the output results at the pixel level and the patch level are obtained through two classifiers.

[0046] The parameters of the student model are updated through backpropagation. The parameters of the student model at time t are denoted as The parameters of the student model at time t-1 are denoted as The parameters of the teacher model are updated by means of exponential moving average. The parameters of the teacher model at time t are denoted as The parameters of the teacher model at time t-1 are denoted as The parameter update method of the teacher model is as follows:

[0047]

[0048] where τ∈[0,1], and its value in this experiment is 0.99;

[0049] The construction of the source domain relationship learning loss for the source domain samples in Step 3 includes:

[0050] Step 3a: For the source domain samples in the training set, denoted as X S =x s , for the pixel-level labels of the source domain samples, denoted as Y S =y s ;

[0051] Step 3b: For the target domain samples in the training set, denoted as XT = x t ;

[0052] Step 3c, calculate the source domain segmentation loss:

[0053]

[0054] Among them, the student model or the teacher model is denoted as G, the feature extractor is denoted as F, and the classifier is denoted as C.

[0055] For source domain pixel-level relationship learning: Since the source domain data itself has labels, supervised relationship learning can be performed.

[0056] In step 3, constructing the target domain relationship learning loss for the target domain samples includes:

[0057] Step 4a, for the target domain sample x t , send it into the teacher model to generate its probability score p t ;

[0058] Step 4b, regard all pixels output by the teacher model as nodes, regard the similarity S between nodes as the weight of the edge, and construct the adjacency matrix W of the undirected weighted graph of the pseudo-relationship between pixels tea :

[0059]

[0060] Step 4c, transform the K-classification problem into the problem of cutting the adjacency matrix W tea into the problem of K subgraphs, so as to obtain the target domain pixel-level pseudo-relationship loss.

[0061] Step 4c includes:

[0062] Step 4c1, flatten the output p of the k-th class through a transpose operation to obtain k

[0063] Step 4c2, calculate the edge w of all vertices inside the k-th subgraph k :

[0064] Step 4c3, calculate the edges between all vertices inside the k-th subgraph and the edges existing between other subgraphs

[0065]

[0066] Step 4c4, calculate the cost cut consumed when the k-th subgraph is cut k :

[0067] ​Step 4c5, calculate the total cost cut for partitioning into K subgraphs:

[0068] Step 4c6, add perturbations or noise to the target domain samples and then send them into the student model to generate their probability scores p stu ;

[0069] Step 4c7, calculate the target domain pixel-level pseudo-relationship loss L according to the following formula T :

[0070]

[0071] For target domain pixel-level relationship learning: There is usually similar semantic information between adjacent pixels, and the mapping relationship captured by it can enhance the generalization ability of the model. Since there is no label information for the target domain data, most current self-training methods focus on how to generate reliable pseudo-labels for the target domain, ignoring the relationship between pixels in the target domain. Therefore, the present invention constructs a target domain pixel-level relationship learning loss, and enables the model to obtain the ability to model the spatial relationship between pixels by optimizing the loss.

[0072] The construction of the source domain patch-level pseudo-relationship loss for source domain samples in Step 3 includes:

[0073] Step 5a, cut the source domain samples into 8*8 patches;

[0074] Step 5b, calculate the proportion c of each category in each patch according to the source domain label y s ; (k,pt) ;

[0075] Step 5c, perform clustering on each patch to obtain its category label y (s,pt) ;

[0076] Step 5d, calculate the patch-level pseudo-relationship loss L of the source domain (s,pt) :

[0077]

[0078] where C pt is the patch-level classifier.

[0079] For source domain local patch-level relationship learning: In image segmentation, the relationship learning between local patches can help us better understand the interaction and dependence between local patches, thereby improving the accuracy and robustness of segmentation. Therefore, the present invention hopes to extend pixel-level relationship learning to the patch level and realize the relationship learning between local image blocks by constructing a patch-level relationship learning loss.

[0080] Building the target domain patch-level pseudo-relationship loss in step 3 includes:

[0081] Step 6a: Cut the target domain image into 8*8 patches;

[0082] Step 6b: Feed the local patches into the teacher model and obtain the probability scores of the teacher model through the patch-level classifier

[0083] Step 6c: Similar to the source domain, construct the adjacency matrix of the undirected weighted graph of the pseudo-relationship between local patches

[0084] Step 6d: After adding perturbations or noises to the target domain patch images, feed them into the student model to generate their probability scores

[0085] Step 6e: Calculate the target domain patch-level pseudo-relationship loss

[0086]

[0087] For the target domain local patch-level relationship learning: The motivation is the same as that of the source domain local patch-level relationship learning. Since the source domain has labels and the target domain has no labels, the loss construction methods of the two are different. The former is achieved through labels, and the latter is achieved through the output of the teacher model.

[0088] The final optimization objective function obtained in step 3 for optimizing the multi-level pseudo-relationship network model is:

[0089]

[0090] where λ px , λ pt are hyperparameters for calculating their respective losses of patch-level pseudo-relationship learning in the source domain and the target domain.

[0091] Step 4: Repeat step 4 for each batch of training samples, and use the final optimization objective function to optimize the training process during the iteration to obtain a trained multi-level pseudo-relationship network model;

[0092] Step 5: Use the trained multi-level pseudo-relationship network model to extract the features of the optical remote sensing image to be classified, and classify according to the features to obtain the categories of the targets included in the optical remote sensing image to be classified.

[0093] The present invention provides an optical remote sensing image ground object classification device based on multi-level pseudo-relationship learning, including:

[0094] An acquisition module, configured to acquire an optical remote sensing image to be classified from a remote device and acquire a training set composed of multiple optical remote sensing images from a database;

[0095] Wherein, the training set includes source domain samples and target domain samples, the target domain samples are distinguished from the source domain samples by different geographical locations, and the source domain samples carry the target class labels included in the images;

[0096] A construction module, configured to construct a multi-level pseudo-relationship network model, the multi-level pseudo-relationship network model includes a teacher model and a student model with exactly the same structure and initialization parameters, and the parameters between the teacher model and the student model are not shared;

[0097] An optimization training module, configured to add perturbations or noises to the training samples and send them into the student model; send the training samples into the teacher model, and perform iterative training on the student model and the teacher model; during the iterative training process, construct a source domain relationship learning loss and a source domain patch-level pseudo-relationship loss for the source domain samples in sequence, and construct a target domain pixel-level pseudo-relationship loss and a target domain patch-level pseudo-relationship loss for the target domain samples, so as to obtain a final optimization objective function for optimizing the multi-level pseudo-relationship network model; extract the primary features and multi-scale features of the training samples through their respective models, and both output a first output result at the pixel level and a second output result at the patch level;

[0098] An iteration module, configured to repeat the training process of step 3 for each batch of training samples, and use the final optimization objective function to optimize the training process during the iteration to obtain a trained multi-level pseudo-relationship network model;

[0099] A classification module, configured to extract the features of the optical remote sensing image to be classified by using the trained multi-level pseudo-relationship network model, and classify according to the features to obtain the categories of the targets included in the optical remote sensing image to be classified.

[0100] The technical effects of the present invention are further described below in combination with simulation experiments:

[0101] 1. Simulation conditions and content:

[0102] The simulation experiment of the present invention is implemented based on the pytorch platform in a hardware environment with GPU: GeForce GTX 2080, memory 16G; CPU: Intel(R)Xeon(R)CPU E5-2603 v3, memory frequency 1.60GHz and a software environment of Ubuntu 16.04.

[0103] The datasets used for the training and testing of the network model of the present invention are the ISPRS 2D Semantic Labeling Challenge Dataset and the LoveDA Dataset. The present invention conducts ablation experiments on the number of clusters and different hyperparameters, and comparative experiments on pixel-level pseudo-relations and patch-level pseudo-relations in the above two different datasets.

[0104] 2. Analysis of simulation results:

[0105] In the experiments of the present invention, MIOU is selected as the experimental evaluation index. It can be seen from the experimental results of the two datasets that the traditional self-training method itself has advantages compared with the alignment-based method. Refer to Figure 4 , Figure 4 Figure 10 is the segmentation visualization result diagram of different methods under the ISPRS 2D Semantic Labeling Challenge Dataset. In the experiments of the ISPRS 2D Semantic Labeling Challenge Dataset, its MIOU index is approximately 18 higher than that of the alignment-based method, which indicates that the self-training strategy can better mine the specific knowledge of the target domain than the alignment strategy. In the experimental results, the learning results of pixel-level pseudo-relations all reach the optimal. In the transfer experiment of the ISPRS 2D Semantic Labeling Challenge Dataset, as shown in Table 1, the MIOU index of pixel-level pseudo-relation learning is 2 higher than that of the self-training method. Among them, the segmentation effects of three ground object categories, namely impervious surface, building, and low vegetation, reach the best. Among them, the MIOU indexes of the impervious surface and the building are 1.16 and 1.15 higher than those of the self-training method respectively, and the segmentation result of the low vegetation ground object category is improved significantly, and its MIOU index is 8.15 higher than that of the self-training method.

[0106] Table 1 Comparison table between the present invention and the traditional scheme in the ISPRS 2D Semantic Labeling Challenge Dataset

[0107] Method Imp.Sur Build Tree Car Lowvege MIOU Adaseg 52.32 43.64 50.63 9.26 28.97 36.96 BDL 42.23 39.09 51.19 8.82 32.10 34.68 Self-training 71.14 76.17 56.99 32.58 37.04 54.78 Source only 44.49 53.92 47.48 11.50 27.33 36.94 Pixel-level pseudo-relationship (Ours) 72.30 77.32 56.84 32.55 45.19 56.85

[0108] Figure 3 Figure 11 is the visualization result diagram under each method. Taking the transfer experiment from Potsdam to Vaihingen in the ISPRS 2D Semantic Labeling Challenge Dataset as an example, different rows correspond to different images, and different columns represent different experimental methods. From left to right, they represent the original image, label image, Adaseg segmentation result image, BDL segmentation result image, self-training segmentation result image, and multi-level pseudo-relation segmentation result of each experimental image. Among them, different colors represent different ground object categories. In the experiment from Potsdam to Vaihingen, there are 5 different ground object categories, which are white representing the ground surface, blue representing buildings, green representing trees, yellow representing vehicles, and representing low vegetation. It can be seen from the visualization result diagram that the method proposed in this chapter predicts more accurately for categories with non-close geographical positions on the image and is not easily affected by some noise and texture features compared with other methods.

[0109] In the transfer experiment on the LoveDA dataset, as shown in Table 2, the MIOU metric of the traditional self-training method is approximately 12 higher than that of the alignment-based method. The MIOU metric of pixel-level pseudo-relationship learning is 2.99 higher than that of the self-training method. Among them, the segmentation effects of the three land cover classes of buildings, roads, and barren land reach the best. The segmentation results of buildings and roads are improved significantly, and their MIOU metrics are 14.85 and 17.40 higher than those of the self-training method respectively. Since the LoveDA dataset collects rural and urban remote sensing images, the class differences are large. For classes such as buildings and roads, their distributions in rural remote sensing images are less. Especially for classes like roads without local rules, pseudo-relationship learning can better capture their long-distance relationships in the image. Also, because these classes are more distributed in urban images and belong to the head classes in the target domain, the classification effect of pseudo-relationship learning on them is better. It should be noted that classes such as forests and agricultural land have strong regional styles. Due to their huge differences in rural and urban distributions, when these classes belong to the long-tail classes in the target domain, the classification effect of the model on them will be greatly affected by the inter-domain differences.

[0110] Table 2 Comparison table of the present invention and traditional schemes on the LoveDA dataset

[0111] Method Bac grd Build Road Water Barren Forest Argicul Miou Adaseg 42.35 23.73 15.61 81.95 13.62 28.70 22.05 32.68 BDL 43.41 25.42 13.75 79.25 13.71 30.44 25.80 33.11 Self-training 36.74 32.21 31.21 67.71 42.66 44.15 58.25 44.70 Source only 38.53 28.19 17.76 65.54 40.75 42.72 56.08 41.37 Pixel-level relationship (Ours) 34.83 47.06 48.61 67.67 43.31 40.11 52.93 47.69

[0112] In the comparative experiment on the ISPRS 2D semantic labeling challenge dataset, as shown in Table 3, the MIOU metric of multi-level pseudo-relationship learning is on average approximately 1.2 higher than that of pixel-level pseudo-relationships. Except for the land cover class of vehicles, multi-level pseudo-relationship learning has varying degrees of improvement for other classes. After adding patch-level pseudo-relationship learning, although there is no improvement in the vehicle class, the fluctuation range is not large. This is because the pixel proportion of the vehicle class itself is small and the distribution is discrete, and vehicles cannot become the main class in local patch blocks, so the model cannot learn the patch-level relationships of the vehicle class. As can be seen from Table 4, in the transfer experiment on the LoveDA dataset, the MIOU metric of the multi-level pseudo-relationship learning method is 1.19 points higher than that of pixel-level pseudo-relationships, enhancing the segmentation effects of all classes in the dataset. The experimental results on the two datasets show that the relationship learning between local patches is also important for land cover classification, and the model performance has varying degrees of improvement when facing data with different distributions.

[0113] Table 3 Comparison table of multi-level and single-level pseudo-relationships on the ISPRS 2D semantic labeling challenge dataset

[0114] Method Imp.Sur Build Tree Car Low vege MIOU Pixel-level pseudo-relationship (Ours) 72.30 77.32 56.84 32.55 45.19 56.85 Multi-level pseudo-relationship (Ours) 73.17 79.45 58.03 32.54 47.07 58.04

[0115] Table 4 Comparison Table of Multi-level and Single-level Pseudo-relations in LoveDA Dataset

[0116] Method Bac grd Build Road Water Barren Forest Argicul Miou Pixel-level relationship (Ours) 34.83 47.06 48.61 67.66 43.31 40.11 52.93 47.69 Multi-level pseudo-relationship (Ours) 35.70 49.19 49.80 67.66 45.19 40.75 53.91 48.88

[0117] Table 5 shows the results when only local patch relationship learning is added to the source domain data. The experiments show that when the number of clustering centers is 10 clusters and the local patch hyperparameter is 0.1, the model has the best performance, and its MIOU score reaches 56.59. When the pathc-level relationship learning is not added to the source domain, the experimental index reaches 56.85. After only adding patch pseudo-level relationship learning to the source domain, the index drops slightly, indicating that only multi-level pseudo-relationship learning on the source domain cannot enhance the generalization performance of the model on the target domain.

[0118] Table 5 Result Display Table of Adding Local Patch Relationship Learning to the Source Domain Data of the Present Invention

[0119]

[0120] Table 6 shows the results after adding local patch relationship learning in the target domain. At this time, the patch loss hyperparameter in the source domain part is fixed at 1. The experiments show that when the number of clustering centers is 10 clusters and the local patch hyperparameter in the target domain is 0.5, the model has the best performance, and its MIOU score reaches 56.87, which is slightly better than the pixel-level pseudo-relationship result. At the same time, when the number of clusters is increased to 15 and 20, the experimental results show that when the hyperparameters in the source domain part are fixed, increasing the clustering centers does not significantly improve the performance.

[0121] Table 6 Result Display Table of Adding Local Patch Relationship Learning in the Target Domain of the Present Invention

[0122]

[0123] Table 7 shows the results when the hyperparameters of the local patch in the source domain are the same as those of the local patch in the target domain after adding local patch relationship learning in the target domain. The experiments show that when the number of clustering centers is 20 clusters and the local patch hyperparameter is 0.5, the model has the best performance, and its MIOU score reaches 58.04. When the parameter settings of the source domain and the target domain are the same, the increase in the number of clusters can improve the performance of the model.

[0124] Table 7 Display Table When the Hyperparameters of the Local Patch in the Source Domain are the Same as Those of the Local Patch in the Target Domain

[0125]

[0126] In summary, through the test experiments on the ISPRS 2D semantic label challenge dataset and the LoveDA dataset, the present invention shows that the proposed multi-level pseudo-relations do indeed make up for the deficiencies in existing self-training methods and greatly improve the generalization of the model in the target domain.

[0127] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0128] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality.

[0129] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An object classification method for optical remote sensing images based on multi-level pseudo-relationship learning, characterized in that, Including: Step 1: Obtain the optical remote sensing image to be classified from a remote device and obtain a training set composed of multiple optical remote sensing images from a database; Among them, the training set includes source domain samples and target domain samples. The target domain samples are distinguished from the source domain samples by different geographical locations, and the source domain samples carry the target category labels contained in the images; Step 2: Construct a multi-level pseudo-relationship network model. The multi-level pseudo-relationship network model includes a teacher model and a student model with exactly the same structure and initialization parameters, and the parameters between the teacher model and the student model are not shared; Step 3: Add perturbations or noises to the source domain samples and target domain samples and send them into the student model; send the source domain samples and target domain samples into the teacher model, and perform iterative training on the student model and the teacher model; during the iterative training process, construct a source domain relationship learning loss and a source domain patch-level pseudo-relationship loss for the source domain samples in sequence, and construct a target domain pixel-level pseudo-relationship loss and a target domain patch-level pseudo-relationship loss for the target domain samples, so as to obtain a final optimization objective function for optimizing the multi-level pseudo-relationship network model; extract the primary features and multi-scale features of the source domain samples and target domain samples through their respective models, and both output a first output result at the pixel level and a second output result at the patch level; Step 4: Repeat the training process of Step 3 for each batch of source domain samples and target domain samples, and use the final optimization objective function to optimize the training process during the iteration to obtain a trained multi-level pseudo-relationship network model; Step 5: Use the trained multi-level pseudo-relationship network model to extract the features of the optical remote sensing image to be classified, and classify according to the features to obtain the category of the target contained in the optical remote sensing image to be classified; Both the teacher model and the student model in Step 2 include a feature extractor, two ASPP networks, a patch-level classifier, and a pixel-level classifier; The output of the feature extractor is connected to the input of the two ASPP networks. The output of one ASPP network is connected to the input of the patch-level classifier, and the output of the other ASPP network is connected to the input of the pixel-level classifier.

2. The object classification method for optical remote sensing images based on multi-level pseudo-relationship learning according to claim 1, characterized in that, In Step 3, constructing a source domain relationship learning loss for the source domain samples includes: Step 3a, for the source domain samples in the training set, denoted as X S = x s , for the pixel-level labels of the source domain samples, denoted as Y S = y s ; Step 3b, for the target domain samples in the training set, denoted as X T = x t ; Step 3c: Calculate the source domain segmentation loss: Among them, the student model or the teacher model is denoted as G, the feature extractor is denoted as F, and the classifier is denoted as C.

3. The object classification method for optical remote sensing images based on multi-level pseudo-relationship learning according to claim 2, characterized in that, In Step 3, constructing a target domain relationship learning loss for the target domain samples includes: Step 4a, for the target domain sample x t , send it into the teacher model to generate its probability score p t ; Step 4b: Consider all pixels output by the teacher model as nodes, and consider the similarity S between nodes as the weight of the edge, and construct the adjacency matrix W of the undirected weighted graph of the pseudo-relations between pixels tea : Step 4c, convert the K-classification problem into a K-subgraph problem of cutting the adjacency matrix W tea to obtain the target domain pixel-level pseudo-relationship loss.

4. The object classification method for optical remote sensing images based on multi-level pseudo-relationship learning according to claim 3, characterized in that, Step 4c includes: Step 4c1, the output p of the k-th class k is flattened by a transpose operation to obtain Step 4c2, calculate the edge w of all vertices inside the k-th subgraph k : Step 4c3, calculate the edges existing between the edges of all vertices inside the k-th subgraph and other subgraphs Step 4c4, calculate the cost cut consumed when the k-th sub-graph is cut k : Step 4c6, add perturbations or noise to the target domain samples and then input them into the student model to generate its probability score p stu ; Step 4c7, calculate the pixel-level pseudo-relationship loss L for the target domain according to the following formula T :[[]]END]] 5. The object classification method for optical remote sensing images based on multi-level pseudo-relationship learning according to claim 4, characterized in that, In Step 3, constructing a source domain patch-level pseudo-relationship loss for the source domain samples includes: Step 5a: Cut the source domain samples into 8*8 patches; Step 5b, according to the source domain label y s calculate the proportion c of each category in each patch (k,pt) ; Step 5c, clustering each patch to obtain its class label y (s,pt ) Step 5d, calculate the patch-level pseudo-relationship loss L of the source domain (s,pt ) Among them, C pt is a patch-level classifier.

6. The method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning according to claim 5, wherein, In Step 3, constructing a target domain patch-level pseudo-relationship loss for the target domain includes: Step 6a: Cut the target domain image into 8*8 patches; Step 6b, sending the local patch into the teacher model and obtaining the probability score of the teacher model through the patch-level classifier Step 6c, construct an adjacency matrix for the undirected weighted graph of pseudo-relations between local patches similar to the source domain Step 6d, adding perturbations or noises to the target domain patch images and then feeding them into the student model to generate their probability scores Step 6e, calculate the target-domain patch-level pseudo-relationship loss 7. The method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning according to claim 6, wherein, The final optimization objective function for optimizing the multi-level pseudo-relationship network model obtained in Step 3 is: Among them, λ px , λ pt are hyperparameters for patch-level pseudo-relationship learning to calculate their respective losses on the source domain and the target domain.

8. The method for classifying ground objects in optical remote sensing images based on multi-level pseudo-relationship learning according to claim 7, wherein, During the iterative training process, the model changes its internal parameters, so as to achieve the optimization purpose of the final optimization objective function. The steps of updating the internal parameters include: Step 7a, the learning model updates the internal parameters through backpropagation; Step 7b, the teacher model updates the internal parameters by means of exponential moving average based on the parameters of the learning model; Among them, the parameters of the student model at time t are denoted as The parameters of the student model at time t - 1 are denoted as The parameters of the teacher model at time t are denoted as The parameters of the teacher model at time t - 1 are denoted as The formula for the moving average is as follows: where τ ∈ [0, 1].

Citation Information

Patent Citations

  • Deep learning method for predicting surface coverage category of label-free remote sensing image

    CN111898507A

  • High-resolution remote sensing image impervious surface extraction method based on deep learning

    CN113591608A