A method and device for remote sensing image feature recognition and classification based on weakly supervised learning
By adopting a weakly supervised learning method in the recognition and classification of remote sensing images, and using partial annotation data and pseudo-notation for model training, the problem of lack of annotation data is solved, and the recognition accuracy and robustness of the model are improved.
Patent Information
- Application Number
- CN202111421623.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-26
AI Technical Summary
Existing remote sensing image land object recognition and classification methods rely on a large amount of labeled data, resulting in high manual labeling costs and lack of labeled data, which limits the wide application of the method and the recognition accuracy.
Using a method based on weakly supervised learning, a machine learning model is established using partially labeled multi-source remote sensing images, and a pseudo-notation is generated through the teacher model, and a model training is carried out in combination with labeling and pseudo-notation to improve the accuracy of land objects recognition and classification.
It significantly improves the accuracy of land objects identification and classification, reduces the need for manual annotation, saves labor costs, and improves the generalization and robustness of the model.
Smart Images

Figure CN114399686B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of geographic information, ecological environment science, and remote sensing technology. Specifically, it relates to a method and device for identifying and classifying ground objects in remote sensing images based on weakly supervised learning. Background Art
[0002] The identification and classification of ground objects in remote sensing images mainly utilize images obtained from aerial or satellite earth observations. Through a machine learning model, the category to which each pixel in the image belongs is identified, thereby realizing land type identification, forest change monitoring, road extraction, building detection, etc. It has a wide range of applications in fields such as resource investigation, land management, urban planning, and topographic mapping, and is of great significance for the sustainable development of humanity.
[0003] Current methods for identifying and classifying ground objects in remote sensing images are mainly based on supervised learning methods. By using remotely sensed images with labeled pixel categories to train a machine learning model, and using the trained model to classify each pixel in the unlabeled image, the identification and classification of ground objects in remote sensing images are realized. The supervised learning method requires a large amount of labeled data for model training, and manually labeling each pixel of a large number of remote sensing images requires huge human and material resources. Therefore, there is a significant lack of high-quality labeled remote sensing images in actual application scenarios, which makes it difficult to effectively improve the accuracy of identifying and classifying ground objects in remote sensing images and limits the widespread application of this method. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for identifying and classifying ground objects in remote sensing images based on weakly supervised learning. The present invention uses partially labeled multi-source remote sensing images to establish a machine learning model and uses the established model to identify ground object types, significantly improving the accuracy of identifying and classifying ground object elements.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] A method for identifying and classifying ground objects in remote sensing images based on weakly supervised learning, the steps of which include:
[0007] 1. Read partially labeled multi-source remote sensing images and construct a labeled sample data set and an unlabeled sample data set;
[0008] 2. Establish a labeled training set and a labeled validation set from the labeled sample data set;
[0009] 3. Establish a teacher model and a student model;
[0010] 4. Input the labeled training set and the labeled validation set, and pre-train the teacher model to obtain a trained teacher model;
[0011] 5. Input the unlabeled sample dataset into the trained teacher model to obtain the prediction results of the unlabeled data as pseudo-labels;
[0012] 6. Read the unlabeled sample dataset and the pseudo-labels to construct a pseudo-labeled training set;
[0013] 7. Input the labeled training set, the labeled validation set, and the pseudo-labeled training set, perform random data augmentation, and train the student model;
[0014] 8. Use the student model as the new teacher model and repeat steps 5 to 7;
[0015] 9. Input the prediction dataset into the trained student model to obtain the results of ground object recognition and classification.
[0016] Further, the multi-source remote sensing images in step 1 include radar remote sensing data and / or optical remote sensing data. Preferably, the multi-source remote sensing images include at least 1000 remote sensing images.
[0017] Further, the radar remote sensing data in step 1 includes ground images obtained by synthetic aperture radar (SAR), etc. The storage file format of the images includes GeoTIFF, JPG, etc. The width of each image is W pixels, the height is H pixels, and the resolution is R. Each image includes one or more channels, and the number of channels is C R .
[0018] Further, the optical remote sensing data in step 1 is a ground image obtained by an optical sensor such as a CCD, and includes one or more spectral bands with different wavelengths such as panchromatic, visible light, near-infrared, short-wave infrared, and thermal infrared. Among them, the visible light includes one or more visible spectral bands with different wavelengths such as red, green, and blue. The storage file format of the images is GeoTIFF, JPG, HDF, NetCDF, etc. The width of each image is W pixels, the height is H pixels, and the resolution is R. Each image includes one or more channels, and the number of channels is C O . Each channel corresponds to a spectral band. Preferably, the optical remote sensing data includes at least visible light and near-infrared spectral bands.
[0019] Further, the partially labeled multi-source remote sensing images in step 1 are a set of multiple input images, and the storage format of the image files is GeoTIFF, PNG, JPG, etc. Each image X includes multiple channels, which are stacked by the channels of the radar remote sensing image X 1 and the optical remote sensing image X 2 with the number of channels being C R +C O . Among them, I 1The input image A is labeled to obtain the corresponding labeled image A', and its storage file format is GeoTIFF, PNG, JPG, etc. Each labeled image includes one channel, and each pixel value therein represents the class label of the geographical area range corresponding to the pixel. The input image A and its corresponding labeled image A' are used as the labeled sample data set, and the remaining I 2 input images B are used as the unlabeled sample data set.
[0020] Furthermore, there are I 1 groups of images in the labeled sample data set described in step 2. Randomly select n t groups of images and set them as the labeled training set, and the remaining I 1 -n t groups of images are set as the labeled validation set, where 1 < n t <I 1 . The images in the labeled training set and the labeled validation set are not repeated. Preferably, the labeled training set includes at least I 1 *80% groups of images, and the labeled validation set includes at least I 1 *10% groups of images.
[0021] Furthermore, the teacher model and the student model described in step 3 are machine learning models, and their model structures can be the same or different. The input data of the model is the input images in the labeled sample data set and the unlabeled sample data set described in step 1; the output result is an image with the same size as the input image, and the number of channels is the same as the number of predicted classes. Each pixel value therein represents the confidence that the geographical area range corresponding to the pixel belongs to each class.
[0022] Furthermore, the output results of the teacher model and the student model for the i-th input image x i are respectively expressed as: Among them, the function t represents the teacher model, and the function s represents the student model.
[0023] Furthermore, step 4 includes the following steps:
[0024] (1) Randomly read m groups of images (1 ≤ m ≤ n t ) without repetition from the labeled training set, use the teacher model to calculate the output result, and use the labeled image to calculate the objective function value;
[0025] (2) Update the model parameters according to the objective function value;
[0026] (3) Repeat the above steps (1) to (2), each time randomly reading m groups of images without repetition from the labeled training set, calculating the output result and the objective function value, and optimizing the model parameters until all the images in the labeled training set complete one training;
[0027] (4) Read the labeled validation set, use the teacher model to calculate the prediction results, and use the labeled images to calculate the evaluation metrics;
[0028] (5) Repeat the above steps (1) to (4), read the labeled training set, calculate the output results and the objective function values; optimize the model parameters; read the labeled validation set, calculate the prediction results and the evaluation metrics until the termination conditions are met. The termination conditions are at least one of the following: the model evaluation metrics reach the expectation, the number of iterations is greater than the maximum number of iterations.
[0029] Further, the objective function described in step 4 is defined as: where: m is the number of samples in a training batch, L is the training loss function, R is the regularization term, y i is the labeled image corresponding to the i-th input image, is the output result of the model for the i-th input image. The regularization term includes L1 regularization, L2 regularization, etc. The objective function may not contain the regularization term. Preferably, the training loss function is the cross-entropy loss function without the regularization term.
[0030] Further, the model evaluation metrics described in step 4 include at least one of the following: sensitivity (Recall), specificity (Specificity), precision (Precision), accuracy (Accuracy), intersection over union (IoU), F1 score, Dice coefficient, Jaccard coefficient, error rate, etc. For class c, the pixels of the image are divided into positive samples and negative samples. The pixels belonging to class c are positive samples, and the pixels not belonging to class c are negative samples; the number of pixels labeled as positive samples and predicted as positive samples is TP, the number of pixels labeled as positive samples and predicted as negative samples is FN, the number of pixels labeled as negative samples and predicted as positive samples is FP, and the number of pixels labeled as negative samples and predicted as negative samples is TN. The sensitivity is defined as: TPR = TP / (TP + FN); the specificity is defined as: TNR = TN / (TN + FP); the precision is defined as: PPV = TP / (TP + FP); the accuracy is defined as: ACC = (TP + TN) / (TP + TN + FP + FN); the F1 score and the Dice coefficient are the same, and their definition is: F1 = Dice = 2TP / (2TP + FP + FN); the intersection over union and the Jaccard coefficient are the same, and their definition is: IoU = Jaccard = TP / (TP + FP + FN); the error rate is defined as: Err = C err / C total where C err is the total number of pixels with prediction errors, C totalis the total number of pixels. Preferably, the model evaluation metric is the mean intersection over union of all classes, and the termination condition is that the mean intersection over union of the labeled validation set reaches the maximum.
[0031] Further, the pseudo-label in step 5 is the prediction result B' of the trained teacher model for each input image B in the unlabeled sample dataset I 2 The prediction result B' can be the class label to which each pixel in the input image B belongs, or the confidence of the class label to which it belongs. Preferably, the prediction result B' is the class label to which each pixel in the input image B belongs.
[0032] Further, the pseudo-labeled training set in step 6 is a set of I 2 groups of images, each group including 2 images, namely the input image B and the pseudo-label B'.
[0033] Further, step 7 includes the following steps:
[0034] (1) Combine the labeled training set and the pseudo-labeled training set as the student training set.
[0035] (2) Randomly and without repetition read m' groups of images (1 ≤ m' ≤ n t +I 2 ) from the student training set. After randomly augmenting these images, use the student model to calculate the output result, and calculate the objective function value using the labeled images and the pseudo-labels;
[0036] (3) Update the model parameters according to the objective function value;
[0037] (4) Repeat the above steps (2) to (3), each time randomly and without repetition read m groups of images from the student training set, calculate the output result and the objective function value, and optimize the model parameters until all the images in the student training set complete one training.
[0038] (5) Read the labeled validation set, use the student model to calculate the prediction result, and calculate the evaluation metric using the labeled images;
[0039] (6) Repeat the above steps (2) to (5), read the student training set, calculate the output result and the objective function value; optimize the model parameters; read the labeled validation set, calculate the prediction result and the evaluation metric, until the termination condition is met. The termination condition is at least one of the following: the model evaluation metric reaches the expectation, the number of iterations is greater than the maximum number of iterations. Preferably, the model evaluation metric is the mean intersection over union of all classes, and the termination condition is that the mean intersection over union of the labeled validation set reaches the maximum.
[0040] Further, the random data augmentation described in step 7 includes image processing methods such as image rotation, shearing, flipping, automatic contrast, equalization, color perturbation, brightness perturbation, image sharpening, and blurring.
[0041] Further, in step 8, if the evaluation index of the student model is better than that of the teacher model, the student model is used as the new teacher model, and steps 5 to 7 are repeated until the evaluation index of the student model reaches the maximum.
[0042] Further, the prediction dataset described in step 9 includes radar remote sensing data and optical remote sensing data for prediction, and each image therein has the same width, height, resolution, storage file format, and number of channels as the input image in the sample dataset described in step 1.
[0043] Further, the result of the ground object recognition and classification described in step 9 is an image corresponding one-to-one to each image in the prediction dataset, with the same width, height, and resolution as the input image. Each image includes one channel, and each pixel value in the image represents the prediction result of the category label of the geographical area range corresponding to the pixel.
[0044] A device for ground object recognition and classification of remote sensing images based on weak supervision learning, comprising:
[0045] A sample dataset acquisition unit, configured to read multi-source remote sensing images and construct a sample dataset using radar remote sensing data and optical remote sensing data;
[0046] A training and validation data establishment unit, configured to establish a training dataset and a validation dataset according to the sample dataset;
[0047] A model setting unit, configured to establish a teacher model and a student model;
[0048] A model training unit, configured to input the training dataset and the validation dataset, and train the teacher model and the student model to obtain the trained models;
[0049] A ground object type recognition unit, configured to input the prediction dataset into the trained student model to obtain the recognition result of the ground object type.
[0050] An electronic device, comprising a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the above-mentioned method.
[0051] Compared with the prior art, the positive effects of the present invention are:
[0052] The method provided by the present invention uses remote sensing images to perform intelligent recognition of ground object types, generates pseudo-labeled images on unlabeled images using a pre-trained teacher model, achieves the purpose of augmenting labeled data, overcomes the difficulty of lacking labeled data in ground object classification of remote sensing images, eliminates the need for a large amount of manual labeling, and saves huge labor costs and expenses. Moreover, a student model jointly trained with pseudo-labeled images and labeled images is used to replace the teacher model to generate pseudo-labeled images of higher quality, effectively improving the recognition ability and classification accuracy of the model. At the same time, random data augmentation is used to train the student model, significantly improving the generalization ability of the model and its robustness and stability to noise, with good effects and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of a weakly supervised learning framework for ground object classification of remote sensing images provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The present invention will be further described below through specific embodiments in conjunction with the drawings.
[0055] The process framework of a method for ground object recognition and classification of remote sensing images based on weakly supervised learning in this embodiment is as Figure 1 shown. Taking the land type recognition using Sentinel-1 satellite SAR radar data and Sentinel-2 satellite multispectral data as an example, a detailed description will be given below.
[0056] First step, read part of the labeled multi-source remote sensing images, and establish a labeled sample data set and an unlabeled sample data set. The multi-source remote sensing images in this embodiment include Sentinel-1 satellite SAR radar image data from 2016 to 2017 and Sentinel-2 satellite multispectral image data. Among them, the Sentinel-1 satellite SAR radar images include 2 channels of VV and VH, and the Sentinel-2 satellite multispectral images include 13 channels such as visible light, near-infrared, and short-wave infrared. The input images include 15 channels. The first to second channels are Sentinel-1 satellite SAR radar images, and the third to fifteenth channels are Sentinel-2 satellite multispectral images. The unlabeled sample data set includes 180,662 groups of images, and each group of images includes 1 input image. The labeled sample data set includes 6,114 groups of images, and each group of images includes 2 images, namely the input image and the labeled image. The labeled image is a single-channel land classification data image. Each image has a width of 256 pixels, a height of 256 pixels, a resolution of 10m, and the image file format is GeoTIFF.
[0057] Second step, the labeled sample data set obtained in the first step includes 6,114 groups of images. Randomly select 10% of the groups of images and set them as the labeled validation set x’, about 611 groups of image data; the remaining 5,503 groups of images are set as the labeled training set x.
[0058] In the third step, a teacher model and a student model are established. The model structure uses a UNet encoder-decoder architecture. Among them, the encoder of the teacher model uses a ResNet-RS-101 residual network structure, and the encoder of the student model uses a ResNet-RS-152 residual network structure.
[0059] In the fourth step, the labeled training set x and the labeled validation set x' are used to train the teacher model to obtain a trained teacher model. The training loss function is the cross-entropy loss function without a regularization term. In other embodiments of the present invention, other forms of loss functions and regularization terms can also be used. The specific steps of the training process are as follows:
[0060] (1) Randomly read 16 groups of images without repetition from the labeled training set x, and calculate the output result and the objective function value;
[0061] (2) Update the model parameters;
[0062] (3) Repeat the above steps (1) to (2) until one training of the entire training data set is completed;
[0063] (4) Read the labeled validation set x', and calculate the prediction result and the accuracy;
[0064] (5) Repeat the above steps (1) to (4), read the labeled training set, calculate the output result and the objective function value; optimize the model parameters; read the labeled validation set, calculate the prediction result and the mean intersection over union until the mean intersection over union reaches the maximum value or the number of iterations is greater than 1000 times.
[0065] In the fifth step, the trained teacher model is used to input the unlabeled sample data set. The model reads the input image and outputs the prediction result of the unlabeled data, that is, the land type to which each pixel in the unlabeled input image belongs, as a pseudo-label.
[0066] In the sixth step, the unlabeled sample data set and the pseudo-label are read to establish a pseudo-labeled training set x", including 180,662 groups of images. Each group of images includes 2 images, namely the input image in the unlabeled sample data set and the pseudo-label.
[0067] In the seventh step, the labeled training set x, the labeled validation set x', and the pseudo-labeled training set x" are used to train the student model to obtain a trained student model. The training loss function is the cross-entropy loss function without a regularization term. The model evaluation index is the mean intersection over union. In other embodiments of the present invention, other forms of loss functions, regularization terms, and evaluation indexes can also be used. The specific steps of the training process are as follows:
[0068] (1) Combine the labeled training set \(x\) and the pseudo-labeled training set \(x''\) as the student training set, which includes 186,165 groups of images.
[0069] (2) Randomly read 16 groups of images without repetition from the student training set and perform random data augmentation, including: image rotation, horizontal shear, vertical shear, horizontal flip, and vertical flip. Calculate the output result of the student model and the objective function value;
[0070] (3) Update the model parameters;
[0071] (4) Repeat the above steps (2) to (3) until one training of the entire training data set is completed;
[0072] (5) Read the labeled validation set \(x'\) and calculate the prediction result and the mean intersection over union;
[0073] (6) Repeat the above steps (2) to (5), read the student training set, calculate the output result and the objective function value; optimize the model parameters; read the labeled validation set, calculate the prediction result and the mean intersection over union until the mean intersection over union reaches the maximum value or the number of iterations is greater than 1000 times.
[0074] In the eighth step, if the mean intersection over union of the student model on the labeled validation set is better than that of the teacher model, then use the student model as the new teacher model and repeat the fifth step to the seventh step until the mean intersection over union of the student model on the labeled validation set reaches the maximum.
[0075] In the ninth step, use the trained student model to input the prediction data set, that is, a set of input images, where each image includes 15 channels. The 1st - 2nd channels are Sentinel - 1 satellite SAR radar images, and the 3rd - 15th channels are Sentinel - 2 satellite multispectral images. Each image has a width of 256 pixels, a height of 256 pixels, a resolution of 10m, and the image file format is GeoTIFF. The model reads the input image and outputs the recognition result of the land type.
[0076] According to the above - mentioned embodiments, training the model can achieve the following improvement effects: Compared with the teacher model trained only on the labeled data set, the student model trained with weak - supervised learning on the labeled data set and the unlabeled data set has the average prediction accuracy of the land type on the validation data set increased to 97.6% and the intersection over union increased to 77.4%.
[0077] Based on the same inventive concept, another embodiment of the present invention provides a remote - sensing image ground object recognition and classification device based on weak - supervised learning, which includes:
[0078] A sample data set acquisition unit for reading multi - source remote - sensing images and constructing a sample data set using radar remote - sensing data and optical remote - sensing data;
[0079] A training and validation data establishment unit for establishing a training data set and a validation data set according to a sample data set;
[0080] A model setting unit for establishing a teacher model and a student model;
[0081] A model training unit for inputting the training data set and the validation data set to train the teacher model and the student model to obtain a trained model;
[0082] A ground feature type recognition unit for inputting a prediction data set into the trained student model to obtain an identification result of the ground feature type.
[0083] Based on the same inventive concept, another embodiment of the present invention provides an electronic device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, and the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the method of the present invention.
[0084] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc), and the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, each step of the method of the present invention is implemented.
[0085] In the specific steps of the solution of the present invention, there may be other alternative ways or deformation ways, for example:
[0086] 1. In step one, in addition to reading multi-source remote sensing images, digital elevation DEM data can also be read.
[0087] 2. In step two, in addition to establishing a training set and a validation set, a test set can also be established. Randomly extract n t groups of images from the labeled sample data set and set them as the training set, n v groups of images are set as the validation set, and the remaining images are set as the test set. The images in the training set, the validation set, and the test set do not repeat.
[0088] 3. For the teacher model and the student model in step three, machine learning models of other types such as support vector machines, random forests, and gradient boosting trees, as well as deep learning semantic segmentation models with other structures, can also be used.
[0089] 4. The training loss function in step four may also include the model evaluation metrics, that is: F1 score, Dice coefficient, intersection over union, Jaccard coefficient, etc.
[0090] 5. In step five, the confidence of each pixel in the unlabeled input image belonging to the land type by the teacher model can also be used as a pseudo label.
[0091] 6. In step seven, other image processing methods such as automatic image contrast, histogram equalization, color perturbation, brightness perturbation, sharpening, and blurring can also be used for random data augmentation.
[0092] 7. In steps seven and eight, other evaluation metrics such as sensitivity, specificity, accuracy, F1 score, Dice coefficient, Jaccard coefficient, and error rate can also be used.
[0093] 8. In step nine, a test set can also be input into the trained model to obtain the prediction results and test accuracy of the model.
[0094] Obviously, the embodiments described above are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art fall within the scope of protection of the present invention.
Claims
1. A method for identifying and classifying ground objects in remote sensing images based on weakly supervised learning, characterized in that, it includes the following steps: 1) Read partially annotated multi-source remote sensing images to construct an annotated sample data set and an unannotated sample data set; 2) Establish an annotated training set and an annotated validation set from the annotated sample data set; 3) Establish a teacher model and a student model; 4) Input the annotated training set and the annotated validation set to pre-train the teacher model to obtain a trained teacher model; 5) Input the unannotated sample data set into the trained teacher model to obtain the prediction results of the unannotated data as pseudo-labels; 6) Read the unannotated sample data set and the pseudo-labels to construct a pseudo-annotated training set; 7) Input the annotated training set, the annotated validation set, and the pseudo-annotated training set, perform random data augmentation, and train the student model; 8) Use the student model as the new teacher model and repeat steps 5) to 7); 9) Input the prediction data set into the trained student model to obtain the results of ground object identification and classification; Step 1) The partially annotated multi-source remote sensing images are a set of multiple input images. Each image X includes multiple channels and is formed by stacking the channels of a radar remote sensing image X 1 and an optical remote sensing image X 2 corresponding to the same geographical area range; I 1 of the input images A are annotated to obtain the corresponding annotated images A'. Each annotated image includes one channel, and each pixel value therein represents the class label of the geographical area range corresponding to the pixel; the input image A and its corresponding annotated image A' are used as the annotated sample data set, and the remaining I 2 input images B are used as the unannotated sample data set; Step 4) includes: (1) Randomly read m groups of images without repetition from the labeled training set, calculate the output results using the teacher model, and calculate the objective function value using the labeled images; (2) Update the model parameters according to the objective function value; (3) Repeat the above steps (1) to (2), each time randomly reading m m groups of images without repetition, calculate the output result and the objective function value, and optimize the model parameters until all the images in the labeled training set are completed for one training; (4) Read the annotated validation set, calculate the prediction results using the teacher model, and calculate the evaluation metrics using the annotated images; (5) Repeat the above steps (1) to (4), read the annotated training set, calculate the output results and the objective function value; optimize the model parameters; read the annotated validation set, calculate the prediction results and the evaluation metrics until the termination condition is met; The termination condition is at least one of the following: the model evaluation metric reaches the expectation, the number of iterations is greater than the maximum number of iterations; Step 7) includes: (1) Combine the annotated training set and the pseudo-annotated training set as the student training set; (2) Randomly read m’ groups of images without repetition from the student training set. After performing random data augmentation on these images, use the student model to calculate the output results, and calculate the objective function value using the labeled images and pseudo-labels; (3) Update the model parameters according to the objective function value; (4) Repeat the above steps (2) to (3), each time randomly reading m m groups of images without repetition from the student training set, calculate the output result and the objective function value, and optimize the model parameters until all the images in the student training set complete one training; (5) Read the annotated validation set, calculate the prediction results using the student model, and calculate the evaluation metrics using the annotated images; (6) Repeat the above steps (2) to (5), read the student training set, calculate the output results and the objective function value; optimize the model parameters; read the annotated validation set, calculate the prediction results and the evaluation metrics until the termination condition is met; In step 9), the results of the ground object identification and classification are images corresponding one by one to each image in the prediction data set, with the same width, height, and resolution as the input images. Each image includes one channel, and each pixel value in the image represents the prediction result of the class label of the geographical area range corresponding to the pixel; In step 8), if the evaluation metric of the student model is better than that of the teacher model, use the student model as the new teacher model and repeat steps 5) to 7) until the evaluation metric of the student model reaches the maximum.
2. The method according to claim 1, characterized in that, the multi-source remote sensing images in step 1) include radar remote sensing data and / or optical remote sensing data; the radar remote sensing data includes ground images obtained by synthetic aperture radar; the optical remote sensing data is ground images obtained by optical sensors, including one or more spectral bands of different wavelengths among panchromatic, visible light, near infrared, shortwave infrared, and thermal infrared.
3. The method according to claim 1, characterized in that, Randomly select n t groups of images from the labeled sample dataset and set them as the labeled training set, and the remaining I 1 -n t groups of images are set as the labeled validation set, where 1 < n t <I 1 , and the images in the labeled training set and the labeled validation set are not repeated.
4. The method according to claim 1, wherein, the teacher model and the student model in step 3) are machine learning models, and their model structures are the same or different.
5. The method according to claim 1, wherein, The objective function described in step 4) is defined as: , where: M is the number of samples in a training batch, L is the training loss function, R is the regularization term, y i is the annotated image corresponding to the i th input image, is the output result of the model for the i th input image; The evaluation metrics described in step 4) include at least one of the following: sensitivity, specificity, precision, accuracy, intersection over union, F1 score, Dice coefficient, Jaccard coefficient, error rate.
6. The method according to claim 1, wherein, The pseudo-label in step 5) is the prediction result B' of the trained teacher model for each input image B in the unlabeled sample dataset I 2 The prediction result B' is the class label to which each pixel in the input image B belongs, or the confidence of the class label; the pseudo-labeled training set in step 6) is I 2 A set of groups of images, each group including 2 images, namely the input image B and the pseudo-label B'.
7. A remote sensing image ground object recognition and classification device based on weak supervision learning adopting the method according to any one of claims 1 to 6, wherein, comprising: a sample data set acquisition unit, configured to read multi-source remote sensing images and construct a sample data set using radar remote sensing data and optical remote sensing data; a training and validation data establishment unit, configured to establish a training data set and a validation data set according to the sample data set; a model setting unit, configured to establish a teacher model and a student model; a model training unit, configured to input the training data set and the validation data set to train the teacher model and the student model to obtain a trained model; a ground object type recognition unit, configured to input a prediction data set to the trained student model to obtain an identification result of the ground object type.
8. An electronic device, wherein, comprising a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Semantic segmentation-based surface feature recognition and classification method and device
CN112464745A
Real-time colonoscope polyp detection device based on deep learning
CN112686856A
Data classification and identification method and device, equipment and readable storage medium
CN112949786A