SAR Image Ground Object Classification Method Based on Self-Supervised Jigsaw Learning
Through the self-supervised puzzle learning method, the pre-trained model is trained using the labelless data design puzzle task, and migrated to the SAR image geographic classification, solving the problems of insufficient sample and insufficient discriminant land feature characteristics, and improving classification performance.
Patent Information
- Application Number
- CN202211374754.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-04
AI Technical Summary
The existing SAR image land object classification method requires a large number of labeled samples and the diverse representations of land object targets lead to insufficient discrimination of features. The self-supervised learning method fails to effectively use labelless data to improve feature representation capabilities.
The self-supervised puzzle learning method is adopted, and the design puzzle task is trained on labelless data as an upstream task. The data characteristics are mined by self-supervised learning, and the upstream task pre-training model is migrated to the downstream scene classification task. A small number of labeled samples are used for fine-tuning to improve feature extraction capabilities.
Better classification performance was achieved in downstream scenario classification tasks, which alleviated the shortcomings of insufficient samples and the diversity of land objects targets, and improved the discriminant and classification accuracy of land objects characteristics.
Smart Images

Figure CN115527071B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent interpretation of radar remote sensing images, and specifically refers to a method for classifying ground objects in SAR images based on self-supervised puzzle learning. Background Technique
[0002] Synthetic Aperture Radar (SAR) is a high-resolution active microwave remote sensing imaging radar, which has the characteristics of all-weather, all-day, not affected by weather conditions such as clouds, rain, and fog, and can penetrate the earth's surface structures such as the ground and leaf clusters. It can effectively detect targets under various camouflage conditions and achieve effective observation in harsh environments. Based on these advantages, SAR images are widely used in many fields, such as environmental protection, disaster monitoring, and geographical mapping. Ground object classification in SAR images is a basic step in the interpretation of synthetic aperture radar. How to effectively extract features is the key to solving the image classification problem. Traditional classification methods, such as support vector machines and random forest classifiers, need to manually extract features and require strong manual experience guidance. Automated deep learning methods require a large number of labeled samples for feature extraction. Using a Convolutional Neural Network (CNN) to solve the problem of ground object classification in SAR images is an idea. The convolutional layer can extract special features from the image without losing other spatial information, and the weight sharing attribute and pooling layer are used to reduce the training parameters required by the network. By designing a multi-layer neural network, the convolutional layer, pooling layer, and activation function layer are connected layer by layer, and finally the ground object classification in SAR images is completed through the fully connected layer, and the effect is further verified in several real large scenes. The disadvantages are that currently, SAR image data is scarce, and the cost of manually labeled data is very high, which will consume a large amount of manpower and financial resources. Moreover, this method exposes the disadvantage of insufficient discriminability of captured ground objects. Therefore, how to make full use of a small amount of labeled data to learn effective ground object feature representations has become a relatively popular solution idea. Self-supervised learning mainly uses auxiliary tasks to mine its own supervision information from large-scale unlabeled data to obtain a pre-trained model, and then fine-tune it for downstream tasks (classification, segmentation, object detection, etc.). The downstream tasks can obtain better results by only using a small amount of labeled data.
[0003] The core of self-supervised learning is how to automatically generate labels for input data. According to the different designs of upstream tasks, self-supervised learning can be mainly divided into three categories: context-based, time-series-based, and contrast-based. The jigsaw task as a method of upstream task belongs to the context-based category. The image is divided into several blocks, and multiple permutation methods are predefined. Any randomly shuffled permutation is input, and it is expected to learn which of the multiple permutation methods this permutation belongs to, regarding it as a classification problem. By solving the jigsaw task, the relationships between individual image blocks are learned, capturing the semantic features of the scene space. Psychological research also shows that the jigsaw task can be used to evaluate human visual spatial processing. Summary of the Invention
[0004] To overcome the above-mentioned shortcomings of the prior art, the present invention provides a method for classifying ground objects in SAR images based on self-supervised jigsaw learning, which improves the feature extraction ability of the classification model and realizes the effective representation of ground object features; it can alleviate the deficiencies of insufficient labeled training samples and insufficient discriminability of ground object features caused by diverse forms of ground object targets, and improve the classification performance of the scene classification model.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A method for classifying ground objects in SAR images based on self-supervised jigsaw learning uses the idea of self-supervised learning to mine the characteristics of unlabeled data itself as supervision information to realize the effective representation of ground object features; designs a jigsaw as an upstream jigsaw task to train an upstream task pre-trained model on unlabeled data, migrates the upstream task pre-trained model and fine-tunes the downstream scene classification task, uses a small number of labeled samples in the downstream scene classification task, and tests the classification performance of the scene classification model in several real large scenes.
[0007] A method for classifying ground objects in SAR images based on self-supervised jigsaw learning includes the following steps:
[0008] Step 1, input the SAR image, and randomly intercept image blocks of a fixed size for each type of area to form a training set and a validation set;
[0009] Step 2, randomly shuffle the list [1, 2, 3, 4, 5, 6, 7, 8, 9], calculate the maximum Hamming distance after shuffling, and select the top eighty permutations with the maximum Hamming distance as the permutation set;
[0010] Step 3, cut the image blocks obtained in Step 1 into 3x3 small blocks, and adopt a cutting method with gaps during cutting. Normalize the 9 obtained image blocks individually.
[0011] Step 4: Randomly select a permutation from the set of permutations obtained in Step 2 as the permutation order, and send the 9 image patches after data processing in Step 3 into the upstream task network for prediction according to the selected permutation order, with the goal of obtaining the index of the selected permutation.
[0012] Step 5: Randomly select 20 images from each category in the training set samples obtained in Step 1 as the training samples for the downstream scene classification task and perform data processing; use the upstream jigsaw task model as the pre-trained model for the downstream scene classification task for transfer learning, and fine-tune the downstream scene classification task.
[0013] Step 6: Cut the complete SAR image into small blocks of the same size as in Step 1 according to the covered cutting method, classify them using the downstream scene classification task model, assign specific colors to specific categories after classification, and splice them into a large image according to the cutting method.
[0014] Step 7: For the unknown classes in the SAR image, find the corresponding regions of the unknown classes on the label map, cover the corresponding regions on the result map with the colors of the unknown classes, and convert the result map into an index map.
[0015] Step 8: Calculate the evaluation metrics between the obtained result map and the label map of the SAR image, calculate the evaluation metrics CPA, Recall, F1score for each category, and the overall evaluation metrics PA, Kappa, MIoU, FWIoU.
[0016] The present invention has the following advantages compared with the existing technologies for SAR image ground object classification:
[0017] 1. The present invention utilizes the idea of self-supervised learning, designs a feature learning method with the jigsaw task as the upstream task, and mines the inherent representation characteristics of the data as supervision information with the self-supervised idea to enhance the feature extraction ability of the scene classification model, realize the effective representation of ground object features, and enable the downstream scene classification task to learn effective feature representations on a better basis. Compared with the method without using a pre-trained model when using a small number of labeled samples in the downstream scene classification task, better results are obtained, alleviating the defects of SAR images requiring a large number of training samples and the lack of discriminability of ground object features due to the diverse manifestation forms of ground object targets.
[0018] 2. The present invention designs a context-free network, uses the operation of sharing convolutional layer weights for each input image patch in the upstream task network, and designs a better block division method to solve the SAR image jigsaw task.
[0019] 3. The present invention designs a unique set of jigsaw permutations and an independent method for normalizing image patch data to prevent solving problems by taking shortcuts in the process of solving the upstream jigsaw task. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is the flow chart of the present invention.
[0021] Figure 2 is the network structure diagram of the upstream jigsaw task of the present invention.
[0022] Figure 3 is the network structure diagram of the downstream scene classification task of the present invention.
[0023] Figure 4 is the real SAR image used in the present invention.
[0024] Figure 5 is the SAR image label map used in the present invention.
[0025] Figure 6 is the final result map of the SAR image obtained in the present invention.
[0026] Figure 7 is the SAR image result map of the comparative experiment in the present invention. Detailed implementation manners
[0027] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.
[0028] Refer to Figure 1 , a method for classifying ground objects in SAR images based on self-supervised jigsaw learning, comprising the following steps:
[0029] Step 1, input the SAR image, and randomly intercept image patches of a fixed size for each type of region to form a training set and a validation set;
[0030] In this embodiment, the SAR images required for the experiment are input. Taking the SAR image of the Napoli area as an example, 500 image patches are taken from each of the four types of regions of water area, forest, building, and farmland according to the sizes of 50x50, 75x75, 100x100, 120x120, 150x150, and 200x200 to explore the influence of the image size on the ground object classification of SAR images. For each group of experiments, the training set and the validation set are divided according to the ratio of 8:2;
[0031] Step 2, randomly shuffle the list [1, 2, 3, 4, 5, 6, 7, 8, 9], calculate the maximum Hamming distance after shuffling, and select the top eighty permutations of the maximum Hamming distance as the permutation set;
[0032] Set the permutation set for the upstream jigsaw task. For the 3x3 jigsaw permutation problem, there are a total of 9! = 362,880 possible permutations. The choice of the permutation set has a certain impact on the performance of the experiment. Some of these 362,880 permutations are very similar to the original permutation, and no effective feature representation can be learned from them. Therefore, a permutation method with a large difference from the original permutation is needed. In this embodiment, the maximum Hamming distance is adopted. The Hamming distance is used in data transmission error control coding and represents the number of different characters at the corresponding positions of two (same-length) strings. Let d(x,y) represent the Hamming distance between two words x and y. Perform an exclusive OR operation on the two strings and count the number of results that are 1 as the Hamming distance of the permutation.
[0033] Select the eighty permutations with the largest Hamming distance as the permutation set for this experiment according to the idea of the maximum Hamming distance, and calculate the average Hamming distance of the selected permutation set according to the exclusive OR operation, as shown in Table 1.
[0034] Table 1 Average Hamming Distance
[0035] Permutation number Average Hamming distance 80 8.1
[0036] Step 3: Cut the image patches obtained in Step 1 into 3x3 small patches. To better solve the jigsaw problem, a cutting method with gaps is adopted during cutting, and the 9 obtained image patches are normalized separately.
[0037] In this embodiment, the image is cropped according to the Figure 2 cropping method on the left to obtain 9 image patches. The pixel values exceeding the image are set to 255. This cropping method is beneficial to solving the upstream jigsaw task. Then, the image is preprocessed, and each image patch is normalized separately. Calculate the mean and standard deviation of each image patch to prevent the interference of low-level information.
[0038] Step 4: Randomly select a permutation from the permutation set obtained in Step 2 as the permutation order, and send the 9 image patches after data processing in Step 3 into the upstream task network for prediction according to the selected permutation order. The goal is to obtain the index of the selected permutation.
[0039] Figure 2 The structure shown is the structure diagram of the upstream jigsaw task network, which is an improvement on Alexnet. The number of convolutional kernels in each convolutional layer is changed to some extent, and fewer parameters are adopted. In Figure 2The maximum pooling layer and the activation function layer are not shown, and the weights are shared before the first fully connected layer (fc6); randomly select a permutation from the permutation set in step 2, and send the 9 preprocessed image patches in step 3 into the upstream task network in the order of the selected permutation. All image patches have the same operations before passing through the second fully connected layer. After passing through the first fully connected layer (fc6), the outputs of the 9 image patches are integrated together to form the input of fc7. After passing through two more fully connected layers and a softmax classifier, the final output value is obtained. The goal is to predict the index value of the selected permutation, calculate the loss and backpropagate to update the network parameters, and determine whether the termination condition of reaching a certain number of training epochs is satisfied;
[0040] Assign an index to the pre-defined permutation set. Finally, the upstream jigsaw task network returns a vector containing the probability values of each index. The output of the upstream jigsaw task network can be regarded as the conditional probability density function of the object space permutation;
[0041]
[0042] where S represents the selected permutation, and A i represents the i-th image patch, and F i represents the intermediate feature representation, p represents the conditional probability density function, and the goal is to make the feature F i have semantic attributes that can recognize the relative positions between image patches;
[0043] If only one jigsaw puzzle is generated for an image, it is possible to only learn the information of the absolute position, which is obviously not in line with the requirements. Therefore, multiple jigsaw puzzles need to be generated for an image. If a position list is generated for each part in the configuration S, where S=(L1, L2,..., L9), then p(S|F1, F2,..., F9) can be written as:
[0044]
[0045] L i represents the position of the i-th image patch, and F i represents the intermediate feature representation, p represents the conditional probability density function, so that the position of each image patch is determined by the corresponding feature;
[0046] In the upstream jigsaw task network, the cross-entropy loss function is used. Cross-entropy is a concept in information theory. Given two probability distributions d and e, the cross-entropy of representing d by e is:
[0047]
[0048] d represents the probability distribution of the true distribution, e represents the probability distribution of the non-true distribution, and H represents the cross-entropy;
[0049] The cross - entropy loss function used is defined as:
[0050]
[0051] where y i is the label value, is the index value of the selected permutation, and y' i is the predicted index value, and n represents the number of classes;
[0052] Step 5: Randomly select 20 images from each class in the training set samples obtained in Step 1 as the training samples for the downstream scene classification task and perform data processing; use the upstream jigsaw task model as the pre - trained model for the downstream scene classification task for transfer learning, and fine - tune the downstream scene classification task, Figure 3 as shown in the network structure diagram of the downstream scene classification task;
[0053] The training set of the downstream scene classification task is composed of randomly extracting 20 pictures from each class in the training set of the upstream jigsaw task, and the rest are used as the validation set; the part before the fully - connected layer of the downstream scene classification task network is the same as that of the upstream jigsaw task network. Use the weights of the upstream jigsaw task network to initialize the part before the last convolutional layer of the downstream scene classification task network and freeze it, and do not update the gradient during the training of the downstream scene classification task. Let the part of the last convolutional layer participate in the training together with the new fully - connected layer; the downstream scene classification task is regarded as a supervised classification task. Use a small number of labeled samples, fine - tune the downstream scene classification task using the pre - trained weights of the upstream jigsaw task, calculate the loss and update the network parameters, and judge whether the termination condition is met, that is, reaching a certain number of epochs;
[0054] Step 6: Cut the complete SAR image into small pieces of the same size as in Step 1 according to the covered cutting method, classify them using the downstream scene classification task model, assign specific colors to specific classes after classification, and splice them into a large image according to the cutting method;
[0055] In this embodiment, the entire SAR image is cut into small images according to the same size as in Step 1, and is cut in an overlapping manner. The overlapping ranges are different at different sizes, but ensure that the central area size is 20x20; send the cut image patches into the downstream scene classification task model to predict the class of each image patch. For each class, assign it a specific color, as shown in Table 2;
[0056] Table 2: Region category - RGB value correspondence
[0057] Category Water area Forest Building Farmland Unknown class RGB value [0,0,255] [0,255,0] [255,0,0] [255,255,0] [0,0,0]
[0058] After assigning specific colors to the cut image patches, the image patches are stitched in the same way as the large cut image. For each image patch, only the middle 20x20 area is taken for stitching to obtain the final result image;
[0059] Step 7, for the unknown classes in the SAR image, find the corresponding areas of the unknown classes on the label image, cover the corresponding areas on the result image with the colors of the unknown classes, and convert the result image into an indexed image;
[0060] In this embodiment, for the areas of the unknown classes on the SAR image label image, find the positions of the unknown class areas in the original SAR image, find the corresponding positions on the result image obtained in step 6 and assign the colors of the unknown classes, and convert the result image into the corresponding indexed image for calculating evaluation metrics with the label image;
[0061] Step 8, calculate the evaluation metrics between the obtained result image and the label image of the SAR image, calculate the evaluation metrics CPA, Recall, F1score for each class, as well as the overall evaluation metrics PA, Kappa, MIoU, FWIoU.
[0062] The relevant evaluation metrics and their definitions are as follows:
[0063] There are k + 1 classes (including one unknown class), p ij represents the number of pixels that originally belonged to class i but were predicted as class j. In this experiment, the unknown class does not participate in the calculation of the evaluation metrics;
[0064] Precision (CPA): The proportion of correctly predicted positives among all predicted positives, defined as follows:
[0065]
[0066] p ii represents the number of pixels that originally belonged to class i but were predicted as class i, p ij represents the number of pixels that originally belonged to class i but were predicted as class j, k represents the class, and C represents the precision;
[0067] Recall: The proportion of correctly predicted positives among all positive samples, defined as follows:
[0068]
[0069] p ii represents the number of pixels that originally belonged to class i but were predicted as class i, p ji represents the number of pixels that originally belonged to class j but were predicted as class i, k represents the class, and R represents the recall;
[0070] F1score: The harmonic mean based on recall and precision, defined as follows:
[0071]
[0072] C represents the precision rate in the above formula, R represents the recall rate in the above formula, and F represents the F1 score;
[0073] Overall accuracy (PA): The ratio of the correctly labeled pixels to the total pixels, which is defined as follows:
[0074]
[0075] p ii represents the number of pixels that originally belong to class i but are predicted as class i, and p ij represents the number of pixels that originally belong to class i but are predicted as class j, where k represents the class;
[0076] Kappa coefficient: Used for consistency testing, punishing the "bias" of the experimental model to obtain a more fair experimental model, which is defined as follows:
[0077]
[0078]
[0079] Among them, p o represents the proportion of the number of correctly classified samples in each class to the total samples, which is equivalent to PA; a1, a2,..., a c represent the number of true samples in each class; b1, b2,..., b c represent the number of predicted samples in each class; the number of classes is C, and the total number of samples is n; p e is calculated from the combination of true samples, predicted samples, and the total number of samples;
[0080] Mean Intersection over Union (MIoU): The sum average of the ratio of the intersection to the union of the predicted results and the true values for each class, which is defined as follows:
[0081]
[0082] p ii represents the number of pixels that originally belong to class i but are predicted as class i, and p ij represents the number of pixels that originally belong to class i but are predicted as class j, and p ji represents the number of pixels that originally belong to class j but are predicted as class i, where k represents the class;
[0083] Frequency Weighted Intersection over Union (FWIoU): An improvement of MIoU, setting weights according to the frequency of class occurrence, which is defined as follows:
[0084]
[0085] pii Indicates the number of pixels that originally belong to class i but are predicted as class i, p ij Indicates the number of pixels that originally belong to class i but are predicted as class j, p ji Indicates the number of pixels that originally belong to class j but are predicted as class i, where k represents the class;
[0086] Experimental content and results:
[0087] The experiment was operated under the hardware system of Intel Xeon(R) CPU and NVIDIA GeForce RTX 2080Ti. The software environment used was python3.8 and pytorch1.8.0, and the operating system used was Ubuntu 18.04.
[0088] The SAR images used in the experiment were SAR images of the Napoli area in Italy. The polarization mode used was HH polarization, the resolution was 2.5m, the image size was 16000*18332, and it contained five land cover classes, namely water area, forest, building, farmland, and unknown area, Figure 4 is the SAR image of the Napoli area, Figure 5 is the label map corresponding to the SAR image of the Napoli area.
[0089] During training, the accuracy of the upstream jigsaw task, the accuracy of the downstream scene classification task, and the evaluation metrics for large image segmentation were tested respectively under the conditions where the image patch sizes were 50x50, 75x75, 100x100, 120x120, 150x150, and 200x200. Taking the evaluation metrics for large image segmentation as the most important indicator, a comparative experiment of supervised CNN without using pre-trained weights was conducted for the best set of settings. The training parameters set for the upstream task were: the learning rate was 0.01, the batch size of the training data was 64, the batch size of the validation data was 32, the number of training epochs was 300, the learning rate decay criterion adopted was that the learning rate was reduced to 0.5 times the original every 50 generations, and the optimizer used was the SGD optimizer; the training parameters for fine-tuning the downstream scene classification task were: the learning rate was 0.001, the batch size of the training data was 2, the batch size of the validation data was 16, the number of training epochs was 100, the learning rate decay criterion adopted was that the learning rate was reduced to 0.5 times the original every 50 generations, and the optimizer used was the SGD optimizer; the accuracy of the upstream task refers to randomly selecting a permutation from a predefined set of permutations, obtaining the index value of the selected permutation, shuffling the image according to this permutation method, inputting the shuffled image patches into the network, finally obtaining a probability vector, taking the index value of the largest value in the probability vector as the final result, and judging whether it is the same as the selected permutation index value, and calculating the accuracy by comparing with the true permutation index value; the accuracy of the downstream scene classification task refers to the accuracy of scene classification; the evaluation metrics for large image segmentation refer to the comparison result between the result image and the SAR image label map.
[0090] Table 3 Accuracy of Upstream Tasks with Different Sizes
[0091] 50x50 75x75 100x100 120x120 150x150 200x200 255x255 0.728 0.856 0.656 0.894 0.805 0.74 0.89
[0092] When the image sizes are 50x50, 100x100, and 200x200 and the image is divided into 3x3 regions, the image cannot be evenly divided, and a certain gap is set during the division, resulting in a relatively low accuracy of the upstream task; when the image sizes are 75x75, 120x120, 150x150, and 255x255 and the image is divided into 3x3 regions, the image can be evenly divided, achieving a relatively high accuracy in the upstream task.
[0093] Table 4 Accuracy of Downstream Scene Classification Tasks with Different Sizes
[0094] 50x50 75x75 100x100 120x120 150x150 200x200 255x255 0.765 0.779 0.845 0.847 0.867 0.877 0.905
[0095] The accuracy of the downstream scene classification task gradually increases as the size of the small image patches increases; next, a simple analysis of the evaluation metrics for large images was conducted, and more detailed analysis and comparative experiments were carried out after finding the set of experiments with the best metrics.
[0096] Table 5 PA and MIoU of Different Sizes
[0097] 50x50 75x75 100x100 120x120 150x150 200x200 255x255 PA 0.788 0.786 0.784 0.80 0.787 0.789 0.767 MIoU 0.618 0.620 0.63 0.649 0.63 0.625 0.596
[0098] It is found that when observing PA and MioU under different sizes, PA and MIoU reach the maximum values when the image size is 120x120; therefore, a comparative experiment without using the pre-training weights of the jigsaw task is conducted under the image size of 120x120 and the experimental indicators are analyzed.
[0099] Table 6 Accuracy of the Downstream Scene Classification Task of the Present Invention and the Comparative Method under the Image Size of 120x120
[0100] The present invention Comparison method Image size 120x120 0.847 0.759
[0101] Table 7 Evaluation Indicators of Large Image Segmentation of the Present Invention and the Comparative Method under the Image Size of 120x120
[0102]
[0103]
[0104] *In Table 7, for each type of evaluation indicator CPA, recall, and F1score, from left to right, they represent water area, forest, building, and farmland in turn.
[0105] Analysis of Experimental Results: As can be seen from Table 6, the accuracy of the present invention in the downstream scene classification task under the image size of 120x120 is about 9% higher than that of the comparative experiment; when migrated to large images, it is found that using the method of the present invention has a certain improvement in each evaluation indicator compared with the comparative experiment. The pixel accuracy (PA) has an accuracy improvement of 6%, the mean pixel accuracy per class (MPA) has an accuracy improvement of 8%, the mean intersection over union (MIoU) has an accuracy improvement of about 8%, the frequency weighted intersection over union (FWIoU) has an accuracy improvement of about 8%, and the kappa coefficient for consistency test also has an accuracy improvement of about 8%. Figure 6 is the final result diagram obtained by the method of the present invention, Figure 7 is the final result diagram obtained by the comparative method. From Figure 6 and Figure 7 comparison, it can be seen that Figure 6 the result diagram of the present invention is more complete than that of the comparative experiment, and the classification effect between buildings and forests is better, Figure 7 the misclassification shown in the figure is relatively serious and there are many stray points. The experimental results of the present invention make it easier to distinguish between buildings and forests, with better regional consistency and clearer edge information. Through Figure 6 and Figure 7 the effect can be seen more intuitively.
[0106] Based on the analysis of the above simulation results, the method for classifying ground objects in SAR images proposed by the present invention, which is based on self-supervised jigsaw learning, uses self-supervised learning to mine the representation characteristics of the data itself as supervision information by designing the idea of an upstream auxiliary task, realizes the effective representation of ground object features, alleviates the defects of insufficient discriminability of ground object features caused by insufficient labeled samples and diverse forms of expression of ground object targets, and improves the classification performance of the scene classification model.
[0107] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof; when implemented in whole or in part in the form of a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0108] As described above, only the embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for classifying ground objects in SAR images based on self-supervised jigsaw learning, characterized in that, It includes the following steps: Step 1: Input the SAR image, and randomly intercept image patches of a fixed size from each type of region to form a training set and a validation set; Step 2: Randomly shuffle the list [1, 2, 3, 4, 5, 6, 7, 8, 9], calculate the maximum Hamming distance after shuffling, and select the top eighty permutations with the maximum Hamming distance as the permutation set; Step 3: Cut the image patches obtained in Step 1 into 3x3 small patches. When cutting, adopt a cutting method with gaps, and perform normalization processing on the obtained 9 image patches separately; Step 4: Randomly select a permutation from the permutation set obtained in Step 2 as the permutation order, and send the 9 image patches after data processing in Step 3 into the upstream task network for prediction. The goal is to obtain the index of the selected permutation; Step 5: Randomly select 20 images from each type of the training set samples obtained in Step 1 as the training samples for the downstream scene classification task and perform data processing; Use the upstream jigsaw task model as the pre-trained model for the downstream scene classification task for transfer learning, and fine-tune the downstream scene classification task; Step 6: Cut the complete SAR image into small patches of the same size as in Step 1 according to the covering cutting method, use the downstream scene classification task model for classification, assign specific colors to specific classes after classification, and splice them into a large image according to the cutting method; Step 7: For the unknown classes in the SAR image, find the corresponding regions of the unknown classes on the label map, cover the corresponding regions on the result map with the colors of the unknown classes, and convert the result map into an index map; Step 8: Calculate the evaluation metrics between the obtained result map and the label map of the SAR image, calculate the evaluation metrics CPA, Recall, F1score for each class, and the overall evaluation metrics PA, Kappa, MIoU, FWIoU.
2. The method according to claim 1, wherein The specific content of Step 1 is as follows: Input the SAR image, and take 500 image patches from each of the four types of regions: water area, forest, building, and farmland, with sizes of 50x50, 75x75, 100x100, 120x120, 150x150, and 200x200 respectively, to explore the influence of the image size on the ground object classification of the SAR image. For each group of experiments, divide the training set and the validation set according to a ratio of 8:
2.
3. The method according to claim 1, wherein The specific content of Step 2 is as follows: Set a permutation set for the upstream jigsaw task. For the 3x3 jigsaw permutation problem, there are a total of 9! = 362880 possible permutations. Permutation methods with a large difference from the original permutation are required. The maximum Hamming distance is adopted. The Hamming distance is used in data transmission error control coding, which represents the number of different characters at the corresponding positions of two strings of the same length. Let d(x,y) represent the Hamming distance between two words x and y. Perform an exclusive OR operation on the two strings and count the number of results as 1 as the Hamming distance of the permutation.
4. The method according to claim 1, wherein The specific content of Step 3 is as follows: Crop the image to obtain 9 image patches, and set the pixel values outside the image to 255. Preprocess the image and perform individual normalization on each image patch, and calculate the mean and standard deviation of each image patch.
5. The method according to claim 1, characterized in that, The specific steps of step 4 are as follows: The upstream jigsaw task network is an improvement of Alexnet. The number of convolutional kernels in each convolutional layer is changed, and fewer parameters are used. The weights are shared before the first fully connected layer fc6. Randomly select a permutation from the permutation set in step 2, and sequentially feed the 9 preprocessed image patches in step 3 into the upstream task network according to the order of the selected permutation. All image patches have the same operations before passing through the second fully connected layer. After passing through the first fully connected layer fc6, the outputs of the 9 image patches are integrated together to form the input of fc7, and then through two fully connected layers and a softmax classifier to obtain the final output value. The goal is to predict the index value of the selected permutation, calculate the loss and backpropagate to update the network parameters, and determine whether the termination condition of reaching a certain number of training epochs is met. Assign an index to the predefined permutation set. Finally, the upstream jigsaw task network returns a vector containing the probability values of each index. The output of the upstream jigsaw task network is regarded as the conditional probability density function of the object space permutation. where S represents the selected permutation, A i represents the i-th image patch, F i forms an intermediate feature representation, p represents the conditional probability density function, and the goal is to make the feature F i have semantic attributes that can recognize the relative positions between image patches; Generate a position list for each part in configuration S, where S=(L1, L2, …, L9), then p(S|F1, F2, …, F9) is written as: L i represents the position of the i-th image patch, and F i represents the intermediate feature representation, and p represents the probability density function; in this way, the position of each image patch is determined by the corresponding feature; Use the cross-entropy loss function in the upstream jigsaw task network. Cross-entropy is a concept in information theory. Given two probability distributions d and e, the cross-entropy of d represented by e is: d represents the probability distribution of the true distribution, e represents the probability distribution of the non-true distribution, and H represents the cross-entropy. The cross-entropy loss function used is defined as: where y i is the tag value and the index value of the selected permutation, and y' i is the predicted index value, and n represents the category.
6. The method according to claim 1, wherein The specific steps of step 5 are as follows: The training set of the downstream scene classification task is randomly selected from 20 pictures of each category in the training set of the upstream jigsaw task, and the rest are used as the validation set. The part before the fully connected layer of the downstream scene classification task network is the same as that of the upstream jigsaw task network. Use the weights of the upstream jigsaw task network to initialize and freeze the part before the last convolutional layer of the downstream scene classification task network. During the training of the downstream scene classification task, the gradient is not updated, and the part of the last convolutional layer participates in the training together with the new fully connected layer. The downstream scene classification task is regarded as a supervised classification task. Use a small number of labeled samples, fine-tune the downstream scene classification task using the pre-trained weights of the upstream task, calculate the loss and update the network parameters, and determine whether the termination condition of reaching a certain number of training epochs is met.
7. The method according to claim 1, wherein The specific steps of step 6 are as follows: Cut the entire SAR image into small images according to the same size in step 1, and cut it in an overlapping manner. The overlapping ranges are different under different sizes, but ensure that the size of the central area is 20x20; send the cut image patches into the downstream scene classification task model to predict the category of each image patch, and assign a specific color to each category; after assigning specific colors to the cut image patches, splice the image patches in the same way as the large cut image, and only take the middle 20x20 area of each image patch for splicing to obtain the final result image.
8. The method according to claim 1, characterized in that, The specific content of step 7 is as follows: For the areas of unknown classes on the SAR image label map, find the positions of the unknown class areas in the original SAR image, find the corresponding positions on the result image obtained in step 6 and assign the colors of the unknown classes, and convert the result image into a corresponding index map for calculating evaluation metrics with the label map.
9. A computer, characterized in that: It is used to implement a method for classifying ground objects in SAR images based on self-supervised jigsaw learning according to any one of claims 1-8.
Citation Information
Patent Citations
Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field
AU2020103901A4
Scene graph generation method based on self-supervised pre-training
CN112989927A