Remote sensing image fine land utilization classification method based on DeepLabV3 + and ViT
By combining the remote sensing image classification method of DeepLabV3+ and ViT, the problem of insufficient land use classification accuracy and generalization ability in the existing technology is solved, and high-precision and good migration land use classification is achieved, which is suitable for cross-regional applications.
Patent Information
- Application Number
- CN202510260007.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-16
Smart Images

Figure CN120014365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image fine land use classification method based on DeepLabV3+ and ViT. Background Art
[0002] my country has clearly stated that it is necessary to build a new pattern of land space development and protection for high-quality development, and accurate acquisition of land use is the key foundation for achieving this goal. Land use information plays an important role in urban and rural planning, ecological environment monitoring, crop disaster assessment, disaster risk assessment and other fields, and is the primary task for the rational planning of land resources. Remote sensing satellite images provide strong data support for land use research with their advantages such as wide coverage, short update cycle and low data acquisition cost. However, the traditional method of manually annotating remote sensing images has problems such as low efficiency, high cost and accuracy affected by subjective factors, which makes it difficult to meet the needs of large-scale and high-precision land use classification.
[0003] With the development of computer technology, land use classification methods have evolved from traditional manual visual interpretation to supervised and unsupervised classification based on statistical patterns, and then to the current intelligent classification technology dominated by machine learning and deep learning. Early land use classification mainly relied on high-resolution remote sensing data and manual visual interpretation. Although it reduced costs to a certain extent, its accuracy and timeliness could not meet the needs of modern land and resources management. Subsequently, classification methods based on statistical patterns gradually emerged. Among them, supervised classification methods (such as the minimum distance method and the maximum likelihood method) need to rely on prior knowledge and training samples, while unsupervised classification methods (such as K-means clustering and ISODATA algorithm) reduce manual intervention through automatic clustering analysis. However, these methods have limited performance when dealing with complex landform types and nonlinear data, and are difficult to adapt to the needs of high-precision classification.
[0004] In recent years, the introduction of machine learning technology has brought significant improvements to land use classification. Support vector machines (SVM) are widely used due to their excellent performance in processing small samples and high-dimensional data; decision trees and random forest algorithms are favored for their ability to efficiently process large-scale data. In particular, the random forest algorithm significantly improves the robustness and accuracy of classification through an integrated learning strategy. At the same time, the rise of artificial neural networks, especially convolutional neural networks (CNN), has brought revolutionary breakthroughs in land use classification. Deep learning models represented by U-Net and DeepLabV3+ have achieved pixel-level accurate classification in semantic segmentation tasks, significantly improving classification accuracy in complex scenarios.
[0005] In addition, as an emerging technology, the Transformer network has gradually been introduced into the field of computer vision after its great success in the field of natural language processing. Its architecture based on the self-attention mechanism can effectively capture long-distance dependencies in remote sensing images, especially when dealing with large-scale, multi-scale object classification tasks. For example, models such as the Visual Transformer (ViT) and Swin Transformer have demonstrated strong performance in remote sensing image classification, providing a new technical path for land use classification.
[0006] Despite the significant progress made in current technology, land use classification still faces some challenges. First, the remote sensing image classification model based on semantic segmentation has poor transferability on data from different regions and is difficult to be directly applied to cross-regional scenarios. Second, the demand for multi-category refined classification is growing, but the existing models still have the problem of insufficient accuracy when dealing with complex land object types. Therefore, there is an urgent need for a new method that can combine the local feature extraction capabilities of convolutional neural networks and the global context modeling capabilities of Transformer to improve the accuracy and generalization of land use classification.
[0007] The Chinese invention patent with application number 202210617050.9 discloses "A Land Use Classification Method and System", which includes: the trained land use classification model classifies the land use type of each pixel in the input target land image to obtain a first land classification image; the land use classification model includes an encoder, a dual-path attention module, a spatial pyramid pooling module and a decoder; the dual-path attention module includes a first channel attention module and a first spatial position attention module; the channel attention weighted features and spatial attention weighted features are obtained through the dual-path attention module; the spatial pyramid pooling module then fuses the two to obtain a fused feature; the conditional random field classifies the land use type of each pixel in the first land classification image to obtain a second land classification image. Summary of the invention
[0008] In order to solve the technical problems of insufficient accuracy and generalization ability of land use classification in the prior art, the present invention provides a remote sensing image fine land use classification method based on DeepLabV3+ and ViT. The technical solution adopted by the present invention is: The first aspect of the present invention provides a remote sensing image fine land use classification method based on DeepLabV3+ and ViT, the method comprising: Construct a land use classification model based on DeepLabV3+ and ViT; Preprocess the public high-resolution remote sensing image land cover classification dataset to obtain the model dataset; The land use classification model based on DeepLabV3+ and ViT is trained and performance evaluated by using the model data set to obtain a final model; Obtain the Gaofen-1 remote sensing image as the target data to be classified, and preprocess the target data to obtain the target data set; The target data set is input into the final model for pixel-level classification prediction to obtain high-precision classification results.
[0009] As a preferred solution, the method of constructing a land use classification model based on DeepLabV3+ and ViT includes: Build the DeepLabV3+ model framework: The DeepLabV3+ model framework adopts an encoder-decoder structure. The encoder consists of the backbone network ResNet-101 and the dilated spatial convolution pooling pyramid pooling module ASPP. The main structure of ResNet-101 is: first a 7×7 convolution with a step size of 2, followed by a 3×3 maximum pooling with a step size of 2, and then a Bottleneck structure, where the Bottleneck structure is a residual structure composed of a 1×1 convolution, a 3×3 convolution and a 1×1 convolution; the main structure of ASPP is: the feature map output by the backbone network is input into a 1×1 convolution, a dilated convolution with expansion rates of 6, 12, and 18, and a global average pooling module, respectively, to obtain five feature maps, which are spliced and then passed through a 1×1 convolution; the decoder uses a 1×1 convolution to reduce the dimension of the low-level feature map output in the middle of the backbone network, and then upsamples the feature map output by ASPP to obtain two feature maps, which are spliced together and then passed through a 3×3 convolution; Build the Vision Transformer module: The Vision Transformer module includes an Embedding layer, a Transformer Encoder layer, and an MLP Head layer. The Embedding layer is used to process the input image into a token sequence that the Transformer can receive. The input image dimension is set to H×W×C, and the size of each block is P×P. First, the input image is divided into a series of two-dimensional sequences according to the size of the block to obtain N blocks, that is, ,in , and N is the effective length of the input sequence of the Transformer module. Then, each block is mapped to a one-dimensional vector through linear transformation. Then, each one-dimensional vector is concatenated with the corresponding position code to form an image token, and a category code is added at the front end. This category token is used for the final category output, and finally the input to the Transformer Encoder layer is obtained. The expression is: (1) in, Parameters added for encoding of blocks; Parameters added for positional encoding, Code the categories; The Transformer Encoder layer is composed of L stacked Transformer Encoder modules. The structure of each Transformer Encoder module is mainly composed of a multi-head self-attention mechanism and a multi-layer perception mechanism. A normalization module is set before the multi-head self-attention mechanism and the multi-layer perception mechanism, and the residual structure is used to fuse the input features and the output features. The operation expression is: (2) (3) Among them, the MLP module consists of full connection, GELU activation function and Dropout.
[0010] The MLP Head layer is used to classify the image and finally output the classification result. The MLP Head layer is composed of a Layer Norm module and a linear mapping module, and the expression is: (4) In the DeepLabV3+ model framework, the Vision Transformer module is introduced to obtain the land use classification model based on DeepLabV3+ and ViT: The land use classification model based on DeepLabV3+ and ViT adopts an encoder-decoder structure, and the encoder is composed of a backbone network ResNet-101, ASPP and Vision Transformer; the remote sensing image is input into the backbone network ResNet-101, and a high-level feature map is obtained after a deep convolution operation, which is respectively input into the ASPP module and the Vision Transformer module. The ASPP module inputs the obtained high-level feature map into a 1×1 convolution, a dilated convolution with expansion rates of 6, 12, and 18, and a global average pooling module to obtain 5 feature images, which are spliced to obtain a high-level feature map that integrates multi-scale information; the Vision Transformer module processes the obtained high-level feature map through a Patch Embedding layer and then inputs it into the Transformer layer, which is composed of 12 Transformer modules, and the obtained global high-level feature image is processed by Patch Embedding. After processing by the Expanding layer, the image is restored to the size before the input module. The decoder inputs the low-level feature map output by the backbone network into a 1×1 convolution to obtain a high-level feature map, inputs the high-level feature map obtained by the ASPP module that integrates multi-scale information into a 1×1 convolution and performs a 4-fold upsampling to obtain a multi-scale low-level feature map, and inputs the high-level feature map obtained by the Vision Transformer module that extracts global features into a 1×1 convolution and performs a 4-fold upsampling to obtain a global low-level feature map. The three are spliced together. Finally, the spliced feature map is input into a 3×3 convolution and performed a 4-fold upsampling to obtain a predicted image with the same size as the original image.
[0011] As a preferred solution, the method of preprocessing the public high-resolution remote sensing image land cover classification dataset to obtain the model dataset includes: The images of the public high-resolution remote sensing image land cover classification dataset were cropped into 256*256 images so that they can be input into the model for training; By data enhancement, the images of land use types with fewer samples in the public high-resolution remote sensing image land cover classification dataset are expanded to obtain an enhanced dataset; The enhanced data set is divided into a source domain data set and a target domain data set, and then the target domain data set is divided into a target domain training set, a target domain verification set and a target domain test set according to a preset ratio.
[0012] As a preferred solution, the method of training and evaluating the performance of the land use classification model based on DeepLabV3+ and ViT by using the model data set to obtain the final model includes: Inputting the source domain dataset into the land use classification model based on DeepLabV3+ and ViT for training to obtain a pre-trained model; Inputting the target domain training set into the pre-trained model for fine-tuning to obtain a fine-tuned model; Inputting the target domain validation set into the fine-tuning model to optimize the loss function and adjust the hyperparameters to obtain an improved model; The target domain test set is input into the improved model and the model performance is evaluated according to the preset evaluation index, and the model with the best performance is retained as the final model.
[0013] As a preferred solution, the preset evaluation indicators include: IoU, average IoU, overall pixel accuracy, Kappa coefficient, and execution time.
[0014] As a preferred solution, a method for obtaining a GF-1 remote sensing image as target data to be classified and preprocessing the target data to obtain a target data set includes: The image is cropped using overlapping sliding windows and dilation prediction to obtain the target data set, specifically: The right and lower boundaries of the original remote sensing image are filled with 0 to obtain a first preprocessed image, wherein the size of the filled image must be an integer multiple of the sliding prediction window; To perform expansion prediction on the edge data, an outer border is added to the filled image to obtain a second preprocessed image, wherein the outer border is filled with 0, and the filling size is 1 / 2 of the sliding window step size; The second preprocessed image is cropped using a sliding window of size 256×256 and a sliding step of 128, and the coordinates of the cropped image are recorded to obtain a target data set.
[0015] As a preferred solution, the method of inputting the target data set into the final model for pixel-level classification prediction to obtain a high-precision classification result includes: The target data set is input into the final model to perform pixel-level classification prediction and output high-precision classification results. The quality of each scene data is checked manually visually, and the classification results are spliced and finally displayed visually.
[0016] As a preferred solution, the splicing processing method includes: Keep the central area of the classification result image with a size of 128×128 pixels; According to the coordinates of the cropped image, the result image is spliced to the corresponding position to obtain the final prediction result image.
[0017] A second aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned method for fine land use classification of remote sensing images based on DeepLabV3+ and ViT.
[0018] The third aspect of the present invention provides a computer device, including a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, the steps of the aforementioned remote sensing image fine land use classification method based on DeepLabV3+ and ViT are implemented.
[0019] Compared with the prior art, the present invention has the following beneficial effects: The invention discloses a remote sensing image fine land use classification method based on the fusion of DeepLabV3+ and Vision Transformer (ViT). It innovatively combines the local feature extraction capability of the deep convolutional neural network (DCNN) with the global feature modeling advantage of the visual transformer (ViT). By embedding the ViT module in the convolutional neural network, the global context information of the remote sensing image is effectively captured, which solves the limitations of the traditional DCNN in modeling the relationship between pixels. Through the extraction and fusion of multi-scale features, the model's ability to recognize complex land categories is significantly improved, and pixel-level fine land use classification is achieved. By introducing transfer learning and domain adaptation technology, the method of the invention shows good mobility and adaptability on data from different regions, solving the problem of poor mobility of existing models in cross-regional applications. While ensuring high precision, the method of the invention has high computational efficiency, can meet the needs of large-scale remote sensing image processing, and has a wide range of practical application value.
[0020] First, the present invention innovatively combines the local feature extraction capability of deep convolutional neural network (DCNN) with the global feature modeling capability of visual transformer (ViT). By embedding the ViT module in the convolutional neural network, the global context information of remote sensing images is effectively captured, solving the limitations of traditional DCNN in modeling the relationship between pixels. At the same time, the ViT module is used to globally model high-level features, which significantly improves the model's ability to identify complex land categories.
[0021] Second, the present invention can simultaneously capture local details and global structural information in remote sensing images through a multi-scale feature extraction and fusion mechanism, significantly improving the model's ability to classify multi-scale objects.
[0022] Third, the method of the present invention can achieve fine pixel-level land use classification and performs well in complex multi-category remote sensing image scenes. For example, it can identify water body information such as ponds, lakes, and rivers, and achieve high-precision classification in complex scenes such as high-density urban areas and mixed vegetation areas.
[0023] Fourth, by introducing transfer learning and fine-tuning technology, the method of the present invention shows good mobility and adaptability on data from different regions, solving the problem of poor mobility of existing models in cross-regional applications.
[0024] Fifth, the method of the present invention has high computational efficiency while ensuring high precision, can meet the needs of large-scale remote sensing image processing, and has broad practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flow chart of a remote sensing image fine land use classification method based on DeepLabV3+ and ViT is provided for this implementation; Figure 2 The structure diagram of the land use classification model based on DeepLabV3+ and ViT provided for this implementation; Figure 3 The overall structure diagram of the Vision Transformer module provided in this embodiment; Figure 4 The Transformer Encoder module structure diagram provided for this embodiment; Figure 5 The MLP module structure diagram provided for this embodiment; Figure 6 A fine-tuning process diagram provided for this embodiment; Figure 7 This is a diagram of the overlapping sliding windows and expansion prediction process provided in this embodiment. DETAILED DESCRIPTION The accompanying drawings are only used for illustrative purposes and are not to be construed as limiting the present invention; It should be clear that the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.
[0026] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms of "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0027] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.
[0028] In addition, in the description of this application, unless otherwise specified, "multiple" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0029] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0030] Example 1 Please refer to Figure 1 This embodiment provides a remote sensing image fine land use classification method based on DeepLabV3+ and ViT, the method comprising: S1: Construct a land use classification model based on DeepLabV3+ and ViT; It should be noted that the land use classification model based on DeepLabV3+ and ViT is an improvement on the traditional semantic segmentation model DeepLabV3+ model. On the basis of the DeepLabV3+ model, the VisionTransformer module is introduced to capture the long-distance dependencies in remote sensing images through the self-attention mechanism, thereby enhancing the model's ability to understand global contextual information.
[0031] In a specific embodiment, the method for constructing a land use classification model based on DeepLabV3+ and ViT includes: Build the DeepLabV3+ model framework: Please refer to Figure 2 , The DeepLabV3+ model framework adopts an encoder-decoder structure. The encoder consists of the backbone network ResNet-101 and the dilated spatial convolution pooling pyramid pooling module ASPP. The main structure of ResNet-101 is: first a 7×7 convolution with a step size of 2, followed by a 3×3 maximum pooling with a step size of 2, and then a Bottleneck structure, where the Bottleneck structure is a residual structure composed of a 1×1 convolution, a 3×3 convolution and a 1×1 convolution; the main structure of ASPP is: the feature map output by the backbone network is input into a 1×1 convolution, a dilated convolution with expansion rates of 6, 12, and 18, and a global average pooling module, respectively, to obtain five feature maps, which are spliced and then passed through a 1×1 convolution; the decoder uses a 1×1 convolution to reduce the dimension of the low-level feature map output in the middle of the backbone network, and then upsamples the feature map output by ASPP to obtain two feature maps, which are spliced together and then passed through a 3×3 convolution; Build the Vision Transformer module: Please refer to Figure 3 , the Vision Transformer module includes an Embedding layer, a TransformerEncoder layer, and an MLP Head layer; the Embedding layer is used to process the input image into a token sequence that the Transformer can receive; the input image dimension is set to H×W×C, and the size of each block is P×P. First, the input image is divided into a series of two-dimensional sequences according to the size of the block to obtain N blocks, that is, ,in , and N is the effective length of the input sequence of the Transformer module. Then, each block is mapped to a one-dimensional vector through linear transformation. Then, each one-dimensional vector is concatenated with the corresponding position code to form an image token, and a category code is added at the front end. This category token is used for the final category output, and finally the input to the TransformerEncoder layer is obtained. The expression is: (1) in, Parameters added for encoding of blocks; Parameters added for positional encoding, Code the categories; Please refer to Figure 4 , the Transformer Encoder layer is composed of L Transformer Encoder modules stacked together. The structure of each Transformer Encoder module is mainly composed of a multi-head self-attention mechanism and a multi-layer perception mechanism. A normalization module is set before the multi-head self-attention mechanism and the multi-layer perception mechanism, and the residual structure is used to fuse the input features and the output features. The operation expression is: (2) (3) The MLP module consists of full connection, GELU activation function and Dropout; please refer to Figure 5 , the MLP module is similar to the FFN layer in the classic Transformer network.
[0032] The MLP Head layer is used to classify the image and finally output the classification result. The MLP Head layer is composed of a Layer Norm module and a linear mapping module, and the expression is: (4) In the DeepLabV3+ model framework, the Vision Transformer module is introduced to obtain the land use classification model based on DeepLabV3+ and ViT: The land use classification model based on DeepLabV3+ and ViT adopts an encoder-decoder structure, and the encoder is composed of a backbone network ResNet-101, ASPP and Vision Transformer; the remote sensing image is input into the backbone network ResNet-101, and a high-level feature map is obtained after a deep convolution operation, which is respectively input into the ASPP module and the Vision Transformer module. The ASPP module inputs the obtained high-level feature map into a 1×1 convolution, a dilated convolution with expansion rates of 6, 12, and 18, and a global average pooling module to obtain 5 feature images, which are spliced to obtain a high-level feature map that integrates multi-scale information; the Vision Transformer module processes the obtained high-level feature map through a Patch Embedding layer and then inputs it into the Transformer layer, which is composed of 12 Transformer modules, and the obtained global high-level feature image is processed by Patch Embedding. After processing by the Expanding layer, the image is restored to the size before the input module. The decoder inputs the low-level feature map output by the backbone network into a 1×1 convolution to obtain a high-level feature map, inputs the high-level feature map obtained by the ASPP module that integrates multi-scale information into a 1×1 convolution and performs a 4-fold upsampling to obtain a multi-scale low-level feature map, and inputs the high-level feature map obtained by the Vision Transformer module that extracts global features into a 1×1 convolution and performs a 4-fold upsampling to obtain a global low-level feature map. The three are spliced together. Finally, the spliced feature map is input into a 3×3 convolution and performed a 4-fold upsampling to obtain a predicted image with the same size as the original image.
[0033] S2: Preprocess the public high-resolution remote sensing image land cover classification dataset to obtain the model dataset; It should be noted that the GID dataset, a publicly available high-resolution remote sensing image land cover classification dataset, does not meet the training format of the model in S1 and has the problem of unbalanced sample classification. It is necessary to crop and preprocess the dataset for data enhancement. The original remote sensing image is cropped into an image of 256*256 size so that it can be input into the model for training. Then, the images of land use types with fewer samples are expanded by using translation transformation, size transformation, rotation transformation, mirror transformation, hue transformation, saturation transformation, and brightness transformation. The GID dataset has a total of 15 land use types, namely background, industrial land, urban residential, rural residential, transportation land, rice fields, irrigated land, dry land, garden, arbor forest, shrub forest, natural grassland, artificial grassland, river, lake and pond, which are marked with 16 different colors accordingly. Finally, the preprocessed data is divided into a source domain dataset and a target domain dataset, and then the target domain dataset is randomly divided into a training set, a validation set and a test set according to a ratio of 6:2:2.
[0034] In a specific embodiment, a method for preprocessing a public high-resolution remote sensing image land cover classification dataset to obtain a model dataset includes: The images of the public high-resolution remote sensing image land cover classification dataset were cropped into 256*256 images so that they can be input into the model for training; By data enhancement, the images of land use types with fewer samples in the public high-resolution remote sensing image land cover classification dataset are expanded to obtain an enhanced dataset; The enhanced data set is divided into a source domain data set and a target domain data set, and then the target domain data set is divided into a target domain training set, a target domain verification set and a target domain test set according to a preset ratio.
[0035] S3: training and evaluating the performance of the land use classification model based on DeepLabV3+ and ViT using the model data set to obtain a final model; In a specific embodiment, the method of training and evaluating the performance of the land use classification model based on DeepLabV3+ and ViT by using the model data set to obtain the final model includes: Inputting the source domain dataset into the land use classification model based on DeepLabV3+ and ViT for training to obtain a pre-trained model; Specifically, the preprocessed source domain dataset was input into the land use classification model based on DeepLabV3+ and ViT for pre-training, with Batch_size of 16, the SGD algorithm selected as the network parameter update algorithm, the momentum gradient of 0.9, and the regularization term coefficient of 0.0001. The learning rate decay strategy selected the cosine annealing algorithm, the initial learning rate was 0.01, and the minimum learning rate was 0.00001. The loss function selected a combination of cross entropy and Dice loss, with weights of 0.5 and 0.5, respectively. After 100 training iterations, the land use classification pre-training model was obtained.
[0036] Please refer to Figure 6 , input the target domain training set into the pre-trained model for fine-tuning to obtain a fine-tuned model; input the target domain validation set into the fine-tuned model for loss function optimization and hyperparameter adjustment to obtain an improved model; Specifically, the preprocessed target domain training set is input into the pretrained model for fine-tuning. By using transfer learning and fine-tuning techniques, the pretrained model in the source domain is migrated to the target domain, so that the model can adapt to different data distributions. During the fine-tuning process, the data from the target domain is used to further adjust the parameters of the model output layer so that it performs better on the target domain tasks. The fine-tuned model can utilize the knowledge of the source domain while adapting to the characteristics of the target domain, thereby improving the generalization ability and performance of the model. The target domain validation set is used to further improve the classification accuracy of the model through loss function optimization and hyperparameter adjustment.
[0037] The target domain test set is input into the improved model and the model performance is evaluated according to the preset evaluation index, and the model with the best performance is retained as the final model.
[0038] In a specific embodiment, the preset evaluation indicators include: IoU, average IoU, overall pixel accuracy, Kappa coefficient, and execution time.
[0039] Specifically, the prediction results of the test set are evaluated using the model's intersection over union (IoU), mean intersection over union (mIoU), overall pixel accuracy (OA), Kappa coefficient, and execution time (Time) in each category on the test set. See Table 1, Table 2, and Table 3.
[0040] Table 1 Intersection-over-union of land use classification models based on DeepLabV3+ and ViT on the test set
[0041] Table 2 Intersection-over-union of land use classification models based on DeepLabV3+ and ViT on the test set
[0042] Table 3 Overall pixel accuracy, Kappa coefficient and execution time of the land use classification model based on DeepLabV3+ and ViT on the test set
[0043] According to Table 1, Table 2 and Table 3, the method proposed in the present invention has better performance than the original DeepLabV3+ model.
[0044] S4: Obtain the GF-1 remote sensing image as the target data to be classified, and preprocess the target data to obtain a target data set; It should be noted that the GF-1 remote sensing image is used as the target data to be classified. Since the remote sensing image is 13754×14906 pixels, it does not meet the training format of the model, so the image needs to be cropped. The overlapping sliding window and expansion prediction method are used to crop the image to a size of 256*256 to form the target data set.
[0045] In a specific embodiment, a method of obtaining a GF-1 remote sensing image as target data to be classified and preprocessing the target data to obtain a target data set includes: Please refer to Figure 7 , the image is cropped by overlapping sliding windows and dilation prediction to obtain the target data set, which is as follows: The right and lower boundaries of the original remote sensing image are filled with 0 to obtain a first preprocessed image, wherein the size of the filled image must be an integer multiple of the sliding prediction window; To perform expansion prediction on the edge data, an outer border is added to the filled image to obtain a second preprocessed image, wherein the outer border is filled with 0, and the filling size is 1 / 2 of the sliding window step size; The second preprocessed image is cropped using a sliding window of size 256×256 and a sliding step of 128, and the coordinates of the cropped image are recorded to obtain a target data set.
[0046] S5: inputting the target data set into the final model for pixel-level classification prediction to obtain a high-precision classification result; In a specific embodiment, the target data set is input into the final model for pixel-level classification prediction to obtain a high-precision classification result, which includes: The target data set is input into the final model to perform pixel-level classification prediction and output high-precision classification results. The quality of each scene data is checked manually visually, and the classification results are spliced and finally displayed visually.
[0047] In a specific embodiment, the splicing method includes: Keep the central area of the classification result image with a size of 128×128 pixels; According to the coordinates of the cropped image, the result image is spliced to the corresponding position to obtain the final prediction result image.
[0048] Example 2 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of a remote sensing image fine land use classification method based on DeepLabV3+ and ViT described in Embodiment 1 are implemented.
[0049] Example 3 This embodiment provides a computer device, including a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, the steps of a remote sensing image fine land use classification method based on DeepLabV3+ and ViT described in Example 1 are implemented.
[0050] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A remote sensing image fine land use classification method based on DeepLabV3+ and ViT, characterized in that: The method comprises: Construct a land use classification model based on DeepLabV3+ and ViT; Preprocess the public high-resolution remote sensing image land cover classification dataset to obtain the model dataset; The land use classification model based on DeepLabV3+ and ViT is trained and performance evaluated by using the model data set to obtain a final model; Obtain the Gaofen-1 remote sensing image as the target data to be classified, and preprocess the target data to obtain the target data set; The target data set is input into the final model for pixel-level classification prediction to obtain high-precision classification results.
2. The remote sensing image fine land use classification method based on DeepLabV3+ and ViT according to claim 1, characterized in that: The methods for building a land use classification model based on DeepLabV3+ and ViT include: Build the DeepLabV3+ model framework: The DeepLabV3+ model framework adopts an encoder-decoder structure. The encoder consists of the backbone network ResNet-101 and the dilated spatial convolution pooling pyramid pooling module ASPP. The main structure of ResNet-101 is: first a 7×7 convolution with a step size of 2, followed by a 3×3 maximum pooling with a step size of 2, and then a Bottleneck structure, where the Bottleneck structure is a residual structure composed of a 1×1 convolution, a 3×3 convolution and a 1×1 convolution; the main structure of ASPP is: the feature map output by the backbone network is input into a 1×1 convolution, a dilated convolution with expansion rates of 6, 12, and 18, and a global average pooling module, respectively, to obtain five feature maps, which are spliced and then passed through a 1×1 convolution; the decoder uses a 1×1 convolution to reduce the dimension of the low-level feature map output in the middle of the backbone network, and then upsamples the feature map output by ASPP to obtain two feature maps, which are spliced together and then passed through a 3×3 convolution; Build the Vision Transformer module: The Vision Transformer module includes an Embedding layer, a Transformer Encoder layer, and an MLPHead layer. The Embedding layer is used to process the input image into a token sequence that the Transformer can receive. The input image dimension is set to H×W×C, and the size of each block is P×P. First, the input image is divided into a series of two-dimensional sequences according to the size of the block to obtain N blocks, that is, ,in , and N is the effective length of the input sequence of the Transformer module. Then, each block is mapped to a one-dimensional vector through linear transformation. Then, each one-dimensional vector is concatenated with the corresponding position code to form an image token, and a category code is added at the front end. This category token is used for the final category output, and finally the input to the Transformer Encoder layer is obtained. The expression is: (1) in, Parameters added for encoding of blocks; Parameters added for positional encoding, Code the categories; The Transformer Encoder layer is composed of L stacked Transformer Encoder modules. The structure of each Transformer Encoder module is mainly composed of a multi-head self-attention mechanism and a multi-layer perception mechanism. A normalization module is set before the multi-head self-attention mechanism and the multi-layer perception mechanism, and the residual structure is used to fuse the input features and the output features. The operation expression is: (2) (3) Among them, the MLP module consists of full connection, GELU activation function and Dropout. The MLP Head layer is used to classify the image and finally output the classification result. The MLP Head layer is composed of a Layer Norm module and a linear mapping module, and the expression is: (4) In the DeepLabV3+ model framework, the Vision Transformer module is introduced to obtain the land use classification model based on DeepLabV3+ and ViT: The land use classification model based on DeepLabV3+ and ViT adopts an encoder-decoder structure, and the encoder is composed of a backbone network ResNet-101, ASPP and Vision Transformer; the remote sensing image is input into the backbone network ResNet-101, and a high-level feature map is obtained after a deep convolution operation, which is respectively input into the ASPP module and the Vision Transformer module. The ASPP module inputs the obtained high-level feature map into a 1×1 convolution, a dilated convolution with expansion rates of 6, 12, and 18, and a global average pooling module to obtain 5 feature images, which are spliced to obtain a high-level feature map that integrates multi-scale information; the Vision Transformer module processes the obtained high-level feature map through a PatchEmbedding layer and then inputs it into the Transformer layer, which is composed of 12 Transformer modules, and the obtained global high-level feature image is processed by Patch After processing by the Expanding layer, the image is restored to the size before the input module. The decoder inputs the low-level feature map output by the backbone network into a 1×1 convolution to obtain a high-level feature map, inputs the high-level feature map obtained by the ASPP module that integrates multi-scale information into a 1×1 convolution and performs a 4-fold upsampling to obtain a multi-scale low-level feature map, and inputs the high-level feature map obtained by the Vision Transformer module that extracts global features into a 1×1 convolution and performs a 4-fold upsampling to obtain a global low-level feature map. The three are spliced together. Finally, the spliced feature map is input into a 3×3 convolution and performed a 4-fold upsampling to obtain a predicted image with the same size as the original image.
3. The remote sensing image fine land use classification method based on DeepLabV3+ and ViT according to claim 1, characterized in that: The method of preprocessing the public high-resolution remote sensing image land cover classification dataset to obtain the model dataset includes: The images of the public high-resolution remote sensing image land cover classification dataset were cropped into 256*256 images so that they can be input into the model for training; By data enhancement, the images of land use types with fewer samples in the public high-resolution remote sensing image land cover classification dataset are expanded to obtain an enhanced dataset; The enhanced data set is divided into a source domain data set and a target domain data set, and then the target domain data set is divided into a target domain training set, a target domain verification set and a target domain test set according to a preset ratio.
4. The method for fine land use classification of remote sensing images based on DeepLabV3+ and ViT according to claim 3 is characterized in that: The method of training and evaluating the performance of the land use classification model based on DeepLabV3+ and ViT by using the model data set to obtain the final model includes: Inputting the source domain dataset into the land use classification model based on DeepLabV3+ and ViT for training to obtain a pre-trained model; Inputting the target domain training set into the pre-trained model for fine-tuning to obtain a fine-tuned model; Inputting the target domain validation set into the fine-tuning model to optimize the loss function and adjust the hyperparameters to obtain an improved model; The target domain test set is input into the improved model and the model performance is evaluated according to the preset evaluation index, and the model with the best performance is retained as the final model.
5. The method for fine land use classification of remote sensing images based on DeepLabV3+ and ViT according to claim 4 is characterized in that: The preset evaluation indicators include: IoU, average IoU, overall pixel accuracy, Kappa coefficient, and execution time.
6. The remote sensing image fine land use classification method based on DeepLabV3+ and ViT according to claim 1, characterized in that: The method of obtaining the Gaofen-1 remote sensing image as the target data to be classified and preprocessing the target data to obtain the target data set includes: The image is cropped using overlapping sliding windows and dilation prediction to obtain the target data set, specifically: The right and lower boundaries of the original remote sensing image are filled with 0 to obtain a first preprocessed image, wherein the size of the filled image must be an integer multiple of the sliding prediction window; To perform expansion prediction on the edge data, an outer border is added to the filled image to obtain a second preprocessed image, wherein the outer border is filled with 0, and the filling size is 1 / 2 of the sliding window step size; The second preprocessed image is cropped using a sliding window of size 256×256 and a sliding step of 128, and the coordinates of the cropped image are recorded to obtain a target data set.
7. The method for fine land use classification of remote sensing images based on DeepLabV3+ and ViT according to claim 6, characterized in that: The method of inputting the target data set into the final model for pixel-level classification prediction to obtain a high-precision classification result includes: The target data set is input into the final model to perform pixel-level classification prediction and output high-precision classification results. The quality of each scene data is checked manually visually, and the classification results are spliced and finally displayed visually.
8. The remote sensing image fine land use classification method based on DeepLabV3+ and ViT according to claim 7, characterized in that: The splicing processing method includes: Keep the central area of the classification result image with a size of 128×128 pixels; According to the coordinates of the cropped image, the result image is spliced to the corresponding position to obtain the final prediction result image.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a remote sensing image fine land use classification method based on DeepLabV3+ and ViT are implemented as described in any one of claims 1 to 8.
10. A computer device, characterized in that: The invention comprises a storage medium, a processor and a computer program stored in the storage medium and executable by the processor, wherein when the computer program is executed by the processor, the steps of a remote sensing image fine land use classification method based on DeepLabV3+ and ViT are implemented as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Land utilization classification method and system
CN114821340A