A land cover classification model for multi-scale remote sensing images

Through the interactive fusion of ViT and CNN encoders and semantic alignment upsampling, the problem of insufficient generalization of remote sensing land object classification methods in terms of resolution and geographic coverage is solved, high-precision land object classification across resolutions is achieved, and the cost of model development and application is reduced.

CN120411622BActive Publication Date: 2025-10-10ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510480407.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-10-10
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing deep learning-based remote sensing object classification methods have insufficient generalization in practical applications, especially affected by spatial resolution and geographic coverage, resulting in high model development and application costs, and are unable to effectively capture the correspondence between different resolutions.

Method used

The ViT encoder, low-rank fine-tuning module, CNN encoder, decoder, scale encoder and position encoder are used to interactively fuse the Transformer layer with CNN features, combined with the attention mechanism and semantic alignment upsampling to achieve surface cover classification of multi-scale remote sensing images.

Benefits of technology

It realizes automatic interpretation across resolutions of a single model, supports high-precision object classification of remote sensing images with resolutions of 2 to 10 meters, reduces development and application costs, and improves interpretation stability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411622B_ABST
    Figure CN120411622B_ABST
Patent Text Reader

Abstract

The application discloses a kind of surface cover classification models for multi-scale remote sensing image, including ViT encoder, low-rank fine-tuning module, CNN encoder, decoder, scale encoder and position encoder;Multi-resolution remote sensing image is handled after entering Transform layer by ViT encoder block coding layer, then interact and fuse with the CNN feature handled after CNN encoder, the visual feature after fusion is fused with scale coding, position coding by attention mechanism, and then enters decoder, the class probability corresponding to each pixel is obtained by the operation of decoder, and finally the semantic segmentation head determines the ground object class.The application realizes single model cross-resolution automatic interpretation, supports 2 to 10 meter resolution remote sensing image multi-element ground object classification.Visual pre-training large model is introduced to improve the stability of interpretation;And through ViT and CNN dual branch, global context and local detail information are considered;And semantic alignment up-sampling is introduced in the decoding process, to improve the segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing applications, and in particular to a land cover classification model for multi-scale remote sensing images. Background Art

[0002] Remote sensing monitoring is widely used in fields such as natural resource management, urban construction, and emergency disaster relief. Fast, stable, and efficient object classification is the foundation of many higher-level applications. Remote sensing object classification methods based on deep learning offer the advantages of high accuracy and speed, and have become the mainstream approach in this field, far outperforming traditional manual modeling and machine learning methods. Semantic segmentation methods with encoder-decoder structures, represented by convolutional networks such as FCN, DeepLab, and Unet, have led the field of remote sensing object classification in recent years. Furthermore, with the continued advancement of Transformer technology in the field of image processing, an increasing number of Transformer-based remote sensing object classification methods have been proposed.

[0003] While existing methods have achieved promising results on academic datasets, they still suffer from significant generalization issues in practical applications, particularly due to the significant impact of spatial resolution and geographic coverage. A common solution is to train multiple models for different resolutions and restrict their application to specific geographic regions. This strategy results in high model development and application costs and fails to capture the corresponding relationships between different resolutions, limiting the model's ability to achieve better generalization. Summary of the Invention

[0004] To solve the existing problems, the present invention provides a land cover classification model for multi-scale remote sensing images. The specific scheme is as follows:

[0005] A land cover classification model for multi-scale remote sensing images, including a ViT encoder, a low-rank fine-tuning module, a CNN encoder, a decoder, a scale encoder, and a position encoder;

[0006] The multi-resolution remote sensing image is processed by the ViT encoder block encoding layer and then enters the Transformer layer, and then interactively fused with the CNN features processed by the CNN encoder. The fused visual features are fused with the scale code and position code through the attention mechanism and then enter the decoder. The decoder obtains the category probability corresponding to each pixel through multiple steps, and finally the semantic segmentation head determines the ground object category.

[0007] Preferably, the ViT encoder adopts the ViT-L architecture, including 1 block coding layer and 24 Transformer layers; the block coding layer converts the input image Split into Image blocks, where H×W×3 represents the length and width of the image, 3 channels, and each image block is mapped to a high-dimensional vector of 1024 dimensions. Specifically, the block coding layer is implemented by a convolutional layer with a kernel size of 16 and a step size of 16. After the block coding layer, the feature map size is After that, the ViT encoder output F is obtained through 24 Transformer layers. ViT has a barrel structure, and the Transformer layer does not change the size and dimension of the feature map; the 24 Transformer layers are equally divided into 4 groups, and each group interacts with the CNN features through the feature fusion module FFM.

[0008] Preferably, the low-rank fine-tuning module is used to reduce the amount of downstream scene parameter updates to adapt to downstream scene training. During the training process, the ViT parameters are frozen and only the low-rank fine-tuning module parameters are updated, thereby reducing sample requirements and model update time. The low-rank fine-tuning module consists of two fully connected layers, which are inserted into the self-attention calculation process of the Transformer layer. The fully connected layers at the insertion positions Q and V are inserted. The forward process after insertion is shown as follows: x'=Wx+BAx, where represents the fully connected layer parameters, C in and C out are the input and output channel dimensions, respectively, and To fine-tune the module parameters for low rank, it is implemented through the fully connected layer, and r is much smaller than C in and C out , to reduce training time and sample requirements.

[0009] Preferably, the CNN encoder is used to introduce multi-scale features; the CNN encoder is developed based on the residual network ResNet18, including a neck layer and 4 groups of residual connection blocks, each block includes multiple convolutional layers, and a feature fusion module FFM is inserted after each residual connection block. A connection is established between the ViT feature F and the convolution feature through a bidirectional attention mechanism, and its output is the ViT feature after feature interaction update. and CNN features and Feature fusion is performed through the FinalFuse module, which includes a convolution layer with a stride of 2 and a feature splicing operation. After the convolutional layer, Feature splicing is performed to obtain E, which is then fused with the scale code S and position code L through a cross-attention mechanism. The feature is then fed into the decoder, and the object classification result is obtained after the pyramid pooling module PPM and step-by-step upsampling.

[0010] Preferably, the decoder adopts step-by-step upsampling, and uses the semantic alignment upsampling module AlignUp to Complete cross-level feature fusion; the output of the bottom-level cross-level feature fusion is P3 and After fusion again, P2 and After fusion, we get P2 and P3 are upsampled to the size of P1 through bilinear interpolation. After the three are spliced ​​in the channel dimension, they are subjected to a depthwise separable convolution and a 1×1 convolution to obtain the category probability corresponding to each pixel; the multi-scale pooling module PPM enhances the model's multi-scale information extraction capability by setting multiple pooling layers of different sizes; the low-level feature X l Through upsampling operation and high-level feature X h Size alignment, upsampling and feature concatenation to obtain X c ;X c Enter two branches to learn sampling deviation Δ∈R 2×H×W , each branch consists of 3 layers of convolution; the sampling offset and input features are obtained by deviating from the sampling function U to obtain the semantic alignment feature X' l 、X' h , the two are added element by element to get the final output X'; the deviation sampling function U is shown as follows:

[0011]

[0012] Preferably, the scale encoder supports 2 to 10 meter resolution image feature extraction, specifically 4 resolutions of 2 meters, 4 meters, 8 meters, and 10 meters. The resolution information is represented as a high-dimensional vector Feature fusion is achieved with the visual feature E through the attention mechanism. S is initialized to 0 and the parameters are updated as the model is trained.

[0013] Preferably, the position encoder uses the SatCLIP algorithm to associate the image geographic location with the visual features and convert the image latitude and longitude coordinates into a high-dimensional vector Reflects local natural geography and land cover characteristics; L realizes feature fusion with visual features E through the attention mechanism.

[0014] The beneficial effects of the present invention are:

[0015] This invention achieves automatic interpretation across resolutions using a single model, supporting multi-element object classification in remote sensing imagery with resolutions from 2 to 10 meters. It introduces a large pre-trained visual model to improve interpretation stability, while using Vision Inference (ViT) and CNN as dual branches to balance global context and local details. Furthermore, semantically aligned upsampling is introduced during the decoding process to improve segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 It is the overall principle block diagram of the model of the present invention;

[0018] Figure 2 It is the Transformer layer structure;

[0019] Figure 3 Schematic diagram of the cross attention mechanism;

[0020] Figure 4 Schematic diagram of the bidirectional attention mechanism;

[0021] Figure 5 It is a structural diagram of the multi-scale pooling module PPM;

[0022] Figure 6 This is a structural diagram of the alignment upsampling module AlignUp;

[0023] Figure 7 This is a rendering after being processed by the model of the present invention. DETAILED DESCRIPTION

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0025] The present invention belongs to the field of remote sensing application technology. Based on deep learning semantic segmentation and a large visual pre-training model, a remote sensing image feature classification method that integrates scale and geocoding is proposed for optical satellite imagery. The method realizes pixel-level classification of 12 categories, overcomes the shortcomings of existing methods in spatial scale and geographic migration, and can output highly consistent automatic interpretation results at multiple resolutions, significantly reducing the development and application costs of remote sensing automatic interpretation.

[0026] The present invention proposes a surface cover classification model for multi-scale remote sensing images, which can be efficiently embedded in a large visual pre-training model and achieves high-precision ground object classification of remote sensing images at various resolutions within the range of 2 to 10 meters by introducing scale and position encoding.

[0027] The model adopts a classic encoder-decoder structure, which includes two encoders.

[0028] The ViT encoder comes from a large visual pre-training model with an original parameter of 1 billion. It was self-supervised pre-trained on more than 100 million online images, and then the number of model parameters was reduced to 300 million through knowledge distillation.

[0029] The CNN encoder uses a lightweight residual network ResNet18, on which deformable convolution is introduced to increase the model's receptive field.

[0030] The decoder adopts a bottom-up feature aggregation strategy, alleviates the semantic misalignment problem introduced by feature upsampling through the semantic alignment module, and improves the receptive field and enhances the context parsing capability through the dual attention module.

[0031] The decoder is connected to a semantic segmentation head at the end, which maps the high-dimensional feature map generated by the codec into the classification probability of each pixel position. The one with the highest probability is the ground object category of the corresponding pixel.

[0032] The scale code is represented by a high-dimensional vector, initialized to 0, and updated during training. Its feature dimension is consistent with the encoder output feature.

[0033] Position encoding is implemented via SatCLIP, which converts the latitude and longitude positions into high-dimensional vectors and passes them through a projection layer to keep the dimension consistent with the encoder output.

[0034] Both scale encoding and position encoding are embedded in the encoder features through the bidirectional attention mechanism and enter the decoder together to finally generate the semantic segmentation results.

[0035] Specifically, if Figure 1 , a surface cover classification model for multi-scale remote sensing images, including a ViT encoder, a low-rank fine-tuning module, a CNN encoder, a decoder, a scale encoder, and a position encoder.

[0036] The multi-resolution remote sensing image is processed by the ViT encoder block encoding layer and then enters the Transformer layer, and then interactively fused with the CNN features processed by the CNN encoder. The fused visual features are fused with the scale code and position code through the attention mechanism and then enter the decoder. The decoder obtains the category probability corresponding to each pixel through multiple steps, and finally the semantic segmentation head determines the ground object category.

[0037] The ViT encoder adopts the ViT-L architecture, which includes 1 block coding layer and 24 Transformer layers; the block coding layer converts the input image Split into Image blocks, where H×W×3 represents the length and width of the image, 3 channels, and each image block is mapped to a high-dimensional vector of 1024 dimensions. Specifically, the block coding layer is implemented by a convolutional layer with a kernel size of 16 and a step size of 16. After the block coding layer, the feature map size is After that, the ViT encoder output F is obtained through 24 Transformer layers. ViT has a barrel structure, and the Transformer layer does not change the size and dimension of the feature map. The structure of each Transformer layer is as follows Figure 2 As shown in the figure, the 24 layers are equally divided into 4 groups, and each group interacts with the CNN features through the feature fusion module FFM.

[0038] ViT parameters are large, while downstream scenarios usually have limited sample size and computing resources, which cannot support ViT's full parameter training. Therefore, a low-rank fine-tuning strategy is used to reduce the amount of parameter updates in downstream scenarios. During training, ViT parameters are frozen and only the low-rank fine-tuning module parameters are updated, thereby reducing sample requirements and model update time.

[0039] The low-rank fine-tuning module consists of two fully connected layers, which are inserted into the self-attention calculation process of the Transformer layer. The self-attention structure of the low-rank fine-tuning module is as follows: Figure 2 As shown. The fully connected layer at the insertion position is Q and V. The forward process after insertion is as follows:

[0040] x'=Wx+BAx, where represents the fully connected layer parameters, C in and C out are the input and output channel dimensions, respectively, and To fine-tune the module parameters for low rank, it is implemented through the fully connected layer, and r is much smaller than C in and C out , to reduce training time and sample requirements.

[0041] The CNN encoder is used to introduce multi-scale features. ViT excels at capturing semantic connections between pixels at different locations in an image, but its barrel-like structure is not conducive to dense visual tasks such as semantic segmentation. The multi-scale design common in CNNs often achieves good results in downstream remote sensing tasks. The introduction of a multi-scale CNN structure effectively supplements the lack of multi-scale information in images provided by ViT features.

[0042] The CNN encoder is developed based on the residual network ResNet18, which includes a neck layer and 4 groups of residual connection blocks. Each block includes multiple convolutional layers. A feature fusion module FFM is inserted after each residual connection block. The connection is established between the ViT feature F and the convolution feature through the bidirectional attention mechanism. Its output is the ViT feature after the feature interaction update. and CNN features and Feature fusion is performed through the FinalFuse module, which includes a convolution layer with a stride of 2 and a feature splicing operation. After the convolutional layer, The feature splicing is performed to obtain E, which is then fused with the scale code S and position code L through the cross attention mechanism. After that, it enters the decoder and obtains the object classification result after the pyramid pooling module PPM and step-by-step upsampling. The cross attention calculation process is as follows: Figure 3 As shown. The schematic diagram of the bidirectional attention mechanism is as follows Figure 4 shown.

[0043] The decoder adopts step-by-step upsampling. The visual feature E generated by the encoder is first fused with the scale feature S and the position feature L through the attention mechanism, and then passes through the multi-scale pooling module PPM, and then is combined with the semantic alignment upsampling module AlignUp. Complete cross-level feature fusion. The output of the bottom-level cross-level feature fusion is P3 and After fusion again, P2 and After fusion, we get P2 and P3 are upsampled to the size of P1 through bilinear interpolation. After the three are spliced ​​in the channel dimension, they are subjected to a depthwise separable convolution and a 1×1 convolution to obtain the category probability corresponding to each pixel. The multi-scale pooling module PPM enhances the model's multi-scale information extraction capability by setting multiple pooling layers of different sizes. The structure of the multi-scale pooling module PPM is as follows: Figure 5 shown.

[0044] The structure of the semantic alignment upsampling module AlignUp is as follows Figure 6 As shown. Low-level features X l Through upsampling operation and high-level feature X h Size alignment, upsampling and feature concatenation to obtain X c ;X c Enter two branches to learn sampling deviation Δ∈R 2×H×W, each branch consists of 3 layers of convolution; the sampling offset and input features are obtained by deviating from the sampling function U to obtain the semantic alignment feature X' l 、X' h , the two are added element by element to get the final output X'. The deviation sampling function U is shown as follows:

[0045]

[0046] The scale encoder supports ground feature extraction from 2 to 10 meter resolution images, specifically 2 meters, 4 meters, 8 meters, and 10 meters, covering most common domestic and foreign civil and commercial satellites. The resolution information is represented as a high-dimensional vector Feature fusion is achieved with the visual feature E through the attention mechanism. S is initialized to 0 and the parameters are updated as the model is trained.

[0047] The position encoder uses the SatCLIP algorithm to link the image geographic location with visual features and convert the image latitude and longitude coordinates into a high-dimensional vector Reflecting local natural geography and land cover characteristics. L achieves feature fusion with visual features E through the attention mechanism.

[0048] like Figure 7 , which is the effect diagram after being processed by the model of the present invention. The present invention has developed a method to use the high generalization ability of the pre-trained large model to improve the accuracy and stability of the model in remote sensing downstream tasks. At the same time, the spatial resolution and imaging position are introduced into the model decision process through scale and position encoding learning, avoiding the classification errors and instability problems that may be caused by relying solely on visual clues. In addition, according to the characteristics of remote sensing land object classification tasks, a multi-scale feature aggregation and semantic alignment module is designed, and a remote sensing semantic segmentation framework with a codec structure is built. On this framework, the fusion embedding of visual pre-training and downstream tasks is realized, and remote sensing land object classification with high accuracy and high stability is achieved.

[0049] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0050] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A land cover classification model for multi-scale remote sensing images, characterized by: It includes a ViT encoder, a low-rank fine-tuning module, a CNN encoder, a decoder, a scale encoder, and a position encoder. The multi-resolution remote sensing image is processed by the block coding layer of the ViT encoder and then enters the Transformer layer to obtain the ViT encoder output F. The multi-resolution remote sensing image is processed by the CNN encoder to obtain CNN features. The ViT encoder output F is then interactively fused with the CNN features. The fused visual features are fused with the scale code and position code through the attention mechanism and then enter the decoder. The scale encoding is to represent the resolution information as a high-dimensional vector, the position encoding is to convert the latitude and longitude coordinates of the image into a high-dimensional vector, the decoder obtains the category probability corresponding to each pixel through multiple steps, and finally the semantic segmentation head determines the category of the ground object; the low-rank fine-tuning module is used to reduce the amount of downstream scene parameter updates to adapt to downstream scene training. The ViT parameters are frozen during training, and only the low-rank fine-tuning module parameters are updated.

2. The model according to claim 1, characterized in that: The ViT encoder adopts the ViT-L architecture, which includes 1 block coding layer and 24 Transformer layers; the block coding layer converts the input image Split into Image blocks, where H×W×3 represents the length, width, and three channels of the image, and each image block is mapped to a high-dimensional vector of 1024 dimensions; the block coding layer is implemented by a convolutional layer with a kernel size of 16 and a stride of 16; After the block coding layer, the feature map size is After that, the ViT encoder output F is obtained through 24 Transformer layers. ViT has a barrel structure, and the Transformer layer does not change the size and dimension of the feature map; the 24 Transformer layers are equally divided into 4 groups, and each group interacts with the CNN features through the feature fusion module FFM.

3. The model according to claim 1, characterized in that: The low-rank fine-tuning module consists of two fully connected layers, which are inserted into the self-attention calculation process of the Transformer layer. The fully connected layers at positions Q and V are inserted. The forward process after insertion is as follows: x'=Wx+BAx, where represents the fully connected layer parameters, C in and C out are the input and output channel dimensions, respectively, and To fine-tune the module parameters for low rank, it is implemented through the fully connected layer, and r is much smaller than C in and C out , to reduce training time and sample requirements.

4. The model according to claim 1, characterized in that: The CNN encoder is used to introduce multi-scale features. The CNN encoder is developed based on the residual network ResNet18, and includes a neck layer and 4 groups of residual connection blocks. Each block includes multiple convolutional layers. A feature fusion module FFM is inserted after each residual connection block. A connection is established between the ViT feature F and the convolution feature through a bidirectional attention mechanism. Its output is the ViT feature after feature interaction update. and CNN features and Feature fusion is performed through the FinalFuse module, which includes a convolution layer with a stride of 2 and a feature splicing operation. After the convolutional layer, Feature splicing is performed to obtain E, which is then fused with the scale code S and position code L through a cross-attention mechanism. The feature is then fed into the decoder, and the object classification result is obtained after the pyramid pooling module PPM and step-by-step upsampling.

5. The model according to claim 4, characterized in that: The decoder adopts step-by-step upsampling and uses the semantic alignment upsampling module AlignUp to Complete cross-level feature fusion; the feature E that combines scale encoding and position encoding with The output of cross-level feature fusion is P3 and After fusion again, P2 and After fusion, we get P2 and P3 are upsampled to the size of P1 through bilinear interpolation. After the three are spliced ​​in the channel dimension, they are subjected to a depthwise separable convolution and a 1×1 convolution to obtain the category probability corresponding to each pixel. The pyramid pooling module PPM enhances the model's multi-scale information extraction capability by setting multiple pooling layers of different sizes. The structure of the semantic alignment upsampling module AlignUp is as follows: low-level feature X l Through upsampling operation and high-level feature X h Size alignment, upsampling and feature concatenation to obtain X c ;X c Enter two branches to learn sampling deviation Δ∈R 2×H×W , each branch consists of 3 layers of convolution; the sampling deviation and input features are obtained by deviating from the sampling function U to obtain the semantic alignment feature X' l 、X' h , the two are added element by element to get the final output X' l+h ; The deviation sampling function U is shown as follows:

6. The model according to claim 4, characterized in that: The scale encoder supports image feature extraction with a resolution of 2 to 10 meters, specifically 4 resolutions: 2 meters, 4 meters, 8 meters, and 10 meters. The resolution information is represented as a high-dimensional vector Feature fusion is achieved with the visual feature E through the attention mechanism. S is initialized to 0 and the parameters are updated as the model is trained.

7. The model according to claim 4, characterized in that: The position encoder uses the SatCLIP algorithm to link the image geographic location with visual features and convert the image latitude and longitude coordinates into a high-dimensional vector Reflects local natural geography and land cover characteristics; L realizes feature fusion with visual features E through the attention mechanism.

Citation Information

Patent Citations

  • Enhanced semantic-position feature fusion network method and device based on different pre-training feature extraction backbone and application

    CN118196628A

  • Real-Time Scene Understanding System

    US20200302214A1