Landscape pattern classification method combining global-local multi-order deep and shallow layer features
Through the Swin Transformer-CNN hybrid model, the depth and shallow feature cross-fusion module and the global-local attention module, the global and local information are integrated, and the problem of insufficient information integration in landscape pattern classification is solved, the classification accuracy is improved, and it is of great significance to ecosystem monitoring.
Patent Information
- Application Number
- CN202510160235.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-13
AI Technical Summary
It is difficult for the prior art to effectively integrate global and local information in landscape pattern classification, resulting in information loss or conflict, affecting classification accuracy.
Using the Swin Transformer-CNN hybrid model, multi-scale features are integrated and global context information and local detail information are aggregated by introducing a shallow-level feature cross-fusion module and a global-local attention module.
It improves the accuracy of landscape pattern classification, improves the spectral characteristic similarity and intra-class differences between land objects in complex landscape scenarios, and is of great significance to dynamic monitoring of ecosystems.
Smart Images

Figure CN120088551A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of photogrammetric data processing, and particularly relates to a landscape pattern classification method combining global-local multi-order deep and shallow features. Background Art
[0002] Ecosystems play an important role in the global carbon cycle and are of great significance in aspects such as degrading pollution, conserving water sources, regulating the climate, and protecting biodiversity. However, due to the impacts of human activities and natural disasters, ecosystems are facing problems of ecological function degradation and biodiversity reduction. Therefore, it is urgent to protect ecosystems. Obtaining accurate long-term landscape pattern information through dynamic monitoring of ecosystem landscape patterns and analyzing evolution characteristics is crucial for the protection and restoration of ecosystem resources.
[0003] With the rapid development of remote sensing technology and the progress of the earth observation system, a large amount of observation data has been provided for dynamic monitoring of landscape patterns. Remote sensing technology has the advantages of fast acquisition speed, wide coverage, and large amount of information contained, and has become an important means for landscape pattern monitoring. However, the problem of mixed pixels in medium- and low-resolution image data has affected the accuracy of landscape pattern information extraction to a certain extent. Therefore, classifying ecosystem landscape patterns using high-resolution remote sensing images has become a current research hotspot and achieved relatively high classification accuracy.
[0004] The method of using manual interpretation for landscape pattern classification is time-consuming and laborious. Machine learning methods, such as Random Forest (RF), Classification and Regression Tree (CART), and Support Vector Machine (SVM), etc., although they have been widely used in landscape pattern mapping, their accuracy still needs to be improved. With the development of artificial intelligence, methods based on deep learning have achieved remarkable results in remote sensing image segmentation. The method based on Convolutional Neural Network (CNN) performs well in mapping complex landscapes, but has limitations in capturing global context information. The Transformer structure conducts global context modeling through a global self-attention mechanism and shows superiority in image segmentation tasks. However, the Transformer still has deficiencies in capturing local information. Combining the methods of CNN and Transformer can give full play to their advantages in local and global information extraction. However, when these hybrid structure methods integrate features at different levels, they are prone to ignoring the semantic gap between features, resulting in information loss or conflict.
[0005] How to use deep learning technology to obtain more local and global information and optimize the hybrid structure method to improve the accuracy of landscape pattern classification is a major problem in the field of landscape pattern classification at the present stage. Summary of the Invention
[0006] The object of the present invention is to overcome the above problems existing in the traditional technology, and provide a landscape pattern classification method that combines global-local multi-order deep and shallow features, so as to improve the accuracy of landscape pattern classification.
[0007] To achieve the above technical objectives and reach the above technical effects, the present invention is realized through the following technical solutions:
[0008] The present invention provides a landscape pattern classification method that combines global-local multi-order deep and shallow features, including the following steps:
[0009] Step 1: Construction of the Swin Transformer module
[0010] The Swin Transformer-CNN hybrid model (STCNet) uses four stages of Swin Transformer blocks as the encoder to perform multi-scale hierarchical feature extraction by a hierarchical construction method. And a deep and shallow feature cross-fusion module (DSFCF) is introduced between the encoder and the decoder to establish a connection, and cross-fusion is performed between adjacent features output by the encoder to establish feature complementarity and alleviate the semantic confusion problem generated during the fusion of features at different levels.
[0011] Step 2: Construction of the deep and shallow feature cross-fusion module (DSFCF)
[0012] There is often a certain semantic gap between the deep features and shallow features generated in the encoder stage. Simple skip connections are prone to cause semantic information confusion during feature fusion. To alleviate this semantic gap, the present invention introduces a deep and shallow feature cross-fusion module (DSFCF) between the encoder and the decoder to establish a connection, and cross-fusion is performed between adjacent features output by the encoder to aggregate multi-scale features and establish feature information complementarity between adjacent feature layers, effectively alleviating the semantic confusion problem generated during the fusion of features at different levels.
[0013] Step 3: Construction of the global-local attention module (GLAB)
[0014] In remote sensing images, the distribution of landscape pattern features is discrete, the detailed information is redundant, and different categories are intertwined with each other, which is prone to the problem of similar features between different classes and different features within the same class. Integrating global and local features can better perceive the spatial correlation of the information in the image and identify the differences in features at multiple scales. Therefore, the present invention introduces three global-local attention modules (GLAB) in the decoder part, fuses the different-level features output by the DSFCF module with the features of the decoder upsampled to restore the resolution, and then transports them to the global-local attention module (GLAB) to obtain a hierarchical expression of the features, and aggregates the global context semantic information and local detailed information at different scales.
[0015] Step 4: Model accuracy verification and analysis
[0016] Use the high-resolution image dataset for experiments, compare and verify the present invention with other semantic segmentation models, and conduct quantitative evaluation.
[0017] The beneficial effects of the present invention are as follows:
[0018] 1. The present invention constructs a hybrid network model SwinTransformer-CNN using a convolutional neural network (CNN) and Swin Transformer. First, a deep and shallow feature cross-fusion module is introduced between the encoder and the decoder to alleviate the semantic gap between different hierarchical features generated by SwinTransformer. By introducing three global-local attention modules (GLAB) in the decoder part, the global context information and local detail information are aggregated, and the discrimination of intra-class feature consistency and inter-class feature difference is strengthened.
[0019] 2. The present invention makes full use of the advantages of both CNN and Swin Transformer in local and global information extraction, aggregates global context information and local fine-grained information, improves the problems of spectral feature similarity between ground object categories and intra-class difference in complex landscape scenes, improves the accuracy of landscape pattern classification, and is of great significance for the dynamic monitoring of ecosystems. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is the overall framework diagram of the Swin Transformer-CNN hybrid model;
[0022] Figure 2 It is the structure diagram of the Swin Transformer block;
[0023] Figure 3 It is the DSFCF structure diagram [(a) and (c) represent the cross-fusion process of two boundary branches and adjacent feature layers, and (b) represents the cross-fusion process diagram of two intermediate branches and adjacent feature layers];
[0024] Figure 4 It is the GLAB framework diagram;
[0025] Figure 5 It is the comparison diagram of the extraction results of the method of the present invention and other model methods; Detailed implementation mode
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] The present invention provides a deep learning method: The purpose of the present invention is to provide a landscape pattern classification method that combines global-local multi-order deep and shallow features, and designs a targeted deep learning algorithm to achieve high-precision and high-efficiency classification of landscape patterns under multi-temporal remote sensing images. Under the guidance of this idea, as Figures 1-5 shown, a landscape pattern classification method that combines global-local multi-order deep and shallow features is designed, including the following steps:
[0028] Step 1: Construction of the Swin Transformer module
[0029] The standard Transformer block uses the multi-head self-attention mechanism (MSA) to perform global self-attention calculations on all tokens of the image, with high computational complexity. The Swin Transformer block replaces the traditional MSA with the window-based Transformer block (W-MSA) and the window-sliding-based Transformer block (SW-MSA) to improve computational efficiency. The image is divided into non-overlapping windows, and self-attention calculations are performed within the local windows. Information connections are established between the windows through sliding windows, enhancing the global feature modeling ability. As Figure 2 shown, the Swin Transformer block designed by the present invention consists of a window-based Transformer block and a window-sliding-based Transformer block. Among them, the window-based Transformer block includes W-MSA, multi-layer perceptron (MLP), and layer normalization (LN). The window-sliding-based Transformer block includes SW-MSA, multi-layer perceptron (MLP), and layer normalization (LN). Both apply residual connections internally. The calculation process of the Swin Transformer block is represented by the formula as:
[0030]
[0031] Among them, st l-1 represents the input feature of W-MSA, and represent the output features of W-MSA and SW-MSA, respectively. st l and stl+1 Represents the output features after LN and MLP.
[0032] The encoder is composed of Swin Transformer blocks in four stages to gradually extract higher-level abstract features. Specifically, the input remote sensing image is first segmented into non-overlapping blocks with a size of 4×4. The flattened feature dimension of one block is 4×4×3 = 48, and the processed feature dimension of one image is H / 4×W / 4×48. Then, a linear embedding layer projects the H / 4×W / 4×48 tensor to dimension C and feeds it into the Swin Transformer block, keeping the number of output tokens as H / 4×W / 4. The linear embedding layer and the Swin Transformer block form the first stage of the encoder, and the output feature dimension is H / 4×W / 4×C. As the network deepens, a block merge is applied before the Swin Transformer blocks in the 2nd, 3rd, and 4th stages. By merging blocks, the number of tokens becomes 1 / 4 of the original (reducing the feature resolution to 1 / 2 of the original), and the output feature dimension is increased. The output feature dimensions of the 2nd, 3rd, and 4th stages composed of block merge and Swin Transformer blocks are H / 8×W / 8×2C, H / 16×W / 16×4C, and H / 32×W / 32×8C respectively.
[0033] Step 2: Construction of the Deep and Shallow Feature Cross-Fusion Module (DSFCF)
[0034] The feature levels extracted at different stages of the Swin Transformer encoder reflect different feature information. Shallow features have a smaller receptive field and usually provide fine-grained information such as boundaries, shape contours, textures, and positions, while deep features have a larger receptive field and contain more abstract context semantic information. When performing feature fusion, due to the large semantic gap between deep and shallow features, a certain degree of semantic confusion will occur. Therefore, the present invention designs a deep and shallow feature cross module to integrate multi-scale features and reduce the semantic gap.
[0035] First, the feature maps output by the four-level Swin Transformer blocks are aligned in feature channels through channel alignment operations to adjust the dimension and distribution of the feature channels, enabling effective fusion of features at different levels in the same space and reducing computational traffic. The number of channels of the four-level feature maps is reduced from 128, 256, 512, and 1024 to 32 respectively, which can be expressed by the formula:
[0036]
[0037] Where, is the output feature of the encoder, is the output feature of CA channel alignment.
[0038] Then, the previous branch and the next branch of the second and third level branches in the middle are downsampled and upsampled respectively to convert them into the current branch features. Figure 1 The next branch of the first-level branch of the boundary and the previous branch of the fourth-level branch are upsampled and downsampled respectively. This process can be expressed by the formula:
[0039]
[0040] Among them, i is the i-th level feature, Down is the downsampling operation, and Up is the 2-fold bilinear interpolation upsampling operation. They represent the output features of the previous and next branches of the current branch after CA channel alignment, Respectively represent the output features after downsampling and upsampling operations; then, the features of the previous and next branches are concatenated with the features of the current branch to complete cross-fusion, and the cross-fused features of each level are subjected to 3×3 convolution operations to further mine fine-grained information; finally, the coordinate attention mechanism with spatial position perception is used to capture long-distance dependencies in space. This process can be expressed as:
[0041]
[0042] in, Indicates the characteristics of the current branch, They represent the feature maps of the previous layer and the next layer respectively. Concat is the feature concatenation operation. CBR is the 3×3 convolution, BN and Relu operations. A represents the coordinate attention mechanism. CF i It represents the result of cross-fusion of features at different levels.
[0043] Step 3: Global-local attention module (GLAB) construction
[0044] The global-local attention module (GLAB) is introduced in the decoder part of the model. GLAB consists of a local branch composed of multi-scale convolutions and a global branch of a window-based multi-head self-attention mechanism, which extract local detail information and global spatial dependency semantic features respectively. The local branch uses convolution groups with kernel sizes of 5×5, 3×3, and 1×1 to capture spatial and spectral features of different scales, thereby fully extracting local fine-grained features. The formula is:
[0045]
[0046] Among them, X L Represents the output features of the local branch.
[0047] In the global branch, first use a 1×1 convolution to expand the number of channels C of the input feature map X i ∈R (B*C*H*W) by 3 times. To reduce the computational complexity brought by the self-attention module, perform a window splitting operation on the input features. The feature map is divided into several non-overlapping windows, and the self-attention mechanism is executed within each window to reduce the overall computational cost and retain the global context information. Within the window, map the input sequence to Query (Q), Key (K), and Value (V) vectors respectively, which can be expressed as:
[0048] Q = XW Q , K = XW K , V = XW V (9)
[0049] where X is the representation of the input sequence, and W Q , W K , W V are learnable fully-weighted matrices. Obtain the attention weights by calculating the dot product between Q and K, and normalize the dot product result using the Softmax function to generate the attention weights. Multiply the attention weights by V to get each attention score, which can be expressed as:
[0050]
[0051] where d k is the channel dimension of the K vector, is the scaling factor used to alleviate the vanishing gradient problem. Divide the Q, K, and V vectors into h heads (i.e., different subspaces), and calculate the self-attention of each head respectively, which can be expressed as:
[0052]
[0053] where Q i , K i , V i are the projections of Q, K, and V on the i-th attention head. Concatenate the results of all attention heads and obtain the final multi-head self-attention output through a linear transformation, which can be expressed as:
[0054] Multihead(Q, K, V) = Concat(head 1 , head 2 ,..., head h )W 0
[0055] where W 0 is the linear transformation matrix of the science department.
[0056] By fusing the global context information obtained from the global branch and the local features extracted from the local branch, the comprehensive expression of multi-scale information is realized, which can be expressed as:
[0057]
[0058] Among them, X GL represents the result of the fusion of the global branch and the local branch.
[0059] Residual connections are applied inside GLAB, and its process can be expressed as:
[0060] X t = GLA(BN(X t-1 )) + X t-1 (13)
[0061] X t+1 = MLP(BN(X t )) + X t (14)
[0062] Among them, GLA represents the global-local attention part composed of the global and local branches, X t-1 represents the input feature of GLA and BN, and X t-1 represents the output feature of MLP and BN.
[0063] Step 4. Model accuracy verification and analysis
[0064] The accuracy of the classification result of the method of the present invention is verified by using three evaluation indexes of F1, OA, and IoU, and the specific formulas are as follows:
[0065]
[0066] In the formula, TP, FP, FN, and TN respectively represent the numbers of true positive predictions, false positive predictions, false negative predictions, and true negative predictions.
[0067] Specifically, the present invention provides an implementation case. The Swin Transformer-CNN proposed by the present invention and another 7 semantic segmentation models (DeepLabV3+, SegNet, SegFormer, HRFormer, UNetFormer, BANet, DCSwin) are experimented on the Shengjin Lake high-resolution image dataset ( Figure 5 ), and Table 1 and Table 2 list the overall and category evaluation index situations of each model.
[0068] The average IoU, average F1, and OA of the classification results of Swin Transformer-CNN on the high-resolution dataset of Shengjin Lake are 78.12%, 87.05%, and 93.23% respectively. In the segmentation results of each category, except for the forest land category, the F1 scores and mIoUs of other categories of Swin Transformer-CNN are much higher than those of SegNet.
[0069] The present invention adopts a strategy of cross-fusion of deep and shallow features and a global-local attention mechanism, which can more effectively improve the extraction of semantic information of the whole and details in the image. Compared with DeepLabV3+, the mIoU is 10.23% higher and the F1 is 7.89% higher.
[0070] The three indicators of SegFormer, HRFormer, UNetFormer, and DCSwin are higher than the previous CNN-based methods. There are still gaps in the segmentation accuracy of different categories compared with Swin Transformer-CNN. Compared with the sub-optimal method DCSwin, the average F1 score of the method of the present invention is 0.46% higher, the mIoU is increased by 0.57%, and the OA is increased by 0.15%.
[0071] Table 1 Comparison of results of different model methods
[0072]
Claims
1. A landscape pattern classification method combining global and local multi-order deep and shallow features, characterized in that: The following steps are involved: Step 1: Swin Transformer module construction; Step 2: Construction of deep and shallow feature cross-fusion module; Step 3: Construction of global-local attention module; Step 4: Model accuracy verification and analysis.
2. A landscape pattern classification method combining global and local multi-order deep and shallow features according to claim 1, characterized in that: The Swin Transformer block in step 1 is composed of a window-based Transformer block and a window-sliding Transformer block. The Swin Transformer block uses the window-based Transformer block and the window-sliding Transformer block to perform global self-attention calculation on all tokens of the image, divides the image into non-overlapping windows, performs self-attention calculation in the local window, and establishes information connection between windows through the sliding window; The window-based Transformer block includes W-MSA, a multi-layer perceptron, and layer normalization, and the window sliding-based Transformer block includes SW-MSA, a multi-layer perceptron, and layer normalization, both of which apply residual connections internally; The Swin Transformer block calculation process is expressed by the formula: Among them, st l-1 represents the input features of W-MSA, and represents the output features of W-MSA and SW-MSA, st l and st l+1 Represents the output features after LN and MLP.
3. A landscape pattern classification method combining global and local multi-order deep and shallow features according to claim 2, characterized in that: The encoder is composed of four stages of Swin Transformer blocks to gradually extract higher-level abstract features; Specifically, the input remote sensing image is first divided into non-overlapping blocks with a block size of 4×4. The feature dimension of a block after flattening is 4×4×3=48, and the feature dimension of one image after processing is H / 4×W / 4×48. The H / 4×W / 4×48 tensor is projected to dimension C using a linear embedding layer and then fed into the Swin Transformer block to keep the number of output tokens at H / 4×W / 4. The linear embedding layer and the Swin Transformer block form the first stage of the encoder, and the output feature dimension is H / 4×W / 4×C. As the network deepens, a block merge is applied before the Swin Transformer blocks in the second, third, and fourth stages to reduce the number of tokens to 1 / 4 of the original by merging blocks. The output feature dimensions of the second, third, and fourth stages composed of Patch Merging and Swin Transformer Block are H / 8×W / 8×2C, H / 16×W / 16×4C, and H / 32×W / 32×8C, respectively.
4. A landscape pattern classification method combining global and local multi-order deep and shallow features according to claim 3, characterized in that: In the step 2, the Swin Transformer encoder includes a deep-shallow feature crossover module for integrating multi-scale features and reducing the semantic gap.
5. A landscape pattern classification method combining global and local multi-order deep and shallow features according to claim 4, characterized in that: The steps of integrating multi-scale features and reducing the semantic gap in the deep and shallow feature cross module are as follows: First, the feature maps output by the four-level Swin Transformer blocks are aligned through the channel alignment operation to adjust the dimension and distribution of the feature channels so that the features of different levels can be effectively fused in the same space and the computational traffic can be reduced. The number of channels of the four-level feature maps is reduced from 128, 256, 512, and 1024 to 32 respectively, which can be expressed as: in, is the output feature of the encoder, Output features of CA channel alignment; Then, the previous branch and the next branch of the second and third level branches in the middle are downsampled and upsampled respectively to convert them into the same size as the feature map of the current branch. The next branch of the first level branch at the boundary and the previous branch of the fourth level branch are upsampled and downsampled respectively. This process can be expressed by the formula: Among them, i is the i-th level feature, Down is the downsampling operation, and Up is the 2-fold bilinear interpolation upsampling operation. and They represent the output features of the previous and next branches of the current branch after alignment through the CA channel, and Respectively represent the output features after downsampling and upsampling operations; then, the features of the previous and next branches are concatenated with the features of the current branch to complete cross-fusion, and the cross-fused features of each level are subjected to 3×3 convolution operations to further mine fine-grained information; finally, the coordinate attention mechanism with spatial position perception is used to capture long-distance dependencies in space. This process can be expressed as: in, Indicates the characteristics of the current branch, They represent the feature maps of the previous layer and the next layer respectively. Concat is the feature concatenation operation. CBR is the 3×3 convolution, BN and Relu operations. A represents the coordinate attention mechanism. CF i It represents the result of cross-fusion of features at different levels.
6. A landscape pattern classification method combining global and local multi-order deep and shallow features according to claim 5, characterized in that: In step 3, a global-local attention module is introduced in the decoder part of the model. The global-local attention module consists of a local branch composed of multi-scale convolutions and a global branch of a window-based multi-head self-attention mechanism, which respectively extracts local detail information and global spatial dependency semantic features. The local branch uses convolution groups with convolution kernel sizes of 5×5, 3×3, and 1×1 to capture spatial and spectral features of different scales, thereby fully extracting local fine-grained features. The formula is: Among them, X L Represents the output features of the local branch; In the global branch, we first use a 1×1 convolution to transform the input feature map X i ∈R (B*C*H*W) The dimension C is expanded to 3 times. In order to reduce the computational complexity brought by the self-attention module, the input features are window-split and the feature map is divided into several non-overlapping windows. The self-attention mechanism is executed in each window to reduce the overall computational overhead and retain the global context information. In the window, the input sequence They are mapped to Query (Q), Key (K), and Value (V) vectors respectively, which can be expressed as: Q=XW Q ,K=XW K ,V=XW V (9) Where X is the representation of the input sequence, W Q , W K , W V It is a learnable full-weight matrix. The attention weight is obtained by calculating the dot product between Q and K. The dot product result is normalized using the Softmax function to generate the attention weight. The attention weight is multiplied by V to obtain each attention score, which can be expressed as: Among them, d k is the channel dimension of the K vector, Is a scaling factor used to alleviate the gradient vanishing problem. The Q, K, V vectors are divided into h heads, and the self-attention of each head is calculated separately, which can be expressed as: Among them, Q i , K i , V i It is the projection of Q, K, V on the i-th attention head. The results of all attention heads are concatenated and the final multi-head self-attention output is obtained through a linear transformation, which can be expressed as: Multihead(Q,K,V)=Concat(head1,head2,...,head h )W0 Among them, W0 is the linear transformation matrix of the science department.
7. A landscape pattern classification method combining global and local multi-order deep and shallow features according to claim 6, characterized in that: By fusing the global context information obtained by the global branch and the local features extracted by the local branch, a comprehensive expression of multi-scale information is achieved, which can be expressed as: Among them, X GL Represents the result of the fusion of global branches and local branches; The global-local attention module applies residual connection internally, and its process can be expressed as: X t =GLA(BN(X t-1 ))+X t-1 (13) X t+1 =MLP(BN(X t ))+X t (14) Among them, GLA represents the global-local attention part composed of global and local branches, X t-1 represents the input features of GLA and BN, X t+1 Represents the output features of MLP and BN.
8. A landscape pattern classification method combining global and local multi-order deep and shallow features according to claim 7, characterized in that: In step 4, the accuracy of the classification results of the method of the present invention is verified using the three evaluation indicators of F1, OA, and IoU. The specific formula is as follows: Where TP, FP, FN and TN represent the number of true positive predictions, false positive predictions, false negative predictions and true negative predictions, respectively.
Citation Information
Patent Citations
Semantic segmentation method and device for medical image
CN114972756A
SAR image ship target detection method based on Transform
CN115565066A
Retinal vessel segmentation method fusing swing-transfer and channel attention
CN118379496A
Remote sensing image semantic segmentation method, device and system, and storage medium
CN118470327A
Lightweight sand dune form type identification model based on partial convolution and Transform
CN119295890A
Cited By
Image analysis early warning method and system based on visual large model
CN120655982A
Image analysis and early warning method and system based on visual large model
CN120655982B