Urban scene classification method based on high-resolution multi-view remote sensing image

By using high-resolution multi-view remote sensing images and MSVTNet models in urban scene classification, the problem of unclear expression of stereo attributes and features is solved, and higher classification accuracy is achieved.

CN119942316AActive Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510034669.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The existing urban market scene classification method does not fully consider the three-dimensional attributes, and the local-global feature expression of the image is not clear, resulting in low classification accuracy.

Method used

High-resolution multi-view remote sensing image, combined with multi-scale pyramid attention and visual Transformer, MSVTNet model is established, and multi-view, multi-spectral and multi-angle stereo features are extracted and fused through adaptive selectors and feature fusion modules.

Benefits of technology

It significantly improves the LCZ classification accuracy, effectively utilizes multi-view angle and multi-spectral information, and enhances the network's ability to capture three-dimensional structure and global information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942316A_ABST
    Figure CN119942316A_ABST
Patent Text Reader

Abstract

The invention relates to an urban scene classification method based on a high-resolution multi-view remote sensing image, and belongs to the technical field of intelligent extraction of urban scene information. The method comprises the following steps: acquiring satellite remote sensing image data including multispectral and multi-view images, and preprocessing the data; carrying out pixel value extraction on the multi-view image to construct multi-angle three-dimensional features; establishing an LCZ sample set for a local climate region LCZ scene classification system; an MSVTNet city scene classification model fusing the attention and vision Transform of the multi-scale pyramid is established; and training the city scene classification model, and performing LCZ scene classification on the real-time remote sensing image after the training is completed. According to the method, the height difference change and the vertical structure feature of the ground object are extracted by using the multi-view image, the plane feature can be effectively supplemented, the proposed network model can effectively extract and fuse the global-local feature of the image, and the scene classification precision is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of intelligent extraction of urban scene information and relates to an urban scene classification method based on high-resolution multi-view remote sensing images. Background Art

[0002] Remote sensing image scene classification is to conduct an overall analysis of all land objects in a certain area in the image, combine the natural attributes and social function characteristics of the main land objects in the area, assign semantic labels, and achieve high-level semantic understanding of remote sensing images. Among them, Local Climate Zones (LCZ) is a detailed scenario classification framework proposed for studying urban climate-related issues. It divides the surface into 17 categories, including LCZ1 (compact high-rise building zone), LCZ2 (compact medium-rise building zone), LCZ3 (compact low-rise building zone), LCZ4 (open high-rise building zone), LCZ5 (open medium-rise building zone), LCZ6 (open low-rise building zone), LCZ7 (lightweight building zone), LCZ8 (large low-rise building zone), LCZ9 (very open low-rise building zone), LCZ10 (heavy industrial zone), LCZA (dense tree zone), LCZB (sparse tree zone), LCZC (shrub zone), LCZD (low vegetation zone), LCZE (hardened ground zone), LCZF (bare sand zone), and LCZG (water zone), according to building height and density, urban functions, and natural cover. It has good application prospects in urban planning and management, disaster risk assessment, air quality and environmental monitoring, and climate change research.

[0003] In LCZ classification, medium-resolution satellite images (such as Landsat and Sentinel-2) are often used. Although they have advantages in spatial coverage and temporal frequency, they only provide single-view images due to their limited spatial resolution, making it difficult to capture the three-dimensional structural characteristics of the objects. This limitation makes it difficult to distinguish LCZ types, especially when distinguishing different building types that are carefully divided according to building height and density. In complex urban scenes, it is still a great challenge to carry out detailed, accurate and large-scale LCZ mapping. At present, there is no generally accepted solution for how to effectively improve the classification accuracy in such scenarios.

[0004] Existing remote sensing image scene classification methods can be divided into methods based on artificially designed features and methods based on deep features. Methods based on artificially designed features mainly rely on extracting features such as color, texture and shape from images to provide decision support for classification. These methods perform well when processing images with similar texture structures or simple scenes, but when facing complex scenes, the representation ability of artificially designed features is often limited, affecting the classification accuracy. To overcome these shortcomings, methods based on deep learning have gradually become mainstream. Deep learning models can more effectively capture multidimensional information in complex scenes by automatically extracting high-level features. This automated feature extraction method breaks through the limitations of traditional methods and opens up a new direction for remote sensing image scene classification. In recent years, various deep learning methods have become research hotspots. Methods based on transfer learning mainly apply models trained on public datasets (such as ImageNet) directly or indirectly to other image classification tasks, but for complex scenes, it is difficult to learn and apply features in other data. The method of using attention mechanism to enhance features embeds the attention mechanism into the existing network architecture, strengthens the local semantic information of remote sensing images, and enhances the ability of the network to extract fine-grained features, but many studies lack the acquisition and utilization of global information. The scene classification method based on graph convolutional neural network establishes the spatial context dependency of different targets in remote sensing images and improves the adaptability of the network to the target, but it is easy to ignore the detailed representation of local information. With the advent of Transformer, the natural image recognition ability demonstrated by visual Transformer has great potential in remote sensing image scene classification.

[0005] Current remote sensing image scene classification mainly relies on planar features extracted from single-view images, ignoring the important role of three-dimensional spatial attributes in urban information extraction; in each unique scene, there is also a lack of detailed feature description of local images, and the global semantic connection is also insufficient, making it difficult to provide effective classification capabilities.

[0006] The three-dimensional features provided by multi-view images can supplement the two-dimensional plane information obtained from a single view, thereby further improving the classification performance. The present invention aims to solve the problems that the existing methods do not fully consider the three-dimensional properties, the local-global features of the images are not clearly expressed, and the scene classification accuracy is not high. Summary of the invention

[0007] In view of this, the purpose of the present invention is to invent an efficient and high-precision scene classification method using high-resolution multi-view remote sensing images to address the problems of lack of consideration of three-dimensional information, unclear expression of local-global features of images, and low model classification accuracy in current urban scene classification. The method is applied to the LCZ classification task to verify the effectiveness of the proposed method.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] A method for classifying urban scenes based on high-resolution multi-view remote sensing images, the method comprising the following steps:

[0010] S1. Obtain satellite remote sensing image data, including multi-spectral and multi-view images, and pre-process the data;

[0011] S2, extracting pixel values ​​from multi-view images to construct multi-angle stereo features;

[0012] S3, establishing a LCZ classification sample set for the preprocessed multispectral and multi-view images;

[0013] S4. Establish the urban scene classification model MSVTNet which integrates multi-scale pyramid attention and visual Transformer;

[0014] S5. Train the urban scene classification model, and use it for scene classification of real-time remote sensing images after training.

[0015] Wherein, S4 includes the following sub-steps:

[0016] S41, establishing a multi-spectral image branch, a multi-view image branch and a multi-angle feature branch, which are used to process the multi-spectral image, the multi-view image and the multi-angle feature respectively;

[0017] S42, the image input of each branch is processed by importance weighting through an adaptive selector;

[0018] S43, inputting the importance-weighted features into the convolution branch and the visual Transformer branch for local feature extraction and global feature extraction respectively;

[0019] S44, using a feature fusion module to dynamically fuse the local features and global features obtained from the two branches using a dual-branch bidirectional fusion strategy;

[0020] S45. Combine the local branch feature and the global branch feature results to obtain a final classification value.

[0021] Furthermore, in step S1, the acquired multi-view images refer to images including front view, front view and rear view; the preprocessing operations performed on the acquired remote sensing image data include: orthorectification, image registration and image fusion.

[0022] Further, in step S2, the extracted pixel values ​​include: difference multi-angle index DMAI, ratio multi-angle index RMAI, normalized multi-angle index NMAI, maximum normalized multi-angle index MMNAI, and their calculation formulas are:

[0023]

[0024] Among them, max / min / mean / var represent the maximum value, minimum value, mean value, and variance respectively, and x n Indicates the pixel value of the n-th viewing angle image.

[0025] Further, in step S3, with the help of Google Earth and Amap building data, samples are selected through visual interpretation and manual annotation, and field sampling is performed to construct a LCZ classification sample set. The constructed LCZ classification sample set is: In step S3, with the help of Google Earth and Amap building data, samples are selected through visual interpretation and manual annotation, and field sampling is performed to construct a LCZ classification sample set. The constructed LCZ classification sample set is: compact high-rise building area LCZ1, compact medium-rise building area LCZ2, compact low-rise building area LCZ3, open high-rise building area LCZ4, open medium-rise building area LCZ5, open low-rise building area LCZ6, light building area LCZ7, large low-rise building area LCZ8, very open low-rise building area LCZ9, heavy industrial area LCZ10, dense tree area LCZA, sparse tree area LCZB, bush area LCZC, low vegetation area LCZD, hardened ground area LCZE, bare sand area LCZF, and water area LCZG.

[0026] Further, in step S42, the adaptation selector uses the torch.nn.Parameter() method to add learnable adaptive weights to the three sets of inputs, and the process is as follows:

[0027] Initialize the weights of the three branches according to their contribution and effectiveness to the classification of remote sensing image scenes;

[0028] The initial weights are trained and optimized. During the training process, the Softmax function is used to normalize the weights. The specific calculation formula is as follows:

[0029]

[0030] where a i represents the initial weight of the i-th branch, w i is the normalized weight of the i-th branch;

[0031] The image of each branch is multiplied by the corresponding weight to obtain a weighted input image with assigned importance, which is then passed to different feature extraction branches for subsequent operations.

[0032] Further, in the convolution branch in step S43, the input features are convolved and pooled, and the calculation formula is:

[0033] x=MaxPool(ReLU(BN(Conv(x))))

[0034] Where x is the input feature, Conv(·) is the convolution operation, BN(·) is batch normalization, ReLU(·) is the activation function, and MaxPool(·) is the maximum pooling operation;

[0035] Then the convolutional pooled features are input into the Bottleneck module and the MPAM module for multi-scale feature extraction, where:

[0036] In the Bottleneck module, the image features first enter a 3x3 convolutional layer Conv, which applies a filter to the input data by sliding a window;

[0037] Then enter the batch normalization layer BN to normalize the output of the previous layer; then pass a ReLU activation function to set all negative values ​​to 0 and keep the positive values ​​unchanged;

[0038] Next, a 3x3 convolutional layer is used to further extract features, and batch normalization is performed again. At the same time, a 1x1 convolutional layer is used to transform features without changing the spatial dimension, and batch normalization is performed again;

[0039] Then, an addition operation is performed to add the input to the output after the above convolution and normalization, and finally the result is passed through a ReLU activation function and then output;

[0040] In the MPAM module, the features output by the Bottleneck module are first subjected to an adaptive average pooling operation, which calculates the average value by dynamically dividing the region;

[0041] Then the features obtained in each step are interpolated, and Interpolate bilinear interpolation is used to align feature maps of different sizes so that they can be fused at the same scale. The specific calculation formula is as follows:

[0042] I(x,y)=(1-a)(1-b)I(i,j)+a(1-b)I(i+1,j)+(1-a)bI(i,j+1)+abI(i+1,j+1)

[0043] Where (i, j) is the upper left pixel in the input feature map; a = xi, b = yj is the relative position offset; I (x, y) is the interpolated pixel value;

[0044] Add Fourier denoising operation to the first two pooling results, that is, perform image-frequency domain conversion, filter high-pass noise, convert its matrix domain into frequency domain, and use Gaussian filter to find the weight matrix of each dimension of the feature block:

[0045]

[0046] σ is the filter kernel, f(x,y) represents the two-dimensional Gaussian function on the image, f X (x), f Y (y) represents the one-dimensional Gaussian function in the x and y directions respectively; the weight matrix obtained by the two-dimensional Gaussian function in the above formula is summed and averaged, and the low-frequency components are filtered out and then inversely converted into the matrix domain output;

[0047] Then all feature results, i.e., multi-scale features, will be concatenated;

[0048] Finally, the spliced ​​multi-scale feature results are passed through the channel attention mechanism CAM. The CAM module consists of multiple Conv, normalized BN, linear activation ReLU and Sigmod operations. The specific calculation formula is as follows:

[0049] Out = Sigmod(ReLU(BN(Conv2d(x)))...)

[0050] Finally, the Out result is output to the external FC layer and introduced as the bottleneck layer.

[0051] Furthermore, in the visual Transformer branch in step S43, each input image block is subjected to word embedding and position encoding operations to obtain an E vector, which is calculated as follows:

[0052]

[0053] Among them, e i represents the embedding of the i-th image block, PE represents the position encoding of the same dimension as the input sequence, F is the input feature map, p represents the Patch size, H and W represent the height and width of the entire image respectively.

[0054] Then E is input into the Transformer Encoder module to obtain global features. The Transformer Encoder module includes normalization Norm, multi-head attention mechanism Muti-Head Attention and multi-layer perceptron MLP. Among them, Norm normalizes all input features;

[0055] The multi-head attention mechanism calculates the relevance of each patch in the entire sequence through the self-attention mechanism. The calculation formula is as follows:

[0056]

[0057] Among them, Q, K, V represent the query vector, key vector and value vector generated by the input features through different projection matrices, T represents the transposition operation, and d k represents the dimension of the key vector;

[0058] The multi-layer perceptron MLP is implemented through a series of linear, ReLU and Dropout operations;

[0059] Finally, the result of the multi-head attention mechanism is passed through the MLP_Head layer, that is, two linear layers Linear to obtain the final result cls_token_final.

[0060] Further, in step S44, the feature fusion module includes a self-attention feature fusion mechanism module SACA, a convolution branch local feature processing module, and a ViT branch global feature processing module, wherein:

[0061] The SACA module includes a global average pooling module GAP, a soft selection mechanism, and a dynamic weight generation module DWG. First, GAP compresses each feature to obtain a feature summary:

[0062] G i =GAP(X i )

[0063] Among them, X i Represents the characteristics of the i-th channel of the two branches;

[0064] Then enter the Soft selection mechanism, which uses FC operation, including multiple Linear layers, ReLU layers, and Sigmod layers. The implementation formula is as follows:

[0065] S i =Sigmod(ReLU(Linear(G i ))...)

[0066] Then, dynamic weight generation is performed through the DWG module. The specific operation formula is as follows:

[0067] w i =Sigmod(MLP(S i ))

[0068] The final dynamic weight w i Weighted with the input features to get the result X SACA :

[0069]

[0070] By dynamically obtaining the different contributions of the two branch features, the result generation of the dual-branch fusion strategy is assisted;

[0071] In the convolution branch local feature processing module, the two-dimensional shape of the convolution branch feature is first retained, and Conv, Flatten stretching and fusion operations are performed. The results are calculated as follows:

[0072] X Mi =F M (X M(i-1) )+SACA V-M (X V(i-1) )

[0073] Where X M and X V represents the convolution branch features and ViT branch features, F M Represents the convolution feature branch Conv, Flatten operation, SACA represents the feature fusion attention mechanism;

[0074] In the ViT branch global feature processing module, the residual fully connected layer Linear operation is performed on the ViT branch features, and the result is calculated as follows:

[0075] X Vi =X V(i-1) +F V (X V(i-1) )+SACA M-V (F M (X M(i-1) ))

[0076] Where X M and X V represents the convolution branch features and ViT branch features, F M Represents the convolution branch feature Conv, Flatten operation, F V It represents the linear operation of the residual fully connected layer of the ViT branch feature, and SACA represents the feature fusion attention mechanism.

[0077] Further, in step S45, the fusion result of local branch features and global branch features is calculated as follows:

[0078] X Out =Conv(X Mi +X Vi )

[0079] Finally, the fusion result X Out The fully connected layer FC and Softmax output are performed to obtain the final classification value. Among them, the fully connected layer FC maps the output feature vector to the category space, assists in obtaining the category probability distribution of the subsequent Softmax function, and takes the type with the maximum probability to indicate the LCZ category to which the image belongs.

[0080] Further, in step S5, the network parameters for training the urban scene classification model include configuration parameters, loss function, and accuracy evaluation index, wherein the configuration parameters include batch size, learning rate, and number of iterations;

[0081] The loss function uses the cross entropy loss function L CE :

[0082]

[0083] Where N is the number of samples, M is the number of classification categories; y ij Indicates whether the category is the same as the sample i category, if they are the same, it is 1, if they are different, it is 0; p ij It represents the predicted probability that the observed sample i belongs to category j;

[0084] The accuracy evaluation indicators include overall accuracy OA, mapping accuracy PA, user accuracy UA, and harmonic average score F1-Score.

[0085] The beneficial effects of the present invention are:

[0086] (1) Based on high-resolution multi-view remote sensing images, this paper proposes a deep learning network MSVTNet that integrates multi-scale pyramid attention (MPAM) and visual transformer (ViT), which fully exploits the plane and three-dimensional spatial information for urban scene classification. Experimental results show that this method uses multi-view images to characterize the height difference changes and vertical structural characteristics of objects, which can effectively supplement the plane features, enhance the inter-class differences, and significantly improve the LCZ classification accuracy, thereby proving the effectiveness of the proposed method for remote sensing image scene classification. In particular, by introducing an adaptive selector, the network can adaptively adjust the weights, thereby more effectively utilizing multi-view, multi-spectral images and multi-angle stereo features.

[0087] (2) The design of the multi-scale pyramid attention module (MPAM) in the present invention improves the network's ability to capture image features in local details. Adaptive average pooling allows the network to dynamically adjust the size of the output feature map according to the size of the input data, increasing the flexibility of the model. At the same time, Fourier transform is introduced to achieve high-precision noise suppression. A channel attention mechanism is added to the final fused features to filter out irrelevant or noise features, allowing the network to focus more on key features and increase inter-class differences.

[0088] (3) The design of the Visual Transformer (ViT) module in this invention improves the network’s ability to capture image features in a global structure. Based on the representation of image blocks, it can effectively capture long-distance dependencies in the image. At the same time, the Transformer has a built-in dense self-attention mechanism that continuously focuses on the most relevant parts of the image, improving the model’s ability to capture contextual relationships and global information.

[0089] (4) The dual-branch bidirectional fusion strategy in the present invention uses the self-attention feature fusion mechanism to dynamically enhance the image category-related information, suppress irrelevant information, and enhance the model's ability to distinguish categories. At the same time, different processing methods are used for the convolution branch and the ViT branch to enrich the feature diversity and provide sufficient features for subsequent classification.

[0090] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0092] Figure 1 It is a schematic diagram of the structure of the urban scene classification model MSVTNet that integrates multi-scale pyramid attention and visual Transformer of the present invention;

[0093] Figure 2 Schematic diagram of the structure of the Bottleneck module in the convolution branch of the MSVTNet model of the present invention;

[0094] Figure 3 It is a schematic diagram of the structure of the MPAM module in the convolution branch of the MSVTNet model of the present invention;

[0095] Figure 4 It is a schematic diagram of the structure of the visual Transformer branch of the MSVTNet model of the present invention;

[0096] Figure 5 It is a schematic diagram of the structure of the Transformer Encoder module in the visual Transformer branch of the MSVTNet model of the present invention;

[0097] Figure 6 Schematic diagram of the structure of the MLP module in the visual Transformer branch of the MSVTNet model of the present invention;

[0098] Figure 7 It is a structural schematic diagram of the feature fusion module of the MSVTNet model of the present invention;

[0099] Figure 8 It is a schematic diagram of the experimental flow under the embodiment;

[0100] Fig. 9 It is a schematic diagram of the experimental results under the embodiment, wherein, Fig. 9 (a) is the real image, Fig. 9 (b) is the classification result. DETAILED DESCRIPTION

[0101] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0102] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0103] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0104] See also Figures 1 to 9, is a method for urban scene classification based on high-resolution multi-view remote sensing images. To address the problems in current urban scene classification, such as lack of consideration of three-dimensional information, unclear expression of local-global features of images, and low model classification accuracy, an efficient and high-precision scene classification method is proposed using high-resolution multi-view remote sensing images. The method is applied to the LCZ classification task to verify the effectiveness of the proposed method.

[0105] Example

[0106] This embodiment provides a detailed implementation process of the urban scene classification method based on high-resolution multi-view remote sensing images of the present invention, which includes the following steps:

[0107] S1. Acquire satellite remote sensing image data, including multi-spectral and multi-view images, and pre-process the data.

[0108] This embodiment mainly obtains remote sensing image data from the Ziyuan-3 satellite. Ziyuan-3 (ZY-3) is one of China's high-resolution stereo mapping satellite series. Its main task is to conduct land and resources surveys, monitoring and management, and to provide updated data for the national basic geographic information database. The multi-view images obtained refer to images of perspectives including front view, front view and rear view. Preprocessing operations performed on the acquired remote sensing image data include: orthorectification, image registration and image fusion.

[0109] S2, extracting pixel values ​​from multi-view images to construct multi-angle stereo features;

[0110] Specifically, the extracted pixel values ​​include: Difference Multi-Angular Index (DMAI), Ratio Multi-Angular Index (RMAI), Normalized Multi-Angular Index (NMAI), and Maximum Normalized Multi-Angular Index (MNMAI), and their calculation formulas are:

[0111]

[0112] Among them, max / min / mean / var represent the maximum value, minimum value, mean value, and variance respectively, and x n Indicates the pixel value of the n-th viewing angle image. For the ZY-3 dataset, it includes the front view, rear view, and front view images.

[0113] S3. Establish an LCZ classification sample set for the preprocessed multispectral and multi-view images.

[0114] Specifically, this embodiment uses Google Earth and Amap building data, selects samples through visual interpretation and manual annotation, and conducts field sampling to ensure sample reliability to construct an LCZ classification sample set. The constructed LCZ classification sample set is: LCZ1 (compact high-rise building area), LCZ2 (compact medium-rise building area), LCZ3 (compact low-rise building area), LCZ4 (open high-rise building area), LCZ5 (open medium-rise building area), LCZ6 (open low-rise building area), LCZ7 (lightweight building area), LCZ8 (large low-rise building area), LCZ9 (very open low-rise building area), LCZ10 (heavy industrial area), LCZA (dense tree area), LCZB (sparse tree area), LCZC (shrub area), LCZD (low vegetation area), LCZE (hardened ground area), LCZF (bare sand area), LCZG (water area), a total of 17 categories.

[0115] S4. Establish the urban scene classification model MSVTNet that integrates multi-scale pyramid attention and visual Transformer.

[0116] Specifically, if Figure 1 The overall architecture of the MSVTNet model shown in Figure 1 integrates the Multi-scale Pyramid Attention Module (MPAM) and the Vision Transformer (ViT) module. The specific processing process includes the following steps:

[0117] S41. First, a multispectral image branch, a multi-view image branch, and a multi-angle feature branch are established, which are used to process multispectral images, multi-view images, and multi-angle features, respectively. The multispectral image branch contains data from a multispectral sensor, which is used to extract image spectral information. The multi-view image branch contains images taken from different perspectives, providing a variety of spatial coverage. The multi-angle stereo feature branch contains angle difference change information extracted from remote sensing images of different perspectives.

[0118] S42. The image input of each branch is first processed by the adaptive selector for importance weighting; the adaptive selector uses the torch.nn.Parameter() method to add learnable adaptive weights to the three groups of inputs to distribute information of different importance to the convolution and ViT branches respectively. Previous studies have shown that spectral features are of indispensable importance in scene classification. At the same time, with the help of high-resolution multi-view remote sensing images, multi-angle stereo features are extracted, which has a significant effect on LCZ classification. Therefore, in the classification of remote sensing image scenes, the initial weights are divided according to the ratio of [0.3, 0.3, 0.4] according to the contribution and effectiveness of the classification, and the weights will be optimized in training.

[0119] During the training process, the Softmax function is used to normalize the weights to ensure that the sum is 1, to avoid imbalance caused by excessive weight of a branch. The specific calculation formula is as follows:

[0120]

[0121] where a i represents the initial weight of the i-th branch, w i is the normalized weight.

[0122] The image of each branch is multiplied by the corresponding weight to obtain a weighted input image with assigned importance, which is then passed to different feature extraction branches for subsequent operations.

[0123] S43. The importance-weighted features are respectively input into the convolution branch and the visual Transformer branch for local feature extraction and global feature extraction.

[0124] Specifically, in the convolution branch, the input features are first subjected to a 7-core convolution plus pooling operation. The specific calculation formula is:

[0125] x=MaxPool(ReLU(BN(Conv(x))))

[0126] Where x is the input feature, Conv(·) is the convolution operation, BN(·) is batch normalization, ReLU(·) is the activation function, and MaxPool(·) is the maximum pooling operation.

[0127] Then enter the Bottleneck+MPAM module for multi-scale feature extraction. The combination of the two, Bottleneck helps to reduce the amount of extraction calculations, and MPAM improves the quality of extracted features, ensuring the efficient extraction of multi-scale features.

[0128] Figure 2The structure diagram of the Bottleneck module is shown. Bottleneck is a common network module in deep learning. By combining with ResNet, a Bottleneck module with residual is generated. In this module, the image enters a 3x3 convolution layer (Conv), applies a filter to the input data by sliding a window, and enters a batch normalization layer (BatchNormalization, BN) to normalize the output of the previous layer, making the output distribution more stable. Then there is a ReLU activation function, which sets all negative values ​​to 0 and keeps positive values ​​unchanged, increasing the nonlinearity of the network to learn more complex features. Next is another 3x3 convolution layer to further extract features and perform batch normalization again. At the same time, a 1x1 convolution layer is used to transform features without changing the spatial dimension, and batch normalization is performed again. Then, an addition operation is performed to add the input (before the 1x1 convolution) to the output after the above convolution and normalization. Finally, the result is obtained through a ReLU activation function to add nonlinearity to the network.

[0129] Figure 3 The schematic diagram of the MPAM module is shown. First, the AdaptiveAvgPool (adaptive average pooling) operation is performed. Unlike the AvgPool (ordinary average pooling) operation, the output size of AdaptiveAvgPool is fixed, that is, the user specifies the target size, and the pooling window size and step size are automatically calculated and dynamically adapted. And unlike AvgPool, which calculates the average value within a fixed window, AdaptiveAvgPool calculates the average value by dynamically dividing the area.

[0130] After that, the features obtained in each step are interpolated, and Interpolate bilinear interpolation is used to align feature maps of different sizes so that they can be fused at the same scale. The specific calculation formula is as follows:

[0131] I(x,y)=(1-a)(1-b)I(i,j)+a(1-b)I(i+1,j)+(1-a)bI(i,j+1)+abI(i+1,j+1)

[0132] Among them, (i, j) is the upper left corner pixel in the input feature map; a=xi, b=yj are the relative position offsets; I(x, y) is the interpolated pixel value.

[0133] In addition, the Fourier denoising (D_FFT) operation is added to the first two pooling results, that is, image-frequency domain conversion is performed, high-pass noise is filtered, its matrix domain is converted to frequency domain, and the weight matrix of each dimension of the feature block is obtained using a Gaussian filter:

[0134]

[0135] σ is the filter kernel, f(x,y) represents the two-dimensional Gaussian function on the image, f X (x), f Y (y) represents the one-dimensional Gaussian function in the x and y directions respectively. In many implementations, the two-dimensional Gaussian filter is decomposed into two one-dimensional Gaussian filters. The weight matrix obtained by the two-dimensional Gaussian function in the above formula is summed and averaged, and the low-frequency components are filtered out and then inversely converted to the matrix domain output.

[0136] Then all feature results, i.e. multi-scale features, will be concatenated. Finally, the concatenated multi-scale feature results will be passed through the channel attention mechanism CAM. The CAM module consists of multiple Conv, normalized BN, linear activation ReLU and Sigmod operations. The specific calculation formula is as follows:

[0137] Out = Sigmod(ReLU(BN(Conv(x)))...)

[0138] Finally, the Out result is output to the external FC layer in order to introduce it as a bottleneck layer. This not only reduces the number of parameters, but also forces the model to learn more compact and representative feature representations.

[0139] Then, in the visual Transformer branch, the structural diagram of the ViT branch module is as follows: Figure 4 As shown in the figure, each input image block is subjected to word embedding and position encoding operations to obtain the E vector, so that the subsequent Transformer processing can capture the semantic relationship and contextual information between words. The specific calculation formula is as follows:

[0140]

[0141] Among them, e i represents the embedding of the i-th image block, PE represents the position encoding of the same dimension as the input sequence, F is the input feature map, p represents the Patch size, H and W represent the height and width of the entire image respectively.

[0142] Then put E into Transformer Encoder to obtain global features, such as Figure 5 The figure shows the structure of the Transformer Encoder, which is mainly composed of normalization Norm, multi-head attention mechanism Muti-Head Attention and multi-layer perceptron MLP. Among them, Norm normalizes all input features.

[0143] Muti-Head Attention, that is, through the self-attention mechanism, calculates the relevance of each patch in the entire sequence and captures the long-distance dependencies between feature blocks. The specific calculation formula is as follows:

[0144]

[0145] Where Q, K, and V represent the query vector, key vector, and value vector generated by the input features through different projection matrices. T represents the transposition operation, d k Indicates the dimension of the key vector, used for scaling.

[0146] For the multi-layer perceptron MLP, its structure is as follows Figure 6 As shown in the figure, this module is implemented through a series of linear, ReLU and Dropout operations. Finally, the result of the multi-head attention mechanism is passed through the MLP_Head layer, that is, two linear layers to obtain the final result cls_token_final.

[0147] S44, using a feature fusion module to dynamically fuse the local features and global features obtained from the two branches using a dual-branch bidirectional fusion strategy.

[0148] The structural diagram of the feature fusion module is shown in Figure 7, which includes a self-attention feature fusion mechanism (Self-Attention Convolutional Adaptive Fusion, SACA) module, a convolution branch local feature processing module, and a ViT branch global feature processing module. Among them, the SACA module includes a global average pooling (Global Average Pooling, GAP) module, a Soft selection mechanism, and a dynamic weight generation (Dynamic Weight Generation, DWG) module.

[0149] First, GAP compresses each feature to obtain a feature summary:

[0150] G i =GAP(X i )

[0151] Among them, X i Represents the features of the i-th channel of the two branches.

[0152] Then enter the Soft Selection mechanism, which allows the model to select or combine different components or paths in a smooth manner, by assigning weights to reflect the importance of different options. When applied to feature channels, it can dynamically emphasize certain channels and suppress other channels, thereby improving the expressiveness of the model.

[0153] The soft selection mechanism used in this embodiment is FC operation, which includes multiple Linear layers, ReLU layers, and Sigmod layers. The specific implementation formula is as follows:

[0154] S i =Sigmod(ReLU(Linear(G i ))...)

[0155] Then, dynamic weight generation is performed through the DWG module. The specific operation formula is as follows:

[0156] w i =Sigmod(MLP(S i ))

[0157] The final dynamic weight w i Weighted with the input features to get the result X SACA :

[0158]

[0159] This method can dynamically obtain the different contributions of the two branch features, and its addition can more effectively assist the result generation of the dual-branch fusion strategy.

[0160] The convolution branch local feature processing module first retains the two-dimensional shape of the convolution branch feature, performs Conv, Flatten stretching and fusion operations, and the result is calculated as follows:

[0161] X Mi =F M (X M(i-1) )+SACA V-M (X V(i-1) )

[0162] Where X M and X V represents the convolution branch features and ViT branch features, F M Represents the convolution feature branch Conv, Flatten operation, and SACA represents the feature fusion attention mechanism.

[0163] The ViT branch global feature processing module performs a residual fully connected layer Linear operation on the ViT branch features, and the result is calculated as follows:

[0164] X Vi =X V(i-1) +F V (X V(i-1) )+SACA M-V (F M (X M(i-1) ))

[0165] Where XM and X V represents the convolution branch features and ViT branch features, F M Represents the convolution branch feature Conv, Flatten operation, F V It represents the linear operation of the residual fully connected layer of the ViT branch feature, and SACA represents the feature fusion attention mechanism.

[0166] S45, then combine the local branch feature and the global branch feature results to obtain the final classification value; wherein the fusion result is calculated as follows:

[0167] X Out =Conv(X Mi +X Vi )

[0168] Finally, the fusion result X Out The fully connected layer FC and Softmax output are performed to obtain the final classification value. The role of FC is to map the feature vector outputted previously to the category space, assist in obtaining the category probability distribution of the subsequent Softmax function, and take the type with the maximum probability to indicate the LCZ category to which the image belongs.

[0169] S5. Train the urban scene classification model MSVTNet, and after training, use it for scene classification of real-time remote sensing images.

[0170] Network parameter settings include various parameters, loss functions, and accuracy evaluation indicators that need to be adjusted and configured when training a neural network.

[0171] Various parameters include: batch size batch_size is 32, cross entropy loss function is used, Adam optimizer is used on the dataset, and learning rate is 2*10 -4 , the number of iterations is 300 rounds.

[0172] The loss function uses the cross entropy loss function L CE (Cross-Entropy Loss), the calculation formula is as follows:

[0173]

[0174] Where N is the number of samples, M is the number of classification categories; y ij Indicates whether the category is the same as the sample i category, if they are the same, it is 1, if they are different, it is 0; p ij It represents the predicted probability that the observed sample i belongs to category j.

[0175] The accuracy evaluation includes calculating the accuracy evaluation index by comparing the consistency between the prediction results of the model and the actual value through the confusion matrix. The accuracy evaluation index includes overall accuracy (OA), mapping accuracy (PA), user accuracy (UA), and harmonic mean score (F1-Score). The specific calculation formula is as follows:

[0176]

[0177] Where P is the total number of samples, P ii , P ij , P ji is the value of the i-th row and j-th column of the confusion matrix.

[0178] In this embodiment, a classification model is trained under the parameter conditions with the best accuracy evaluation results, and the LCZ classification results are predicted for the images to obtain the final urban LCZ classification mapping results.

[0179] This example is based on the LCZ classification dataset of Resource No. 3, which contains 16,309 sample images of 125*125m and covers five typical cities in China, and experiments are conducted on the above process. The experimental process is as follows: Figure 8 As shown, we get Fig. 9 An example of LCZ classification results is shown, where Fig. 9 (a) is the real image, Fig. 9 (b) is the classification result.

[0180] The highest OA (overall accuracy) obtained by existing studies on the Resource III dataset is 82.7%. Through the method proposed in the present invention, OA is increased to 86.06%, verifying the effectiveness of MSVTNet for LCZ classification.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A method for classifying urban scenes based on high-resolution multi-view remote sensing images, characterized by: The method comprises the following steps: S1. Obtain satellite remote sensing image data, including multi-spectral and multi-view images, and pre-process the data; S2, extracting pixel values ​​from multi-view images to construct multi-angle stereo features; S3, establishing a LCZ classification sample set for the preprocessed multispectral and multi-view images; S4. Establish the urban scene classification model MSVTNet which integrates multi-scale pyramid attention and visual Transformer; S5. Train the urban scene classification model, and use it for scene classification of real-time remote sensing images after training. Wherein, S4 includes the following sub-steps: S41, establishing a multi-spectral image branch, a multi-view image branch and a multi-angle feature branch, which are used to process the multi-spectral image, the multi-view image and the multi-angle feature respectively; S42, the image input of each branch is processed by importance weighting through an adaptive selector; S43, inputting the importance-weighted features into the convolution branch and the visual Transformer branch for local feature extraction and global feature extraction respectively; S44, using a feature fusion module to dynamically fuse the local features and global features obtained from the two branches using a dual-branch bidirectional fusion strategy; S45. Combine the local branch feature and the global branch feature results to obtain a final classification value.

2. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 1, characterized in that: In step S1, the acquired multi-view images refer to images including front view, front view and rear view; the pre-processing operations performed on the acquired remote sensing image data include: orthorectification, image registration and image fusion.

3. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 1 is characterized by: In step S2, the extracted pixel values ​​include: difference multi-angle index DMAI, ratio multi-angle index RMAI, normalized multi-angle index NMAI, maximum normalized multi-angle index MMNAI, and their calculation formulas are: DMAI=max / min / mean / var RMAI=max / min / mean / var NMAI=max / min / mean / var MNMAI=max / min / mean / var Among them, max / min / mean / var represent the maximum value, minimum value, mean value, and variance respectively, and x n Indicates the pixel value of the n-th viewing angle image.

4. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 1, characterized in that: In step S3, with the help of Google Earth and Amap building data, samples are selected through visual interpretation and manual annotation, and field sampling is carried out to construct the LCZ classification sample set. The constructed LCZ classification sample set is: compact high-rise building area LCZ1, compact medium-rise building area LCZ2, compact low-rise building area LCZ3, open high-rise building area LCZ4, open medium-rise building area LCZ5, open low-rise building area LCZ6, light building area LCZ7, large low-rise building area LCZ8, very open low-rise building area LCZ9, heavy industrial area LCZ10, dense tree area LCZA, sparse tree area LCZB, shrub area LCZC, low vegetation area LCZD, hardened ground area LCZE, bare sand area LCZF, and water area LCZG.

5. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 1, characterized in that: In step S42, the adaptive selector uses the torch.nn.Parameter() method to add learnable adaptive weights to the three sets of inputs. The process is as follows: Initialize the weights of the three branches according to their contribution and effectiveness to the classification of remote sensing image scenes; The initial weights are trained and optimized. During the training process, the Softmax function is used to normalize the weights. The specific calculation formula is as follows: where a i represents the initial weight of the i-th branch, w i is the normalized weight of the i-th branch; The image of each branch is multiplied by the corresponding weight to obtain a weighted input image with assigned importance, which is then passed to different feature extraction branches for subsequent operations.

6. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 5, characterized in that: In the convolution branch in step S43, the input features are convolved and pooled, and the calculation formula is: x=MaxPool(ReLU(BN(Conv(x)))) Where x is the input feature, Conv(·) is the convolution operation, BN(·) is batch normalization, ReLU(·) is the activation function, and MaxPool(·) is the maximum pooling operation; Then the convolutional pooled features are input into the Bottleneck module and the MPAM module for multi-scale feature extraction, where: In the Bottleneck module, the image features first enter a 3x3 convolutional layer Conv, which applies a filter to the input data by sliding a window; Then enter the batch normalization layer BN to normalize the output of the previous layer; then pass a ReLU activation function to set all negative values ​​to 0 and keep the positive values ​​unchanged; Next, a 3x3 convolutional layer is used to further extract features, and batch normalization is performed again. At the same time, a 1x1 convolutional layer is used to transform features without changing the spatial dimension, and batch normalization is performed again; Then, an addition operation is performed to add the input to the output after the above convolution and normalization, and finally the result is passed through a ReLU activation function and then output; In the MPAM module, the features output by the Bottleneck module are first subjected to an adaptive average pooling operation, which calculates the average value by dynamically dividing the region; Then the features obtained in each step are interpolated, and Interpolate bilinear interpolation is used to align feature maps of different sizes so that they can be fused at the same scale. The specific calculation formula is as follows: I(x,y)=(1-a)(1-b)I(i,j)+a(1-b)I(i+1,j)+(1-a)bI(i,j+1)+abI(i+1,j+1)where (i,j) is the upper left pixel in the input feature map; a=xi, b=yj are relative position offsets; I(x,y) is the interpolated pixel value; Add Fourier denoising operation to the first two pooling results, that is, perform image-frequency domain conversion, filter high-pass noise, convert its matrix domain into frequency domain, and use Gaussian filter to find the weight matrix of each dimension of the feature block: σ is the filter kernel, f(x,y) represents the two-dimensional Gaussian function on the image, f X (x), f Y (y) represents the one-dimensional Gaussian function in the x and y directions respectively; the weight matrix obtained by the two-dimensional Gaussian function in the above formula is summed and averaged, and the low-frequency components are filtered out and then inversely converted into the matrix domain output; Then all feature results, i.e., multi-scale features, will be concatenated; Finally, the spliced ​​multi-scale feature results are passed through the channel attention mechanism CAM. The CAM module consists of multiple Conv, normalized BN, linear activation ReLU and Sigmod operations. The specific calculation formula is as follows: Out = Sigmod(ReLU(BN(Conv(x)))...) Finally, the Out result is output to the external FC layer and introduced as the bottleneck layer.

7. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 6, characterized in that: In the visual Transformer branch in step S43, each input image block is subjected to word embedding and position encoding operations to obtain an E vector, which is calculated as follows: Among them, e i represents the embedding of the i-th image block, PE represents the position encoding of the same dimension as the input sequence, F is the input feature map, p represents the patch size, H and W represent the height and width of the entire image respectively; Then E is input into the Transformer Encoder module to obtain global features. The Transformer Encoder module includes normalization Norm, multi-head attention mechanism Muti-Head Attention and multi-layer perceptron MLP. Among them, Norm normalizes all input features; The multi-head attention mechanism calculates the relevance of each patch in the entire sequence through the self-attention mechanism. The calculation formula is as follows: Among them, Q, K, V represent the query vector, key vector and value vector generated by the input features through different projection matrices, T represents the transposition operation, and d k represents the dimension of the key vector; The multi-layer perceptron MLP is implemented through a series of linear, ReLU and Dropout operations; Finally, the result of the multi-head attention mechanism is passed through the MLP_Head layer, that is, two linear layers Linear to obtain the final result cls_token_final.

8. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 7, characterized in that: In step S44, the feature fusion module includes a self-attention feature fusion mechanism module SACA, a convolution branch local feature processing module, and a ViT branch global feature processing module, wherein: The SACA module includes a global average pooling module GAP, a soft selection mechanism, and a dynamic weight generation module DWG. First, GAP compresses each feature to obtain a feature summary: G i =GAP(X i ) Among them, X i Represents the characteristics of the i-th channel of the two branches; Then enter the Soft selection mechanism, which uses FC operation, including multiple Linear layers, ReLU layers, and Sigmod layers. The implementation formula is as follows: S i =Sigmod(ReLU(Linear(G i ))...) Then, dynamic weight generation is performed through the DWG module. The specific operation formula is as follows: w i =Sigmod(MLP(S i )) The final dynamic weight w i Weighted with the input features to get the result X SACA : By dynamically obtaining the different contributions of the two branch features, the result generation of the dual-branch fusion strategy is assisted; In the convolution branch local feature processing module, the two-dimensional shape of the convolution branch feature is first retained, and Conv, Flatten stretching and fusion operations are performed. The results are calculated as follows: X Mi =F M (X M(i-1) )+BAG V-M (X V(i-1) ) Where X M and X V represents the convolution branch features and ViT branch features, F M Represents the convolution feature branch Conv, Flatten operation, SACA represents the feature fusion attention mechanism; In the ViT branch global feature processing module, the residual fully connected layer Linear operation is performed on the ViT branch features, and the result is calculated as follows: X Vi =X V(i-1) +F V (X V(iV1) )+SACA M-V (F M (X M(i-1) )) Where X M and X V represents the convolution branch features and ViT branch features, F M Represents the convolution branch feature Conv, Flatten operation, F V It represents the linear operation of the residual fully connected layer of the ViT branch feature, and SACA represents the feature fusion attention mechanism.

9. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 8, characterized in that: In step S45, the fusion result of local branch features and global branch features is calculated as follows: X Out =Conv(X Mi +X Vi ) Finally, the fusion result X Oyr The fully connected layer FC and Softmax output are performed to obtain the final classification value. Among them, the fully connected layer FC maps the output feature vector to the category space, assists in obtaining the category probability distribution of the subsequent Softmax function, and takes the type with the maximum probability to indicate the LCZ category to which the image belongs.

10. The urban scene classification method based on high-resolution multi-view remote sensing images according to claim 1, characterized in that: In step S5, the network parameters for training the urban scene classification model include configuration parameters, loss function, and accuracy evaluation index, wherein the configuration parameters include batch size, learning rate, and number of iterations; The loss function uses the cross entropy loss function L VE : Where N is the number of samples, M is the number of classification categories; y ij Indicates whether the category is the same as the sample i category, if they are the same, it is 1, if they are different, it is 0; p ij It represents the predicted probability that the observed sample i belongs to category j; The accuracy evaluation indicators include overall accuracy OA, mapping accuracy PA, user accuracy UA, and harmonic average score F1-Score.

Citation Information

Patent Citations

  • Image classification method

    CN115222998A

  • A memory system, an electronic system including the same memory system and a method of opertaing the same electronic system

    KR1020250021811A

Cited By

  • Fourier transform-based large-model adaptive fine-tuning remote sensing land coverage classification method

    CN120510426A

  • High-precision classification method for object-oriented city remote sensing images

    CN120823444A

  • Urban ventilation corridor construction method based on local climate region classification and remote sensing inversion

    CN121052014A

  • CNN-based complex scene field adaptive remote sensing classification method and apparatus

    CN121330371A