Local climate region classification method based on multi-source data and deep learning
By constructing the MMA-Net model, using multi-source data and deep learning methods, combining multi-scale fusion and global attention mechanisms, the problem of insufficient classification accuracy in local climate regions is solved, especially in the identification of urban building types, and higher classification accuracy and flexibility are achieved.
Patent Information
- Application Number
- CN202510437667.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, local climate zone classification accuracy is insufficient, especially in urban dense and heterogeneous environments, the model has low intra-class similarity to building types LCZ 1-10, making it difficult to effectively extract local features.
The local climate zone LCZ classification model MMA-Net is constructed, satellite image information is obtained through multi-source data processing, combined with the multi-scale fusion module MS, the cascade feature aggregation module MLA and the global attention mechanism GAM, multi-scale feature extraction and feature weighting are performed, and deep learning methods are used for classification.
It improves the accuracy and complexity of local climate zone classification, especially in the identification of urban building types, reduces in-class confusion, improves the flexibility and feature extraction capabilities of the model, and achieves higher classification accuracy.
Smart Images

Figure CN120387086A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a local climate zone classification method based on multi-source data and deep learning, belonging to the technical field of local climate zone classification. Background Art
[0002] For the classification of local climate zones (LCZs), neighborhood information includes the formation of local microclimate by the interaction of local building structures, surface materials, and vegetation. How to extract these local features is crucial for improving the LCZ classification results. In deep learning methods, the Residual Neural Network (ResNet) and its variants are superior in local feature extraction, mainly manifested in allowing cross-level information transmission, enhancing feature reusability, and supporting deeper network structures. These advantages have enabled ResNet to achieve significant performance advantages in tasks such as image classification and object detection, and have been successfully applied in the classification of LCZs. However, as cities become more and more dense and heterogeneous, the model's ability to handle complexity and extract fine details is still lacking, resulting in low intra-class similarity for building type LCZ 1-10 based on elevation and density. Summary of the Invention
[0003] The purpose of the present invention is to provide a local climate zone classification method based on multi-source data and deep learning to solve the problem of insufficient classification accuracy of local climate zones in the prior art.
[0004] A local climate zone classification method based on multi-source data and deep learning obtains multi-band information in satellite images, performs cloud removal after resampling to obtain data to be classified in the study area, constructs a local climate zone LCZ classification model MMA-Net, and inputs the data to be classified into MMA-Net to obtain the local climate zone classification result.
[0005] The MMA-Net includes using a multi-scale fusion module MS to perform multi-scale feature extraction on the input feature Input, inputting the stacked feature maps output by MS into a cascaded feature aggregation module MLA. The MLA includes three stage layers, embeds the attention mechanism SENet into the stage layers, and the stage layers also include a convolutional block Conv Block and an identity block Identity Block; a global attention mechanism GAM is added after each stage layer to perform global pooling to obtain multi-dimensional features; the multi-dimensional features are average-pooled and then input into a fully connected layer to obtain the LCZ classification result.
[0006] The MS includes three convolutional layers Conv with convolutional kernels of 5×5, 3×3, and 1×1 respectively, and filter sizes of 16, 32, and 16 respectively. The input features are simultaneously input into the three convolutional layers, and the outputs of the three filters are connected to form a 64-channel stacked feature map Merge.
[0007] The three stage layers are Stage1, Stage2, and Stage3 respectively.
[0008] Each stage layer sequentially includes a first Conv Block, a second Conv Block, a first Identity Block, a third Conv Block, and a second Identity Block.
[0009] The Conv Block sequentially includes an input layer input, a two-dimensional convolutional layer conv2d, a batch normalization layer batchnorm, an activation function relu, conv2d, batchnorm, relu, conv2d, batchnorm, SENet, conv2d, batchnorm, relu. There are shortcut connections connecting the third batchnorm, SENet, and the fourth batchnorm to the last relu, and there is a shortcut connection from input to the last conv2d.
[0010] The Identity Block sequentially includes input, conv2d, batchnorm, relu, conv2d, batchnorm, relu, conv2d, batchnorm, SENet, relu. There are shortcut connections from input, the last batchnorm, and SENet to the last relu.
[0011] The SENet sequentially includes a residual layer Residual, a global pooling layer Global pooling, a fully connected layer fc, relu, fc, an activation function sigmoid, and a scale layer scale. The input features of the SENet are feature fused with the result of scale.
[0012] GAM includes rearranging the input feature F1, inputting it into a multi-layer perceptron, then rearranging it again, and inputting it into the activation function sigmoid to obtain the output feature McF1. Taking McF1 as the input feature F2, performing two convolutional operations with a convolutional kernel of 7×7 in sequence, and then inputting it into the activation function sigmoid to obtain the output feature McF2.
[0013] Performing global pooling includes performing global pooling on the output of the third-stage layer and the outputs of all GAMs respectively, and then fusing them to obtain multi-dimensional features.
[0014] Compared with the prior art, the present invention has the following beneficial effects: The present invention extracts local information of an image from multiple perspectives, makes full use of early features of the model to strengthen the basic information of ground objects, the attention mechanism weights the features to make the model more flexible, and ablation experiments also show the effectiveness of the module and the complementarity between different modules. The present invention also performs excellently in the comparative experiment, achieving a good balance between accuracy and complexity, analyzing the characteristics of the LCZ map in the experimental area, and being of great significance to urban morphology and climate research. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a structural diagram of the local climate zone LCZ classification model MMA-Net;
[0016] Figure 2 It is a structural diagram of the global attention mechanism GAM;
[0017] Figure 3 It is a dissimilarity matrix;
[0018] Figure 4 It is a first confusion matrix;
[0019] Figure 5 It is a second confusion matrix;
[0020] Figure 6 It is the user accuracy of different baseline models;
[0021] Figure 7 It is the specific improvement effect of MS and MLA on the classification accuracy of different LCZ categories. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] A local climate zone classification method based on multi-source data and deep learning, obtains multi-band information in satellite images, performs cloud removal processing after resampling to obtain data to be classified in the study area, constructs a local climate zone LCZ classification model MMA-Net, and inputs the data to be classified into MMA-Net to obtain a local climate zone classification result.
[0024] The MMA-Net is as Figure 1As shown, it includes using the multi-scale fusion module MS to perform multi-scale feature extraction on the input feature Input, inputting the stacked feature maps output by MS into the cascaded feature aggregation module MLA. MLA includes three stage layers, embedding the attention mechanism SENet into the stage layers, and the stage layers also include a convolutional block Conv Block and an identity block IdentityBlock; adding a global attention mechanism GAM after each stage layer, performing global pooling to obtain multi-dimensional features; performing average pooling on the multi-dimensional features, and then inputting them into a fully connected layer to obtain the LCZ classification result.
[0025] MS includes three convolutional layers Conv with convolutional kernels of 5×5, 3×3, and 1×1 respectively, and filter sizes of 16, 32, and 16 respectively. The input feature is input into the three convolutional layers simultaneously, and the outputs of the three filters are connected to form a 64-channel stacked feature map Merge.
[0026] The three stage layers are Stage1, Stage2, and Stage3 respectively.
[0027] Each stage layer sequentially includes a first Conv Block, a second Conv Block, a first Identity Block, a third Conv Block, and a second Identity Block.
[0028] The Conv Block sequentially includes an input layer input, a two-dimensional convolutional layer conv2d, a batch normalization layer batchnorm, an activation function relu, conv2d, batchnorm, relu, conv2d, batchnorm, SENet, conv2d, batchnorm, relu. The third batchnorm, SENet, and the fourth batchnorm are all provided with shortcut connections to the last relu, and input is provided with a shortcut connection to the last conv2d.
[0029] The Identity Block sequentially includes input, conv2d, batchnorm, relu, conv2d, batchnorm, relu, conv2d, batchnorm, SENet, relu. Input, the last batchnorm, and SENet are all provided with shortcut connections to the last relu.
[0030] The SENet successively includes a residual layer Residual, a global pooling layer Global pooling, a fully connected layer fc, relu, fc, an activation function sigmoid, and a scale layer scale. The input features of the SENet are fused with the result of scale.
[0031] The GAM is as Figure 2 shown, which includes rearranging the input features F1, inputting them into a multi-layer perceptron, rearranging again, and inputting the activation function sigmoid to obtain the output features McF1. Taking McF1 as the input features F2, two convolution operations with a convolution kernel of 7×7 are performed in sequence, and then the activation function sigmoid is input to obtain the output features McF2.
[0032] Performing global pooling includes performing global pooling on the output of the third stage layer and the outputs of all GAMs respectively, and then fusing them to obtain multi-dimensional features.
[0033] The present invention uses NVIDIA GeForce RTX 4060Ti as the hardware and Python 3.10.4 and Pytorch as the framework. This framework has a good programming toolbox and is usually used to implement deep learning models. A Dropout layer with a size of 0.2 is added before the fully connected layer to further reduce the risk of overfitting. The Softmax classifier calculates the voting probability of each feature vector for each LCZ class and outputs the highest probability class as the final result. For the training hyperparameters, a batch size of 64 batches is introduced as the data capacity of the input images and LCZ labels in each load, the learning rate is set to 0.001, and the adaptive moment estimation (adam) optimizer is selected to minimize the cross-entropy loss function. To control the training time and avoid overfitting, early stopping with 10 epochs is used, which means that if the validation loss does not decrease within 10 full dataset iterations, the training stops. After training is completed, the model will record the parameters of the network for predicting new data.
[0034] In the evaluation of the model accuracy, four traditional classification accuracy evaluation metrics are used, namely Kappa, OA, OA n and OA nb . The traditional classification accuracy evaluation metrics only consider whether the samples are correctly classified and do not consider that some classes are more similar to each other than others. Since the impact of misclassifications with a very small impact on the whole will be amplified, the accuracy is low. To alleviate this problem, the present invention introduces a weighted accuracy measure (WA). WA introduces a dissimilarity matrix such as Figure 3As described above, penalizing the confusion between different types provides a more comprehensive and effective method for evaluating the LCZ classification effect. This method imposes a lower penalty on the misclassification of physically similar LCZ classes and a higher penalty on the misclassification of physically dissimilar LCZ classes according to the surface characteristics of the local climate, including openness, height, land cover, and surface object type.
[0035] To test the performance of MMA-Net in LCZ classification, a baseline model comparison experiment was carried out in the study. Among the baseline models, ResNet18, ResNet34, ResNet50, and ResNet101 are several common variants of ResNet, with differences in the depth and number of parameters of the network. ResNeXt mainly uses grouped convolution based on ResNet and also adds residual connections. ResNeSt introduces the Split-Attention module on the basis of ResNeXt, enabling the model to extract features separately from multiple convolutional kernel branches in the same layer, making the features extracted by the network more diverse. DenseNet is a densely connected network that eliminates information bottlenecks by introducing dense connection blocks in the network, enhancing the feature reuse ability of the model, and effectively solving the problem of gradient disappearance. All models in the comparison between the baseline model and MMA-Net in the validation experiment were trained from scratch for meaningful comparison. The experimental settings of the present invention are shown in Table 1.
[0036] Table 1. Experimental settings
[0037]
[0038]
[0039] After training, the model can be used for LCZ classification by the sliding window method. To keep the input consistent with the So2SatLCZ42 dataset, the input image is also set to a 32×32 window size with a 10-meter resolution of Sentinel-2 here, and the stride is set to 5. The window slides within the study area, and an LCZ classification result is generated for each window.
[0040] The present invention mainly uses satellite images of Sentinel-2A / B for LCZ classification. The Sentinel-2 images use 8 bands from visible light to near-infrared, omitting the bands (bands 1, 9, and 10) that reflect the coastal aerosol, water vapor, and cirrus cloud conditions because they are of little significance in LCZ classification. To create a synthesis with a consistent image size, all bands are upsampled to a 10m resolution using resampling. Images with a cloud cover of less than 20% in August 2018 in the study area were screened and de-clouded, and finally, the data to be classified with a spatial resolution of 10 meters in the study area was synthesized.
[0041] This paper compares multiple baseline models with the aim of selecting the baseline that best fits the experimental data and improving on it. The precision comparisons of different baseline models are shown in Table 2, and the user precision is as Figure 6 shown.
[0042] Table 2. Precision Comparison of Different Baseline Models
[0043]
[0044] From ResNet18 to ResNet101, as the model depth increases, the performance gradually improves, and ResNet101 exhibits the highest precision. As variants of ResNet, ResNeSt and ResNeXt have made improvements in the design of residual blocks by integrating nested residual blocks and group convolutions into the backbone network. According to the comparative experiments, these variants do not show advantages over the original ResNet series. Their accuracy metrics are generally lower than ResNet50. DenseNet eliminates the information bottleneck through dense connection blocks, and its accuracy is generally between ResNet50 and ResNet101, and slightly lower than ResNet50 in some precision metrics such as OA b and OA nb .
[0045] Although ResNet50 is more concise than ResNet101 in terms of model depth and the number of parameters, its classification performance is not significantly inferior to ResNet101. Regarding the classification precision of the model in different LCZs, ResNet50 shows better recognition effects in LCZ1, LCZ2, LCZ5, and LCZ6, and these types are also the main types of urban built-up area LCZs. Accurately distinguishing these LCZ types has an important impact on the overall classification precision. In addition, due to the moderate depth of the model, the training cost of ResNet50 is also less than that of other complex ResNet networks. Therefore, considering both the model performance and the training cost, ResNet50 is finally selected as the baseline model for classification improvement, and it is found that the most serious confusion occurs in LCZ1, LCZ5, LCZ7, and LCZ10.
[0046] Using the So2Sat LCZ42 training set, MMA-Net is trained, and the performance of the models in the two trainings is evaluated by cross-entropy loss and confusion matrix as Figure 4 and Figure 5As shown in the figure. MMA-Net, which is improved from ResNet50, has an overall cross-entropy loss smaller than that of ResNet50 (G0), with a fast decline rate and excellent performance. From the confusion matrix on the validation set, for G0, the main confusion is concentrated in the built-up area. The confusion between LCZ8 (large low-rise buildings) and LCZ10 (heavy industrial plants) reaches 32.8%. They are both large-volume low-rise buildings with similar background environments and slightly different materials, making it difficult to distinguish them from their forms. The confusions between LCZ2 and LCZ5, and between LCZ3 (compact low-rise) and LCZ6 (open low-rise) are 15.8% and 17.7% respectively. The reasons are that occlusions cause incomplete or blurred information in some areas, the shadows projected by buildings, and the complex appearance of buildings may all lead to confusion with the surrounding environment, making it difficult to accurately classify them as sparse or dense. At the same time, the problem of unclear floor height recognition is also worthy of attention, mainly reflected in the cases of LCZ1 and LCZ2, LCZ2 and LCZ3, and LCZ4 and LCZ5. After using MMA-Net, the recognition of building types and heights has been effectively improved, especially for the confusions that occur between sparse built-up areas (LCZ3, LCZ6) and other classes, the confusions of compactness (the confusion between LCZ5 and LCZ2 is reduced by 8.9%, the confusion between LCZ6 and LCZ3 is reduced by 6.8%, etc.), and the confusions of floor height (the misjudgment of LCZ2 as LCZ3 is reduced by 5.1%, the misjudgment of LCZ2 as LCZ8 is reduced by 5.3%, etc.).
[0047] This invention attempts to improve the LCZ classification accuracy using three methods, SENet, MS, and MLA, and their combinations, as shown in Table 3. Compared with scenario G0, the accuracy is slightly improved after adding SE (scenario G1). On the basis of SE, coupling MLA (scenario G3) has a more significant improvement in accuracy than coupling MS (scenario G2). The Kappa, WA, OA, and OA_b of G3 are improved by 4.4%, 0.6%, 4%, and 6.3% respectively compared with G2. Coupling the three mechanisms (MMA) has the most obvious improvement in classification progress, and its WA and OA_nb exceed 96%.
[0048] Table 3. LCZ Classification Accuracy Results
[0049]
[0050] The specific improvement effects of MS and MLA on the classification accuracy of different LCZ categories can be seen in Figure 7。The accuracy of natural classes is already relatively high. Comparing G2 with G0, the improvement of MS is mainly reflected in distinguishing the differences among LCZ1-6. Especially for LCZ1 (compact high-rise), the accuracy improvement is relatively large. When MS is not added (G0), the objects it is confused with are mainly LCZ2 and LCZ4 (Fig6a). The differences between them and LCZ1 lie in the tightness and building height. Thus, it can be seen that the addition of MS can help the model better identify the background and height of the correct buildings. The addition of MLA (G3) is mainly reflected in the improvement of LCZ7 (lightweight low-rise buildings) and LCZ10 (heavy industrial plants). Since their main confused objects before improvement were both LCZ8 (large low-rise buildings), it shows that the basic information such as object edges, textures, and colors extracted by low-level features in the early stage of the network can better help identify the basic style features such as the structural form of large low-rise houses.
[0051] In the present invention, the GAM attention mechanism is used, and its ability to achieve cross-dimensional information interaction in channels and space to capture long-range dependencies between features is utilized to improve the LCZ classification accuracy. Comparing with the SE, ECA, and CBAM attention mechanism algorithms, the improvement of GAM on the LCZ classification accuracy is shown in Table 4.
[0052] Table 4. Improvement results of GAM on LCZ classification accuracy
[0053]
[0054] Among various attention mechanisms, GAM has an obvious effect on improving the LCZ classification accuracy, and the WA reaches 0.97. The improvements of ECA on Kappa, OA, OA_b, and OA_nb are 79%, 80%, 76%, and 86% respectively, which are less than 84%, 87%, 82%, and 93% of CBAM. The results show that when simultaneously paying attention to channel and spatial information, CBAM can utilize spatial information more comprehensively than ECA which only focuses on channels. The structure, shape, size, and spatial layout of buildings are closely related. By capturing the differences between these features, different types of buildings can be better distinguished. Moreover, spatial information can help the model understand the position, layout, and relationship of buildings in their surrounding environment. These neighborhood information is very important for classification decisions, especially when the appearance or structural features of buildings overlap between different categories. As an advancement of CBAM, GAM performs the best, which fully demonstrates that the optimization of feature extraction by cross-dimensional interaction helps to reduce redundant information and enhance the perception ability of key features, and is suitable for more complex visual tasks.
[0055] The LCZ distribution classified by MMA-Net basically conforms to the urban building and spatial distribution characteristics of the experimental area. The long history of the experimental area is integrated with its rapid development, resulting in a mixture of high and low buildings with obvious height contrast. Large-volume low-rise buildings are extremely widespread in the experimental area. According to the classification results, large low-rise buildings such as warehouses and factories are well separated from other types. At the same time, in the most prosperous areas, high-rise and low-rise buildings are well distinguished. For background recognition, the MMA-Net model is also relatively effective, separating sparse and dense backgrounds well.
[0056] There are obvious misclassifications in the farmland along the Yangtze River. Many farmlands are classified as LCZ9. The main reason is that the 320m spatial resolution is not sufficient to clearly distinguish small-scale buildings and farmlands, and they are merged into similar spectral mixed pixels, making it difficult for the model to accurately distinguish them.
[0057] The present invention discloses a new deep learning framework that integrates MS, MLA, and an attention mechanism to extract multi-scale and multi-level features in LCZ classification and perform LCZ mapping on the experimental area. MS extracts local information of the image from multiple perspectives, MLA makes full use of the early features of the model to strengthen the basic information of ground objects, and in response to changes such as the scale and deformation of ground objects, the attention mechanism weights the features, making the model more flexible. Ablation experiments also demonstrate the effectiveness of the module and the complementarity between different modules. The model also performs excellently in comparative experiments, achieving a good balance between accuracy and complexity. Further, the impact of class imbalance in So2Sat LCZ42 on this model is explored, and it is found that class-weighted training significantly improves the problem of the model ignoring small classes, providing valuable reference for solving the problem of insufficient class samples in LCZ mapping. On this basis, the characteristics of the LCZ map of the experimental area are analyzed, which is of great significance for urban morphology and climate research.
[0058] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for classifying local climate zones based on multi-source data and deep learning, characterized in that, Obtain the information of multiple bands in the satellite image, perform cloud removal after resampling to obtain the data to be classified in the study area, construct the local climate zone LCZ classification model MMA-Net, and input the data to be classified into MMA-Net to obtain the local climate zone classification result.
2. The local climate zone classification method based on multi-source data and deep learning according to claim 1, characterized in that The MMA-Net includes using the multi-scale fusion module MS to perform multi-scale feature extraction on the input feature Iuput, and inputting the stacked feature maps output by MS into the cascaded feature aggregation module MLA. MLA includes three stage layers, and the attention mechanism SENet is embedded into the stage layers. The stage layers also include a convolutional block Conv Block and an identity block Identity Block; Add a global attention mechanism GAM after each stage layer, and perform global pooling to obtain multi-dimensional features; Perform average pooling on the multi-dimensional features, and then input them into the fully connected layer to obtain the LCZ classification result.
3. The method for classifying local climate zones based on multi-source data and deep learning according to claim 2, characterized in that, MS includes three convolutional layers Conv, with convolutional kernels of 5×5, 3×3, and 1×1 respectively, and filter sizes of 16, 32, and 16 respectively. The input feature is input into the three convolutional layers at the same time, and the outputs of the three filters are connected to form a 64-channel stacked feature map Merge.
4. The method for classifying local climate zones based on multi-source data and deep learning according to claim 3, characterized in that The three stage layers are Stage1, Stage2, and Stage3 respectively.
5. The local climate zone classification method based on multi-source data and deep learning according to claim 4, characterized in that Each stage layer sequentially includes a first Conv Block, a second Conv Block, a first Identity Block, a third Conv Block, and a second Identity Block.
6. The method for classifying local climate zones based on multi-source data and deep learning according to claim 5, wherein The Conv Block sequentially includes an input layer input, a two-dimensional convolutional layer conv2d, a batch normalization layer batchnorm, an activation function relu, conv2d, batchnorm, relu, conv2d, batchnorm, SENet, conv2d, batchnorm, relu. The third batchnorm, SENet, and the fourth batchnorm are all provided with shortcut connections connecting to the last relu, and input is provided with a shortcut connection connecting to the last conv2d.
7. The local climate zone classification method based on multi-source data and deep learning according to claim 6, characterized in that, The Identity Block sequentially includes input, conv2d, batchnorm, relu, conv2d, batchnorm, relu, conv2d, batchnorm, SENet, relu. input, the last batchnorm, and SENet are all provided with shortcut connections connecting to the last relu.
8. The method for classifying local climate zones based on multi-source data and deep learning according to claim 7, wherein The SENet sequentially includes a residual layer Residual, a global pooling layer Global pooling, a fully connected layer fc, relu, fc, an activation function sigmoid, and a scale layer scale. The input feature of SENet is fused with the result of scale.
9. The method for classifying local climate zones based on multi-source data and deep learning according to claim 8, wherein The GAM includes rearranging the input feature F1, inputting it into a multi-layer perceptron, rearranging it again, inputting it into the sigmoid activation function to obtain the output feature McF1, using McF1 as the input feature F2, performing two consecutive convolution operations with a convolution kernel of 7×7, and then inputting it into the sigmoid activation function to obtain the output feature McF2.
10. The local climate zone classification method based on multi-source data and deep learning according to claim 9, characterized in that Performing global pooling includes performing global pooling on the output of the third-stage layer and the outputs of all GAMs respectively, and then fusing them to obtain multi-dimensional features.