A Multimodal LCZ Classification Method Based on Text Blending

By integrating remote sensing images and descriptive text information in the LCZ classification method, and using the deep fusion characteristics of self-attention and cross-attention modules, the problem of difficulty in capturing subtle differences in urban complex environments in the prior art is solved, and LCZ classification with higher accuracy and efficiency is achieved.

CN119723329BActive Publication Date: 2025-06-13耕宇牧星(北京)空间科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411766616.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-06-13
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The existing LCZ classification methods mainly rely on remote sensing image data, and it is difficult to fully capture subtle differences and deep semantic information in complex urban environments, resulting in inaccurate classification results.

Method used

A multimodal LCZ classification method based on text fusion is adopted, by integrating remote sensing images and descriptive text information, and using semantic information in the text to integrate into the classification process, a self-attention module and a cross-attention module are constructed to deeply integrate the image and text features.

Benefits of technology

It significantly improves the accuracy and efficiency of LCZ classification, can better capture subtle differences in complex urban environments, and provide a more scientific and reliable basis for urban planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723329B_ABST
    Figure CN119723329B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal LCZ classification method based on text fusion, including: constructing an LCZ classification model based on text fusion and training it to obtain an optimized LCZ classification model; inputting the LCZ remote sensing image to be processed into the optimized LCZ classification model to obtain the corresponding LCZ classification result; the S2 includes: processing the LCZ remote sensing image to be processed through a feature extraction network to obtain text-enhanced multi-modal fusion features; based on the text-enhanced multi-modal fusion features, obtaining the LCZ classification result through a classification network. By fusing text and image data, combining self-attention and cross-attention mechanisms, the accuracy and robustness of classification are significantly improved, and the full process automation of the classification task is realized, improving the application efficiency and providing a more scientific and reliable microclimate analysis basis for urban planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of LCZ classification, and more specifically, to a multi-modal LCZ classification method based on text fusion. Background Art

[0002] Local Climate Zones (LCZ) classification is a general method aimed at understanding and evaluating the microclimate characteristics of cities and their surrounding areas. This classification system is of great significance for urban planning and environmental science research. It can help researchers compare the climate characteristics of different cities, thus supporting urban planners in designing more suitable living environments. With the acceleration of urbanization, problems such as the urban heat island effect, energy consumption, and environmental pollution have become increasingly prominent. LCZ classification provides a scientific basis for effectively managing these problems.

[0003] However, existing LCZ classification methods mainly rely on remote sensing image data and use deep learning models such as convolutional neural networks (CNNs) and Transformers to extract and analyze image features. This method has obvious limitations in capturing complex urban microclimate characteristics. The urban environment is highly heterogeneous and dynamic. Relying solely on image data is difficult to fully reflect subtle and complex differences such as vegetation cover, building density, and human activity patterns. These factors are crucial for understanding local climate conditions, but often the information cannot be deeply mined, resulting in potentially inaccurate classification results, especially when dealing with areas that are visually similar but actually have very different climate characteristics.

[0004] In addition, traditional LCZ classification methods have limited ability to capture deep semantic information. Although remote sensing images can provide rich visual information, they are insufficient for detailed semantic descriptions of specific land uses, building types, material properties, etc. Such information usually requires the combination of geographical information system (GIS) data or text descriptions to obtain a complete understanding. Due to the lack of effective integration of non-image data sources such as text, existing methods fail to fully utilize this contextual knowledge that can supplement and enrich image interpretation, thus limiting the depth of the model's understanding of the environmental background and affecting the overall performance of the classification task.

[0005] Most current LCZ classification technologies are based on a single-modal data source - remote sensing images, which to some extent restricts their application scope and development potential. Single-modal methods are easily restricted by specific data types. For example, factors such as weather conditions and lighting conditions may significantly affect image quality, thereby interfering with classification accuracy. In addition, ignoring other potentially valuable information sources, such as text descriptions and meteorological observation records, makes it difficult for the model to comprehensively grasp the actual situation of the target area and cannot meet the growing needs of urban management and environmental protection.

[0006] Therefore, how to design a multi-modal LCZ classification method based on text fusion, which can capture the subtle differences in the complex urban environment by integrating remote sensing images and descriptive text information and improve the accuracy and efficiency of LCZ classification, is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a multi-modal LCZ classification method based on text fusion, which not only needs to make full use of the spectral and spatial information in the image, but also can integrate the semantic information in the text into the classification process, and can show significant advantages in capturing the subtle differences in the complex urban environment, providing a more scientific and reliable basis for urban planning, so as to effectively cope with the urban heat island effect and other environmental problems.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A multi-modal LCZ classification method based on text fusion, comprising:

[0010] S1. Construct an LCZ classification model based on text fusion and perform training to obtain an optimized LCZ classification model;

[0011] S2. Input the LCZ remote sensing image to be processed into the optimized LCZ classification model to obtain the corresponding LCZ classification result; the S2 includes:

[0012] S21. Process the LCZ remote sensing image to be processed through a feature extraction network to obtain a text-enhanced multi-modal fusion feature;

[0013] S22. Based on the text-enhanced multi-modal fusion feature, obtain the LCZ classification result through a classification network.

[0014] Further, in the S1, the cross-entropy loss function L CE is used for model training; the cross-entropy loss function L CE is expressed as:

[0015]

[0016] where M represents the total number of LCZ categories, y c represents the true value label, represents the probability that the model predicts to belong to category c.

[0017] Further, in the S21, processing the LCZ remote sensing image to be processed through a feature extraction network includes:

[0018] Based on the input remote sensing image Img, the image embedding feature is obtained through a multi-layer perceptron MLP The image embedding feature is expressed as:

[0019]

[0020] Input the image embedding feature into a 12-layer self-attention module to obtain the image attention features of the 2nd layer the image attention features of the 4th layer the image attention features of the 10th layer and the image attention features of the 12th layer

[0021] Furthermore, in S21, when processing the LCZ remote sensing image to be processed through the feature extraction network, it further includes:

[0022] Based on the input remote sensing image Img, generate the descriptive text Text corresponding to the remote sensing image Img through the text generator VLM;

[0023] Map the descriptive text Text into the feature space by combining with the CLIP text encoder to generate text features The text features are expressed as:

[0024]

[0025] Input the text features into the multi-layer perceptron MLP for dimensional transformation, and obtain the text attention features through the self-attention module The text attention features are expressed as:

[0026]

[0027] Input the text attention features into three consecutive cross-attention modules to perform cross-attention calculations with the image attention features of different layers in turn to obtain the fused attention feature CAO T ; The fused attention feature CAO T is expressed as:

[0028]

[0029] Concatenate the fused attention feature CAO T with the image attention features of the 12th layer to obtain the text-enhanced multi-modal fusion feature MMF; The text-enhanced multi-modal fusion feature MMF is expressed as:

[0030]

[0031] Among them, Concat represents the splicing operation.

[0032] Furthermore, the process of the self-attention module processing the image embedding features includes:

[0033] Performing layer normalization on the image embedding features to obtain the layer-normalized vector X E ; The layer-normalized vector X E is expressed as:

[0034]

[0035] where LayerNorm(·) represents the layer normalization function;

[0036] Inputting the layer-normalized vector X E into the multi-head attention structure to obtain the fused information MultiHead(X E ); Among them, in the multi-head attention structure, each attention head independently performs attention calculation, and multiple attention heads process in parallel;

[0037] Performing residual connection processing on the fused information MultiHead(X E ) and the image embedding features , and performing layer normalization processing to obtain the result Output after layer normalization; The result Output after layer normalization is expressed as:

[0038]

[0039] Inputting the result Output after layer normalization into the multi-layer perceptron MLP to obtain the output of the multi-layer perceptron; Performing residual connection processing on the output of the multi-layer perceptron and the result Output after layer normalization, and performing layer normalization processing to obtain the output of the self-attention module The output of the self-attention module is expressed as:

[0040]

[0041] where j represents the number of self-attention modules passed through.

[0042] Furthermore, in the multi-head attention structure, each attention head passes the layer-normalized vector X through three independent fully connected layers EThey are respectively mapped to a query vector Q, a key vector K, and a value vector V; the query vector Q, key vector K, and value vector V in the i-th attention head are respectively expressed as:

[0043]

[0044] where respectively represent independent learnable weight matrices in different fully connected layers;

[0045] Calculate the dot product of the query vector Q i and the key vector K i and perform scaling processing to obtain attention scores;

[0046] Normalize the attention scores and multiply them by the value vector V i to obtain the attention calculation result head i ; the attention calculation result head i is expressed as:

[0047]

[0048] where softmax(·) represents the normalization function, T represents the transpose operation of the matrix, and d k represents the dimension of the key vector;

[0049] Concatenate the information extracted by all attention heads and perform transformation through a linear layer to obtain the fused information MultiHead(X E ); the fused information MultiHead(X E ) is expressed as:

[0050] MultiHead(X E ) = Concat(head 1 , head 2 , …, head h )W O

[0051] where W O represents the learnable weight matrix, and Concat represents the concatenation process.

[0052] Furthermore, the processing process of the self-attention module for the text feature includes:

[0053] Perform layer normalization on the text feature to obtain the layer-normalized vector T E ; the layer-normalized vector T E is expressed as:

[0054]

[0055] Among them, LayerNorm(·) represents the layer normalization function;

[0056] Input the layer-normalized vector T E into the multi-head attention structure to obtain the fused information MultiHead(T E ); among them, in the multi-head attention structure, each attention head independently performs attention calculation, and multiple attention heads process in parallel;

[0057] Perform residual connection processing on the fused information MultiHead(T E ) and the text features , and perform layer normalization processing to obtain the result Output after layer normalization; the result Output after layer normalization is expressed as:

[0058]

[0059] Input the result Output after layer normalization into the multi-layer perceptron MLP to obtain the output of the multi-layer perceptron; perform residual connection processing on the output of the multi-layer perceptron and the result Output after layer normalization, and perform layer normalization processing to obtain the output of the self-attention module The output of the self-attention module is expressed as:

[0060]

[0061] where j represents the number of self-attention modules passed through.

[0062] Furthermore, in the multi-head attention structure, each attention head maps the layer-normalized vector T E to the query vector Q, key vector K, and value vector V respectively through three independent fully connected layers; the query vector Q, key vector K, and value vector V in the i-th attention head are respectively expressed as:

[0063]

[0064] where respectively represent the independent learnable weight matrices in different fully connected layers;

[0065] Calculate the dot product of the query vector Q i and the key vector K i and perform scaling processing to obtain the attention scores;

[0066] Normalize the attention scores and multiply them with the value vector V iMultiply to obtain the attention calculation result head i ; The attention calculation result head i is expressed as:

[0067]

[0068] where softmax(·) represents the normalization function, T represents the transpose operation of the matrix, and d k represents the dimension of the key vector;

[0069] Concatenate the information extracted by all attention heads and perform transformation through a linear layer to obtain the fused information MultiHead(T E ); The fused information MultiHead(T E ) is expressed as:

[0070] MultiHead(T E ) = Concat(head 1 , head 2 , …, head h )W O

[0071] where W O represents the learnable weight matrix, and Concat represents the concatenation process.

[0072] Furthermore, the cross-attention module includes: calculation branches CAM Img and CAM Text ; The processing process of the cross-attention module for image data and text data includes:

[0073] The calculation branches CAM Img and CAM Text map the image data and text data to query vector Q, key vector K, and value vector V respectively through three independent fully connected layers; where, in the calculation branch CAM Img , the query vector Q, key vector K, and value vector V in the i-th attention head are respectively expressed as Q i Img , K i Img and V i Img ; In the calculation branch CAM Text , the query vector Q, key vector K, and value vector V in the i-th attention head are respectively expressed as Q i Text , K iText and V i Text ;

[0074] In the computational branch CAM Img the dot product of the query vector Q i Text and the key vector K i Img is calculated and scaled to obtain attention scores; the attention scores are normalized and multiplied by the value vector V i Img to obtain the attention calculation result The attention calculation result is expressed as:

[0075]

[0076] In the computational branch CAM Text the dot product of the query vector Q i Img and the key vector K i Text is calculated and scaled to obtain attention scores; the attention scores are normalized and multiplied by the value vector V i Text to obtain the attention calculation result The attention calculation result is expressed as:

[0077]

[0078] Based on the attention calculation result and the outputs CAO Img of the computational branch CAM Text and the computational branch CAM Img and CAO Text ;

[0079] The outputs CAO Img and CAO Text are concatenated to obtain the output CAO of the cross-attention module.

[0080] Furthermore, the S22 includes:

[0081] Based on the text-enhanced multimodal fusion features, the ReLU non-linear activation function is used to introduce non-linearity;

[0082] Combined with a 1×1 convolutional layer, the text-enhanced multimodal fusion features with introduced non-linearity are adjusted in feature depth and the feature channels are fused;

[0083] Output the predicted probability of each LCZ category through the softmax function; where, the S22 is expressed as:

[0084] y = softmax(Conv 1×1 (ReLU(MMF)))

[0085] where, y represents the LCZ classification result, softmax represents the softmax function, Conv 1×1 represents a 1×1 convolution operation, ReLU represents the ReLU non-linear activation function, and MMF represents the text-enhanced multi-modal fusion feature.

[0086] It can be seen from the above technical solutions that, compared with the prior art, the technical solutions of the present invention have the following

[0087] beneficial effects:

[0088] 1. Existing LCZ classification mainly relies on remote sensing image data, and uses deep learning models for feature extraction and analysis, but there are limitations in capturing complex urban microclimate features and deep semantics. This method enhances the understanding of scene features in remote sensing images by fusing text and image data and using geographical and environmental background knowledge in the text, improving the accuracy and reliability of classification. Especially when dealing with LCZ categories with fuzzy or overlapping features, multi-modal data fusion effectively reduces the risk of misclassification.

[0089] 2. By building a self-attention module (SAM) and a cross-attention module (CAM), deep feature extraction and fusion of image and text data are realized. The self-attention module can capture the global features and long-range dependencies of the data, while the cross-attention module calculates the cross-modal fusion features and the feature dependency relationships between them, thereby generating a richer and more comprehensive feature representation. This mechanism helps the model better understand complex urban environments and improves the robustness and generalization ability of the classification model.

[0090] 3. This method provides a multi-modal LCZ classification model based on text blending, which can realize the full-process automation of the LCZ classification task in remote sensing images without manual intervention. This not only improves the efficiency of the classification task but also reduces the risk of human errors. In addition, when dealing with urban heat island effects and other environmental problems, this method can provide a more scientific and reliable basis for urban planning and has important practical application value. Description of the Drawings

[0091] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained according to the provided accompanying drawings.

[0092] Figure 1 Flowchart of the multi-modal LCZ classification method based on text fusion provided by the embodiment of the present invention;

[0093] Figure 2 Framework diagram of the LCZ classification model based on text fusion provided by the embodiment of the present invention;

[0094] Figure 3 Framework diagram of the feature extraction network provided by the embodiment of the present invention;

[0095] Figure 4 Framework diagram of the self-attention module provided by the embodiment of the present invention;

[0096] Figure 5 Framework diagram of the multi-head attention structure provided by the embodiment of the present invention;

[0097] Figure 6 Framework diagram of the cross-attention module provided by the embodiment of the present invention. Detailed implementation manners

[0098] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0099] As Figure 1 shown, this embodiment provides a multi-modal LCZ classification method based on text fusion, including:

[0100] S1. Construct an LCZ classification model based on text fusion and train it to obtain an optimized LCZ classification model;

[0101] S2. Input the LCZ remote sensing image to be processed into the optimized LCZ classification model to obtain the corresponding LCZ classification result;

[0102] As Figure 2 shown, the S2 includes:

[0103] S21. Process the LCZ remote sensing image to be processed through a feature extraction network to obtain text-enhanced multi-modal fusion features;

[0104] S22. Based on the text-enhanced multi-modal fusion features, obtain the LCZ classification result through a classification network.

[0105] For this multi-modal LCZ classification method, an LCZ classification model that combines image and text data is constructed and trained; then the LCZ remote sensing image to be processed is input into this model, text-enhanced multi-modal fusion features are generated through a feature extraction network, and the final LCZ classification result is output via a classification network. It not only utilizes the spatial and spectral information of the image data but also incorporates the semantic context provided by the text description, thereby enhancing the understanding and classification accuracy of complex urban environmental features and providing a more scientific and reliable tool for urban planning and environmental research.

[0106] The following further elaborates on the above steps and related technical features in detail:

[0107] In this embodiment S1, an LCZ classification model based on text blending is constructed and trained to obtain an optimized LCZ classification model;

[0108] Constructing an LCZ classification model based on text blending mainly includes: First, the model designs a self-attention module (SAM) and a cross-attention module (CAM). Among them, SAM is used to capture the global features and long-range dependencies of a single data modality, while CAM is used to calculate the cross-modal fusion features and the feature dependence relationships between them to achieve the effective fusion of text and image data. Second, the model also includes a feature extraction module based on text blending. This module uses a pre-trained text generator and a CLIP text encoder to convert the text description corresponding to the remote sensing image into text features, and through multiple cross-attention modules, deeply fuse with the image features to generate a feature representation rich in multi-modal information. Finally, the text-enhanced multi-modal fusion features are processed through a classification network to output the LCZ classification result.

[0109] During the training process, the cross-entropy loss function is used to optimize the model. Through a large number of LCZ remote sensing images and their corresponding text descriptions for training, the model can learn the feature representations of different LCZ categories and gradually improve the classification accuracy and robustness.

[0110] The cross-entropy loss function L CE is expressed as:

[0111]

[0112] where M represents the total number of LCZ categories, and y c represents the true value label. Represents the probability that the model predicts belonging to class c.

[0113] After training is completed, an optimized LCZ classification model is obtained, which can automatically perform the LCZ classification task and achieve an accurate division of the urban microclimate zone.

[0114] In this embodiment S2, the LCZ remote sensing image to be processed is input into the optimized LCZ classification model to obtain the corresponding LCZ classification result; it includes:

[0115] S21. Process the LCZ remote sensing image to be processed through the feature extraction network to obtain text-enhanced multimodal fusion features;

[0116] S22. Based on the text-enhanced multimodal fusion features, obtain the LCZ classification result through the classification network.

[0117] As Figure 3 shown, in step S21, processing the LCZ remote sensing image to be processed through the feature extraction network includes:

[0118] Based on the input remote sensing image Img, obtain the image embedding feature through the multi-layer perceptron MLP The image embedding feature Is expressed as:

[0119]

[0120] Input the image embedding feature into the 12-layer self-attention module to obtain the 2nd layer image attention feature the 4th layer image attention feature the 10th layer image attention feature and the 12th layer image attention feature

[0121] Furthermore, it also includes:

[0122] Based on the input remote sensing image Img, generate the descriptive text Text corresponding to the remote sensing image Img through the text generator VLM;

[0123] Combine the CLIP text encoder to map the descriptive text Text into the feature space to generate the text feature The text feature Is expressed as:

[0124]

[0125] Input the text feature Perform dimensional transformation through a multi-layer perceptron (MLP) and obtain text attention features through a self-attention module The text attention features are expressed as:

[0126]

[0127] Take the text attention features and perform cross-attention calculations with the image attention features of different layers in sequence through three consecutive cross-attention modules to obtain the fused attention feature CAO T ; The fused attention feature CAO T is expressed as:

[0128]

[0129] Take the fused attention feature CAO T and concatenate it with the image attention feature of the 12th layer to obtain the text-enhanced multimodal fusion feature MMF; The text-enhanced multimodal fusion feature MMF is expressed as:

[0130]

[0131] where Concat represents the concatenation operation

[0132] As Figure 4 shown, the processing process of the self-attention module for the embedded features includes:

[0133] Perform layer normalization on the image embedded features to obtain the layer-normalized vector X E ; The layer-normalized vector X E is expressed as:

[0134]

[0135] where LayerNorm(·) represents the layer normalization function

[0136] Input the layer-normalized vector X E into the multi-head attention structure to obtain the fused information MultiHead(X E ); Among them, in the multi-head attention structure, each attention head independently performs attention calculations, and multiple attention heads process in parallel

[0137] For the fused information MultiHead(X E ) and the image embedded features Perform residual connection processing and layer normalization processing to obtain the result Output after layer normalization processing; the result Output after layer normalization processing is expressed as:

[0138]

[0139] Input the result Output after layer normalization processing into a multi-layer perceptron MLP to obtain the output of the multi-layer perceptron; perform residual connection processing on the output of the multi-layer perceptron and the result Output after layer normalization processing, and perform layer normalization processing to obtain the output of the self-attention module The output of the self-attention module Is expressed as:

[0140]

[0141] Among them, j represents the number of self-attention modules passed through.

[0142] As Figure 5 Shown, in the multi-head attention structure, each attention head maps the layer-normalized vector X E To the query vector Q, key vector K, and value vector V respectively through three independent fully connected layers; the query vector Q, key vector K, and value vector V in the i-th attention head are respectively expressed as:

[0143]

[0144] Among them, Respectively represent independent learnable weight matrices in different fully connected layers;

[0145] Calculate the dot product of the query vector Q i And the key vector K i And perform scaling processing to obtain the attention score;

[0146] Normalize the attention score and multiply it by the value vector V i To obtain the attention calculation result head i ; The attention calculation result head i Is expressed as:

[0147]

[0148] Among them, softmax(·) represents the normalization function, T represents the transpose operation of the matrix, and d k Represents the dimension of the key vector; d k Is used to assist in the scaling operation, and the scaling operation helps to avoid generating too large a value before performing the softmax operation.

[0149] Concatenate the information extracted by all attention heads and transform it through a linear layer to obtain the fused information MultiHead(X E );The fused information MultiHead(X E ) is expressed as:

[0150] MultiHead(X E ) = Concat(head 1 , head 2 , …, head h )W O

[0151] Among them, W O represents a learnable weight matrix, and Concat represents the concatenation process.

[0152] Furthermore, the processing process of the self-attention module for the text feature includes:

[0153] Perform layer normalization on the text feature to obtain the layer-normalized vector T E ; The layer-normalized vector T E is expressed as:

[0154]

[0155] Among them, LayerNorm(·) represents the layer normalization function;

[0156] Input the layer-normalized vector T E into the multi-head attention structure to obtain the fused information MultiHead(T E ); Among them, in the multi-head attention structure, each attention head independently performs attention calculation, and multiple attention heads process in parallel;

[0157] Perform residual connection processing on the fused information MultiHead(T E ) and the text feature , and perform layer normalization processing to obtain the result Output after layer normalization; The result Output after layer normalization is expressed as:

[0158]

[0159] Input the result Output after layer normalization into the multi-layer perceptron MLP to obtain the output of the multi-layer perceptron; Perform residual connection processing on the output of the multi-layer perceptron and the result Output after layer normalization, and perform layer normalization processing to obtain the output of the self-attention module The output of the self-attention module is expressed as:

[0160]

[0161] where j represents the number of self-attention modules passed through.

[0162] Furthermore, in the multi-head attention structure, each attention head maps the layer-normalized vector T E to a query vector Q, a key vector K, and a value vector V respectively through three independent fully connected layers; the query vector Q, key vector K, and value vector V in the i-th attention head are respectively expressed as:

[0163]

[0164] where respectively represent independent learnable weight matrices in different fully connected layers;

[0165] Calculate the dot product of the query vector Q i and the key vector K i and perform scaling processing to obtain attention scores;

[0166] Normalize the attention scores and multiply them with the value vector V i to obtain the attention calculation result head i ; the attention calculation result head i is expressed as:

[0167]

[0168] where softmax(·) represents the normalization function, T represents the transpose operation of the matrix, and d k represents the dimension of the key vector;

[0169] Concatenate the information extracted by all attention heads and perform transformation through a linear layer to obtain the fused information MultiHead(T E ); the fused information MultiHead(T E ) is expressed as:

[0170] MultiHead(T E ) = Concat(head 1 , head 2 , …, head h )W O

[0171] where W O represents the learnable weight matrix, and Concat represents the concatenation process.

[0172] In the above step S21, the processing of the LCZ remote sensing image is implemented in detail. First, the input remote sensing image is converted into image embedding features through a multi-layer perceptron, and then these features are fed into a 12-layer self-attention module to obtain image attention features at different levels. At the same time, based on the content of the remote sensing image, a text generator is used to generate descriptive text, and these texts are mapped to the feature space through a CLIP text encoder to form text features. Then, the text features are further processed through a multi-layer perceptron and a self-attention module to obtain text attention features. After that, these text attention features will pass through three consecutive cross-attention modules and be fused with image attention features at different levels to generate fused attention features. Finally, the fused attention features are concatenated with the 12th-layer image attention features to form text-enhanced multi-modal fusion features, providing rich feature representations for subsequent LCZ classification.

[0173] In addition, in the feature extraction process, the self-attention module plays a key role. It first deeply processes the image and text embedding features through layer normalization and a multi-head attention structure to capture the global features and long-range dependencies of the data. In the multi-head attention structure, each attention head independently performs attention calculations, and multiple heads process in parallel, thereby extracting feature information from multiple perspectives. The processed information then undergoes residual connection, layer normalization, and further transformation by a multi-layer perceptron to obtain the output of the self-attention module. This process ensures the richness and diversity of the features, laying a solid foundation for subsequent cross-attention calculations and multi-modal fusion.

[0174] As Figure 6 shown, the cross-attention module includes: calculation branches CAM Img and CAM Text ; the processing process of the cross-attention module for image data and text data includes:

[0175] Calculation branches CAM Img and CAM Text map the image data and text data to query vectors Q, key vectors K, and value vectors V through three independent fully connected layers respectively; among them, in the calculation branch CAM Img , the query vector Q, key vector K, and value vector V in the i-th attention head are respectively represented as Q i Img , K i Img and V i Img ; in the calculation branch CAMText Among them, the query vector Q, key vector K, and value vector V in the i-th attention head are respectively expressed as Q i Text , K i Text and V i Text ;

[0176] In the calculation branch CAM Img Among them, calculate the dot product of the query vector Q i Text and the key vector K i Img and perform scaling processing to obtain attention scores; perform normalization processing on the attention scores, and multiply them by the value vector V i Img to obtain the attention calculation result The attention calculation result is expressed as:

[0177]

[0178] In the calculation branch CAM Text Among them, calculate the dot product of the query vector Q i Img and the key vector K i Text and perform scaling processing to obtain attention scores; perform normalization processing on the attention scores, and multiply them by the value vector V i Text to obtain the attention calculation result The attention calculation result is expressed as:

[0179]

[0180] Based on the attention calculation result and obtain the outputs CAO Img of the calculation branch CAM Text and the calculation branch CAM Img and CAO Text ;

[0181] Connect the outputs CAO Img and CAO Text to obtain the output CAO of the cross-attention module.

[0182] The cross-attention module processes image data and text data by calculating branches separately. Each branch uses three independent fully connected layers to map the input into query vector Q, key vector K, and value vector V. In each attention head, the calculation branch uses image data as the query vector, while text data as the key and value vectors; conversely, the calculation branch uses text data as the query vector and image data as the key and value vectors. Then, the dot product of the query vector and the key vector is calculated and scaled in both branches to obtain attention scores. These scores are normalized and then multiplied by the corresponding value vectors to generate the respective attention calculation results. Finally, the outputs of the two branches are concatenated to form the comprehensive output of the cross-attention module, realizing the effective fusion of information between the image and text modalities and enhancing the model's understanding and classification ability of complex LCZ features.

[0183] In step S22 of this embodiment, based on the text-enhanced multi-modal fusion features, the LCZ classification result is obtained through a classification network. Specifically, it includes:

[0184] Based on the text-enhanced multi-modal fusion features, the ReLU non-linear activation function is used to introduce non-linear characteristics;

[0185] Combined with a 1×1 convolutional layer, the text-enhanced multi-modal fusion features with introduced non-linear characteristics are adjusted in terms of feature depth and feature channels are fused;

[0186] The prediction probability of each LCZ category is output through the softmax function; where S22 is expressed as:

[0187] y = softmax(Conv 1×1 (ReLU(MMF)))

[0188] where y represents the LCZ classification result, softmax represents the softmax function, Conv 1×1 represents the 1×1 convolutional operation, ReLU represents the ReLU non-linear activation function, and MMF represents the text-enhanced multi-modal fusion features.

[0189] In this step, by introducing non-linear characteristics and performing feature adjustment, the recognition ability of the classification network for complex LCZ features is significantly improved, thereby enhancing the accuracy and reliability of the classification.

[0190] This embodiment provides a multi-modal LCZ classification method based on text blending. By constructing and training an LCZ classification model that combines image and text data, precise division of urban microclimate zones is achieved. In the model construction stage, a self-attention module (SAM) and a cross-attention module (CAM) are designed to capture the global features of a single modality and the fusion features across modalities. The text generator and CLIP text encoder pre-trained are used to convert the text description corresponding to the remote sensing image into text features, which are deeply fused to generate a feature representation rich in multi-modal information. During the training process, the cross-entropy loss function is used to optimize the model parameters to ensure that it learns the feature representations of different LCZ categories.

[0191] In the classification stage, the LCZ remote sensing image to be processed is first processed by a feature extraction network to obtain text-enhanced multi-modal fusion features; then the final LCZ classification result is output through a classification network. Specifically, the feature extraction network not only extracts embedded features from the image and processes them through multiple self-attention modules, but also generates text features through the text generator and CLIP text encoder, and deeply fuses them with the image features through multiple cross-attention modules to form text-enhanced multi-modal fusion features. Finally, the classification network introduces a non-linear activation function and a 1×1 convolutional layer to adjust the fusion features, and outputs the prediction probability of each LCZ category through the softmax function, significantly improving the accuracy and reliability of classification, and providing scientific and reliable tool support for urban planning and environmental research.

[0192] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0193] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal LCZ classification method based on text fusion, characterized in that: include: S1. Build and train a LCZ classification model based on text fusion to obtain an optimized LCZ classification model; S2, inputting the LCZ remote sensing image to be processed into the optimized LCZ classification model to obtain the corresponding LCZ classification result; S2 includes: S21. Processing the LCZ remote sensing image to be processed by a feature extraction network to obtain a multimodal fusion feature for text enhancement; wherein processing the LCZ remote sensing image to be processed by a feature extraction network includes: Based on the input remote sensing image Img, the image embedding features are obtained through the multi-layer perceptron MLP The image embedding feature It is expressed as: Embedding images into features Input 12 layers of self-attention modules to obtain the second layer of image attention features Layer 4 image attention features Layer 10 image attention features and the 12th layer image attention features Also includes: Based on the input remote sensing image Img, a descriptive text Text corresponding to the remote sensing image Img is generated through a text generator VLM; Combined with CLIP text encoder, the descriptive text Text is mapped into feature space to generate text features The text features It is expressed as: The text features The multi-layer perceptron MLP is used to transform the dimension and the text attention feature is obtained through the self-attention module The text attention feature It is expressed as: Text Attention Features Through three consecutive cross-attention modules, cross-attention calculations are performed with the image attention features of different layers in turn to obtain the fused attention feature CAO T ; The fusion attention feature CAO T It is expressed as: The fusion attention feature CAO T With the 12th layer image attention feature The splicing process is performed to obtain a text-enhanced multimodal fusion feature MMF; the text-enhanced multimodal fusion feature MMF is expressed as: Among them, Concat represents the concatenation operation; S22. Based on the multimodal fusion features of the text enhancement, an LCZ classification result is obtained through a classification network.

2. The multimodal LCZ classification method based on text fusion according to claim 1, characterized in that: In S1, the cross entropy loss function L is used CE Perform model training; the cross entropy loss function L CE It is expressed as: Where M represents the total number of LCZ categories, y c represents the true value label, represents the probability that the model predicts that it belongs to category c.

3. The multimodal LCZ classification method based on text fusion according to claim 1, characterized in that: The self-attention module embeds features into the image The processing process includes: Embedding features into images Perform layer normalization to obtain layer normalized vector X E ; The layer normalized vector X E It is expressed as: Among them, LayerNorm(·) represents the layer normalization function; The vector X that normalizes the layer E Input the multi-head attention structure to obtain the fused information MultiHead(X E ); In the multi-head attention structure, each attention head performs attention calculation independently, and multiple attention heads are processed in parallel; The fused information MultiHead(X E ) and image embedding features Perform residual connection processing and layer normalization processing to obtain the result Output after layer normalization processing; the result Output after layer normalization processing is expressed as: The result Output after layer normalization processing is input into the multi-layer perceptron MLP to obtain the multi-layer perceptron output; the multi-layer perceptron output is subjected to residual connection processing with the result Output after layer normalization processing, and layer normalization processing is performed to obtain the self-attention module output The self-attention module outputs It is expressed as: Among them, j represents the number of self-attention modules passed.

4. The multimodal LCZ classification method based on text fusion according to claim 3 is characterized in that: In the multi-head attention structure, each attention head transforms the layer-normalized vector X through three independent fully connected layers. E They are mapped to query vector Q, key vector K and value vector V respectively; the query vector Q, key vector K and value vector V in the i-th attention head are expressed as: in, Represent the independent learnable weight matrices in different fully connected layers respectively; Calculate the query vector Q i and the key vector K i The dot product of and scaling is performed to obtain the attention score; The attention scores are normalized and summed with the value vector V i Multiply to get the attention calculation result head i ; The attention calculation result head i It is expressed as: Among them, softmax(·) represents the normalization function, T represents the transposition of the matrix, and d k represents the dimension of the key vector; The information extracted by all attention heads is concatenated and transformed through a linear layer to obtain the fused information MultiHead(X E ); the fused information MultiHead(X E ) is expressed as: MultiHead(X E )=Concat(head1,head2,…,head h )W O Among them, W O Represents a learnable weight matrix, and Concat represents concatenation.

5. The multimodal LCZ classification method based on text fusion according to claim 1, characterized in that: The self-attention module is used to analyze the text features The processing process includes: Text features Perform layer normalization to obtain the layer normalized vector T E ; The layer normalized vector T E It is expressed as: Among them, LayerNorm(·) represents the layer normalization function; The vector T that normalizes the layer E Input the multi-head attention structure to obtain the fused information MultiHead(T E ); In the multi-head attention structure, each attention head performs attention calculation independently, and multiple attention heads are processed in parallel; The fused information MultiHead(T E ) and text features Perform residual connection processing and layer normalization processing to obtain the result Output after layer normalization processing; the result Output after layer normalization processing is expressed as: The result Output after layer normalization processing is input into the multi-layer perceptron MLP to obtain the multi-layer perceptron output; the multi-layer perceptron output is subjected to residual connection processing with the result Output after layer normalization processing, and layer normalization processing is performed to obtain the self-attention module output The self-attention module outputs It is expressed as: Among them, j represents the number of self-attention modules passed.

6. The multimodal LCZ classification method based on text fusion according to claim 5, characterized in that: In the multi-head attention structure, each attention head passes through three independent fully connected layers to normalize the vector T E They are mapped to query vector Q, key vector K and value vector V respectively; the query vector Q, key vector K and value vector V in the i-th attention head are expressed as: in, Represent the independent learnable weight matrices in different fully connected layers respectively; Calculate the query vector Q i and the key vector K i The dot product of and scaling is performed to obtain the attention score; The attention scores are normalized and summed with the value vector V i Multiply to get the attention calculation result head i ; The attention calculation result head i It is expressed as: Among them, softmax(·) represents the normalization function, T represents the transposition of the matrix, and d k represents the dimension of the key vector; The information extracted by all attention heads is concatenated and transformed through a linear layer to obtain the fused information MultiHead(T E ); the fused information MultiHead(T E ) is expressed as: MultiHead(T E )=Concat(head1,head2,…,head h )W O Among them, W O Represents a learnable weight matrix, and Concat represents concatenation.

7. The multimodal LCZ classification method based on text fusion according to claim 1, characterized in that: The cross attention module includes: calculating the branch CAM Img and CAM Text ; The cross attention module is used for image data and text data The processing process includes: Compute Branch CAM Img and CAM Text The image data is respectively transformed into and text data Mapped into query vector Q, key vector K and value vector V; where in the calculation branch CAM Img In the ith attention head, the query vector Q, key vector K, and value vector V are represented as Q i Img , K i Img and V i Img ; In the calculation branch CAM Text In the ith attention head, the query vector Q, key vector K, and value vector V are represented as Q i Text , K i Text and V i Text ; In the calculation branch CAM Img In the example above, we calculate the query vector Q i Text and the key vector K i Img The dot product of and scaling is performed to obtain an attention score; the attention score is normalized and combined with the value vector V i Img Multiply to get the attention calculation result The attention calculation result It is expressed as: In the calculation branch CAM Text In the example above, we calculate the query vector Q i Img and the key vector K i Text The dot product of and scaling is performed to obtain an attention score; the attention score is normalized and combined with the value vector V i Text Multiply to get the attention calculation result The attention calculation result It is expressed as: Based on the attention calculation result and Get the calculation branch CAM Img and compute branch CAM Text Output of CAO Img and CAO Text ; The output CAO Img and CAO Text Perform connection processing to obtain the output CAO of the cross attention module.

8. The multimodal LCZ classification method based on text fusion according to claim 1, characterized in that: The S22 includes: Based on the multimodal fusion features of the text enhancement, a ReLU nonlinear activation function is used to introduce nonlinear characteristics; Combined with the 1×1 convolutional layer, the multimodal fusion features of text enhancement with nonlinear characteristics are introduced to adjust the feature depth and fuse the feature channels; The predicted probability of each LCZ category is output through the softmax function; wherein, S22 is expressed as: y=softmax(Conv 1×1 (ReLU(MMF))) Among them, y represents the LCZ classification result, softmax represents the softmax function, Conv 1×1 represents a 1×1 convolution operation, ReLU represents the ReLU nonlinear activation function, and MMF represents the multimodal fusion feature for text enhancement.

Citation Information

Patent Citations

  • Remote sensing image scene classification method based on multi-mode airspace transformation network

    CN116503753A

  • Remote sensing image-based local climate region classification system and method

    CN118864985A