A large model driven high-resolution remote sensing image interpretation method
By constructing a large-scale intelligent model for multimodal remote sensing information extraction, the accuracy and consistency issues of high-resolution remote sensing image interpretation methods in complex scenarios are solved, achieving efficient and accurate remote sensing image interpretation and supporting urban planning and disaster assessment.
Patent Information
- Application Number
- CN202411733922.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing high-resolution remote sensing image interpretation methods suffer from poor accuracy and consistency when dealing with complex scenes, and are time-consuming and labor-intensive, making it difficult to effectively utilize multimodal remote sensing data.
A large-scale intelligent model for multimodal remote sensing information extraction is constructed, including a visible light information encoder and a thermal infrared information encoder. Through image preprocessing and multimodal feature fusion, a large language model is used for high-resolution remote sensing image interpretation.
It improves the accuracy and efficiency of remote sensing image interpretation, reduces misjudgments and omissions, lowers the cost of manual intervention, provides more accurate remote sensing image interpretation results, and supports urban planning and disaster assessment.
Smart Images

Figure CN119851148B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of large models and deep learning, and particularly relates to a large model driven high-resolution remote sensing image interpretation method. BACKGROUND
[0002] With the rapid development of remote sensing technology, the ability to acquire large-scale remote sensing image data has significantly improved. These data have shown great potential for application in many fields such as environmental monitoring, urban planning, agricultural management, and disaster warning. High-resolution remote sensing image interpretation, which extracts valuable geographic spatial information from these data, converts complex remote sensing images into human-readable and analyzable text information. It requires interpretation methods to understand the spatial relationships and semantic connections between various ground object targets in high-resolution remote sensing images, thereby generating accurate and comprehensive geographic spatial information descriptions. High-resolution remote sensing image interpretation not only builds a bridge between remote sensing data and practical applications, enabling the conversion of raw pixel data into understandable and analyzable information resources, providing a solid foundation for scientific research, policy making, and decision support. Moreover, it plays a crucial role in promoting global sustainable development, optimizing resource allocation, strengthening ecological and environmental protection, and improving disaster warning capabilities. Therefore, exploring efficient and accurate high-resolution remote sensing image interpretation methods is of great significance for promoting the in-depth application and development of remote sensing technology.
[0003] Traditional remote sensing image interpretation methods mainly rely on human experience and traditional image processing techniques. These methods are competent when dealing with small-scale and simple scenes, but they often perform poorly when faced with complex scenes and diverse ground object targets in high-resolution remote sensing images. In addition, manual interpretation is not only time-consuming and labor-intensive, but also susceptible to subjective factors, resulting in poor reliability and consistency of the interpretation results. In recent years, with the rapid development of artificial intelligence and deep learning technologies, deep learning-based image interpretation techniques have made significant progress. These methods can automatically extract feature information from massive data, and then accurately interpret complex scenes through large language models. However, these methods still face challenges such as large data volume, high computational complexity, and long training time when dealing with large-scale remote sensing image data. Currently, in both the image processing field and the natural language processing field, large models have shown great advantages in feature extraction and processing due to their complex structure and large-scale parameters. SUMMARY
[0004] The present application proposes a large model driven high-resolution remote sensing image interpretation method to address the technical problem of poor high-resolution remote sensing image interpretation, in order to realize the interpretation of high-resolution remote sensing images and effectively improve the accuracy of high-resolution remote sensing image interpretation.
[0005] To solve the above technical problems, the technical scheme adopted by the present application is as follows:
[0006] A large model driven high-resolution remote sensing image interpretation method, comprising the following steps:
[0007] Step 1, a large-scale multi-modal high-resolution remote sensing image dataset is constructed, which includes visible light images and thermal infrared images; and the image data in the multi-modal high-resolution remote sensing image dataset is preprocessed to obtain a training set for model training;
[0008] Wherein, the image preprocessing includes the image data format unification of each modality, image size normalization and image cropping; wherein the size and number of each image cropped sub-image are consistent, and then the cropped image is scaled to the size of the cropped sub-image; for each modality, all cropped sub-images and scaled cropped images to the size of the cropped sub-image are regarded as a sample;
[0009] Step 2, a multi-modal remote sensing information extraction intelligent large model is constructed, which is used for multi-modal remote sensing information coding;
[0010] Wherein, the multi-modal remote sensing information extraction intelligent large model includes a visible light information encoder and a thermal infrared information encoder, the visible light information encoder is used to extract visible light coding features of visible light images, and the thermal infrared information encoder is used to extract thermal infrared coding features of thermal infrared images;
[0011] Step 3, a multi-modal remote sensing information fusion module is constructed, which is used for multi-modal feature fusion of visible light coding features and thermal infrared coding features, and outputs the fused multi-modal features;
[0012] Step 4, a multi-modal decoder for high-resolution remote sensing image interpretation is constructed based on the network structure of the encoder of the large language model, the fused multi-modal features are used as the input of the multi-modal decoder, and the interpretation result of the high-resolution remote sensing image is obtained based on the output thereof;
[0013] Step 5, based on the training set composed of multiple samples, the model parameters of the remote sensing image interpretation model composed of the multi-modal remote sensing information extraction intelligent large model, the multi-modal remote sensing information fusion module and the multi-modal decoder are learned and trained, when the preset training convergence condition (the number of times reaches the upper limit or the loss function value set converges) is met, the interpretation result of the input target image is obtained based on the trained remote sensing image interpretation model. The input target image includes visible light images and thermal infrared images of the same target object.
[0014] Further, in step 1, the image cropping includes: based on the specified overlap rate, the image size normalized image is cropped according to the specified cropping order (such as from top to bottom, from left to right order), and the image resolution of the cropped sub-image is consistent, and the image resolution of the cropped sub-image of each image is consistent, and the number of cropped sub-images is consistent. Further, all cropped sub-images and the cropped image scaled to the size of the cropped sub-image are regarded as a sample of each modality.
[0015] For example, when the image size is normalized, the visible light image and the thermal infrared image can be cropped into 2000x2000 resolution images, respectively. For another example, when the image is cropped to obtain a plurality of cropped sub-images, each size normalized image can be cropped into 16-25 sub-images with a resolution of 512x512 according to an overlap rate of 10%-30%.
[0016] Further, in step 2, the visible light information encoder and the thermal infrared information encoder in the multi-modal remote sensing information extraction intelligent large model have consistent structures, each including n parallel global high-order information embedding Transformer branches, where n is the number of cropped sub-images of each image plus 1; and each global high-order information embedding Transformer branch includes a plurality of consecutive feature extraction stages, i.e., each global high-order information embedding Transformer branch includes m cascaded global high-order information embedding Transformer modules, which can be set to 4 in general, and each global high-order information embedding Transformer module corresponds to a stage; and in each stage of the visible light information encoder and the thermal infrared information encoder, the information joint perception of different global high-order information embedding Transformer branches is performed through the cross-region information joint intelligent perception module. That is, the output feature maps of the global high-order information embedding Transformer modules of the same order in the n branches are subjected to information joint perception through the cross-region information joint intelligent perception module, and the information joint perception features of each branch (i.e., n information joint perception features) are output, and then output to the global high-order information embedding Transformer module of the next order of the corresponding branch.
[0017] Further, the global high-order information embedding Transformer module of the present application can improve the perception ability of the model to similar ground objects and constantly changing ground objects, thereby laying a good foundation for realizing accurate interpretation. The global high-order information embedding Transformer module in the present application is specifically set as follows:
[0018] The input vector of the global high-order information embedding Transformer module is subjected to query weight matrix W Q , key weight matrix W K, value weight matrix W V Linear mapping is query vector Q, key vector K and value vector V;
[0019] Covariance pooling operation is performed on query vector Q, key vector K and value vector V respectively, and then the obtained pooled vector and the corresponding vector before pooling are point multiplied to obtain a new vector Q n , K n and V n ;
[0020] The vectors Q n , K n and V n are sent into a multi-head self-attention module for self-attention mechanism calculation to obtain the output of the multi-head self-attention module Wherein, softmax() represents a softmax activation function, d k represents the dimension of vector K n ;
[0021] The output of the multi-head self-attention module and the input vector of the global high-order information embedding Transformer module are subjected to feature superposition through a first superposition & normalization layer, and then normalized to obtain the first self-attention fusion feature;
[0022] The first self-attention fusion feature is sent into a full connection layer, and the output of the full connection layer is further subjected to feature superposition through a second superposition & normalization layer, and then normalized to obtain the second self-attention fusion feature, that is, the output feature of the global high-order information embedding Transformer module.
[0023] Further, the covariance pooling operation is specifically:
[0024] The channel number reduction operation is performed on vectors Q, K and V through a convolution layer with a convolution kernel of 1×1 respectively to obtain new feature vectors Q', K' and V' with the same channel number, and C M represents the reduced channel number;
[0025] The correlation between the channels in each feature vector Q', K' and V' is calculated respectively to obtain correlation matrices M M , M M , M Q with the dimension of C K ×C V ; wherein, the calculation of the feature map correlation can be through Pearson correlation coefficient, cosine similarity, element-wise multiplication, etc., in the present application, the cosine similarity is preferred, that is, each correlation matrix is obtained by calculating the cosine similarity between the feature maps;
[0026] The new feature vectors Q', K' and V' are respectively multiplied by the corresponding correlation matrix M Q , M K , M V to obtain feature vectors S Q , S K , S V .
[0027] The feature vectors S Q , S K , S V are respectively passed through a global average pooling layer and a convolution layer with a 1x1 kernel to obtain three new feature vectors S' Q , S' K , S' V .
[0028] The feature vectors S' Q , S' K , S' V are respectively multiplied by the corresponding vectors Q, K and V to obtain new vectors Q n , K n and V n .
[0029] Further, the cross-region information joint intelligent perception module specifically performs the following process:
[0030] (1) The feature maps of each input global high-order information embedding Transformer branch (i.e. the output feature maps of the n global high-order information embedding Transformer modules of the same order) are respectively subjected to global average pooling and global maximum pooling, and then the two pooled results are element-wise added to obtain the first feature vector of each branch;
[0031] (2) The first feature vectors of all branches are element-wise added and then passed through several (preferably 3) fully connected layers and a ReLU activation function layer to obtain a global feature vector;
[0032] (3) The feature maps of each global high-order information embedding Transformer branch are respectively multiplied by the global feature vector to obtain the information joint perception features of each branch, which are used as the input features of the next stage global high-order information embedding Transformer module of the corresponding branch.
[0033] Further, in the multi-modal remote sensing information fusion module, the visible light encoding features and the thermal infrared features are fused through the cross multi-head self-attention mechanism.
[0034] Further, the multi-modal remote sensing information fusion module comprises a multi-layer multi-head self-attention mechanism layer, two groups of parallel self-attention calculations are performed in each multi-head self-attention mechanism layer, one group takes the visible light encoded feature as the value vector, and takes the thermal infrared encoded feature as the query vector and the key vector; the other group takes the thermal infrared encoded feature as the value vector, and takes the visible light encoded feature as the query vector and the key vector; and after each group of self-attention calculations is completed, the output feature maps of the two groups are added element by element as the input feature of the next layer of multi-head self-attention mechanism layer.
[0035] Further, in step 5, the multi-modal remote sensing information extraction intelligent large model in the remote sensing image interpretation model is first trained by a random size mask self-supervised learning strategy, which specifically comprises:
[0036] One training object is selected in the visible light information encoder and the thermal infrared information encoder, and a general decoder (for example, a general decoder based on a convolutional network) is connected after the training object, which is used to reconstruct the input image of the training object;
[0037] The input image (visible light or thermal infrared image) of the training object is randomly selected in a specified plurality of image scales (for example, three image sizes of 224x224, 512x512 and 1024x1024, which can be set based on actual application scenarios) as the size of the current processing image;
[0038] On the image of this size, a certain proportion of image blocks are randomly selected for mask processing, and the mask matrix M used for mask processing cannot exceed the size of the selected processing image (for example, one of four sizes of 30x30, 90x90, 120x120 and 256x256); preferably, the specific operation of the mask can be to replace the selected image block with all black (pixel value of 0);
[0039] The loss function of the random size mask self-supervised learning strategy of the training object is set as:
[0040]
[0041] Wherein, H, W and C represent the height, width and channel number of the input image respectively; ∑M represents the total number of unmasked elements in the mask matrix M, for example, for the mask operation of replacing the selected image block with all black (pixel value of 0), ∑M is the total number of elements with value 1 in the mask matrix M, I hwc and I and I respectively represent the pixel values of the original image and the reconstructed image at the image position (h, w, c), wherein the original image is the input image of the training object.
[0042] When the preset training convergence conditions are met, the trained visible light information encoder and thermal infrared information encoder are obtained in the form of parameter sharing based on the trained training objects.
[0043] Furthermore, in step 5, the multimodal remote sensing information fusion module and the multimodal decoder in the remote sensing image interpretation model are trained on weight parameters based on the pre-trained multimodal remote sensing information extraction intelligent large model. Preferably, a cross entropy loss function can be used for training.
[0044] The technical solution provided by the present invention brings at least the following beneficial effects:
[0045] (1) By constructing a large-scale model for intelligent extraction of multimodal remote sensing information, this invention achieves the extraction and coordination of multimodal remote sensing features, effectively utilizing multiple information sources in remote sensing images and improving the integrity and accuracy of feature extraction. This technological improvement makes the interpretation of high-resolution remote sensing images more precise and reduces the possibility of misjudgments and missed detections.
[0046] (2) The design of the multimodal remote sensing information fusion module achieves a deep fusion of visible light coding features and thermal infrared coding features, further improving the model's interpretation capabilities. This cross-modal feature fusion method is technologically innovative and provides new ideas for the interpretation of high-resolution remote sensing images.
[0047] (3) This invention reduces the cost and time of manual intervention by improving the accuracy and efficiency of high-resolution remote sensing image interpretation. In the field of remote sensing image analysis, this technological improvement can significantly reduce the consumption of human resources and lower the operating costs of enterprises.
[0048] (4) In terms of urban planning, the present invention can provide more accurate and timely remote sensing image interpretation results, providing a scientific basis for urban planners to optimize urban layout and resource allocation.
[0049] (5) In terms of environmental monitoring and disaster assessment, the application of the present invention can timely discover environmental problems and disaster risks, provide early warning and decision-making support to relevant departments, and protect people’s lives and property. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 A schematic diagram of a flow chart of an embodiment of the present invention;
[0052] Figure 2 A structural diagram of a remote sensing information intelligent extraction large model of an embodiment of the present application;
[0053] Figure 3 A structural diagram of a multi-head attention structure in a global high-order information embedding Transformer of an embodiment of the present application;
[0054] Figure 4 A structural diagram of a cross-region information joint intelligent perception module of an embodiment of the present application;
[0055] Figure 5 A structural diagram of a multi-modal remote sensing information fusion module constructed by an embodiment of the present application. DETAILED DESCRIPTION
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application but not all the embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0057] The embodiments of the present application provide a large model driven high-resolution remote sensing image interpretation method, which aims to solve the problem of poor current remote sensing image interpretation effect.
[0058] As a possible implementation manner, referring to Figure 1 The method provided by the embodiments of the present application includes the following steps:
[0059] Step 1, a large-scale multi-modal high-resolution remote sensing image dataset is constructed, and image preprocessing is performed on the image data in the dataset to construct a training set and a test set. The multi-modal high-resolution remote sensing image dataset includes visible light images and thermal infrared images, and the image preprocessing includes image data format unification, image size normalization, and image cropping of each modality; the constructed training set is used for learning and training of model parameters, and the test set is used for verifying the performance of the trained model.
[0060] In this embodiment, large-scale remote sensing image data is collected from multiple remote sensing data sources, including visible light images and thermal infrared images, and is uniformly converted into a standardized format. In addition, data cleaning and labeling can be performed to remove noise, correct distortion, and improve data quality through data augmentation techniques, preparing for subsequent processing and analysis. The processed data set is then divided into training and test sets according to a predetermined ratio.
[0061] In one embodiment, step 1 specifically includes the following steps:
[0062] Step 1.1: Adjust the resolution of images in all data sets to 2000x2000.
[0063] Step 1.2: All visible light images and thermal infrared images are respectively cropped from top to bottom and left to right, with each original image being cropped into 16-25 sub-images with a resolution of 512x512, and an overlap rate of 10%-30%. In addition, the original image is also individually adjusted to 512x512 size. All cropped sub-images of the original image and the scaled original image are considered as a complete data sample.
[0064] Step 1.3: The above processed samples are divided into training and test sets in a 1:1 ratio, respectively for network model training and testing.
[0065] Step 2, build a multi-modal remote sensing information intelligent extraction large model to realize the extraction and collaboration of multi-modal remote sensing information. The multi-modal remote sensing information intelligent extraction large model is divided into a visible light information encoder and a thermal infrared information encoder, which respectively process visible light images and thermal infrared images.
[0066] In one embodiment, step 2 includes the following steps:
[0067] In this embodiment, the visible light information encoder and the thermal infrared information encoder in step 2 have consistent structures, each including n (rounded to the number of cropped sub-images of each original image plus 1) parallel global high-order information embedding Transformer branches, such as Figure 2 Each image in each data sample in the data set is extracted by an independent global high-order information embedding visual Transformer branch. In each stage of the global high-order information embedding visual Transformer branch, cross-region information joint intelligent perception modules are used to realize the interaction of different branch features.
[0068] Preferably, the global high-order information embedding Transformer described above is a reinforcement of the original ViT, which can capture more rich global information and discriminative information in remote sensing large scenes and ground object targets, thereby improving the perception ability of the model to similar ground objects and changing ground objects and laying a foundation for accurate and comprehensive interpretation. The global high-order information embedding Transformer adds high-order information embedding in the original ViT.
[0069] Referring to Figure 3 , the global high-order information embedding Transformer module in the embodiment is specifically:
[0070] (1) The input vectors of the global high-order information embedding Transformer module (the first-order input vectors are directly the respective cropped subgraphs and the scaled images after the cropped images of the corresponding samples) are respectively linearly mapped into query vectors Q, key vectors K and value vectors V through a query weight matrix W Q , a key weight matrix W K and a value weight matrix W V .
[0071] (2) The covariance pooling operation is performed on the query vectors Q, the key vectors K and the value vectors V respectively, and then the dot product operation is performed on the obtained pooled vectors and the corresponding vectors before the pooling, to obtain new vectors Q n , K n and V n .
[0072] (3) The vectors Q n , K n and V n are sent into the multi-head self-attention module for self-attention mechanism calculation, to obtain the output of the multi-head self-attention module wherein, softmax() represents a softmax activation function, and d k represents the dimension of the vector K n .
[0073] (4) The feature stacking & normalization layer is used to perform feature stacking on the output of the multi-head self-attention module and the input vectors of the global high-order information embedding Transformer module, and then perform normalization operation, to obtain the first self-attention fusion feature.
[0074] (5) The first self-attention fusion feature is sent into the full connection layer, and then the output of the full connection layer is further processed by the second stacking & normalization layer, to perform feature stacking on the output of the full connection layer and the first self-attention fusion feature and then perform normalization operation, to obtain the second self-attention fusion feature, which is the output feature of the global high-order information embedding Transformer module.
[0075] The covariance pooling operation is specifically as follows:
[0076] The channel number reduction operation is performed on the vectors Q, K and V through a convolution layer with a convolution kernel of 1*1, respectively, to obtain new feature vectors Q', K' and V' with the same channel number, and C is defined as M , which represents the reduced channel number.
[0077] The correlation between the channels in each feature vector Q', K' and V' is calculated, respectively, to obtain correlation matrices M Q , M K , M V with dimensions of C M * C M ; wherein the calculation of the feature map correlation can be performed through the Pearson correlation coefficient, the cosine similarity, element-by-element multiplication, etc., and in the present application, the cosine similarity is preferred, that is, each correlation matrix is obtained by calculating the cosine similarity between the feature maps.
[0078] The new feature vectors Q', K' and V' are subjected to matrix multiplication with the corresponding correlation matrices M Q , M K , M V , respectively, to obtain feature vectors S Q , S K , S V .
[0079] The feature vectors S Q , S K , S V are subjected to a global average pooling layer and a convolution layer with a convolution kernel of 1*1, respectively, to obtain three new feature vectors S' Q , S' K , S' V .
[0080] The feature vectors S' Q , S' K , S' V are subjected to dot multiplication with the corresponding vectors Q, K and V, respectively, to obtain new vectors Q n , K n and V n .
[0081] Preferably, referring to Figure 4 , the processing procedure of the joint intelligent perception of different region information is as follows:
[0082] Firstly, the feature maps of each independent global high-order information embedding Transformer branch are respectively subjected to global average pooling and global maximum pooling, the obtained vectors are added pixel by pixel to obtain a vector with a size of 1×1×C'(C' is the set channel number), then the n vectors are added element by element and pass through a plurality of fully connected layers and a ReLU activation function to obtain a new vector S, the global vector also has a size of 1×1×C', and the above process is expressed by the following formula:
[0083]
[0084] Wherein, f represents a fully connected layer, GAP represents global average pooling, F i represents the output feature of the current stage of the i-th global high-order information embedding Transformer branch.
[0085] Next, the output feature maps of each branch at the current stage are point multiplied with the global vector S to obtain new feature maps as the input of the next stage of each branch, and the above process is expressed by the following formula:
[0086] V i = F i ⊙S + F i
[0087] Wherein, V i represents the input feature of the next stage of the i-th global high-order information embedding visual Transformer branch, and represents point multiplication operation.
[0088] Step 3, a multi-modal remote sensing information fusion module is constructed to realize the fusion of visible light coding features and thermal infrared coding features. The coding features from the visible light information encoder and the coding features from the thermal infrared information encoder are fused into remote sensing multi-modal features through the multi-modal remote sensing information fusion module.
[0089] Preferably, the multi-modal remote sensing information fusion module in step 3 is composed of a plurality of (for example, 12) multi-head self-attention mechanism layers. In each self-attention mechanism layer, two groups of parallel self-attention calculations are performed, one group of visible light coding features as Value, and thermal infrared coding features as Query and Key; the other group of thermal infrared coding features as Value, and visible light coding features as Query and Key. After each group of self-attention calculation is completed, the output feature maps of the two groups are added element by element, as shown in the following formula: Figure 5
[0090] Step 4, a multi-modal decoder for high-resolution remote sensing image interpretation is constructed based on the network structure of the large language model encoder, the fused multi-modal features are input into the multi-modal decoder, and the interpretation result of the high-resolution remote sensing image is obtained based on the output of the multi-modal decoder.
[0091] Step 5, learning and training the model parameters of the remote sensing image interpretation model composed of the multi-modal remote sensing information extraction intelligent large model, the multi-modal remote sensing information fusion module and the multi-modal decoder based on the training set composed of multiple samples. When the preset training convergence condition (the number of times reaches the upper limit or the loss function value converges) is met, the interpretation result of the input target image is obtained based on the trained remote sensing image interpretation model. The input target image includes a visible light image and a thermal infrared image of the same target object.
[0092] Embodiments
[0093] The large model driven high-resolution remote sensing image interpretation method provided in this embodiment includes the following steps:
[0094] Step S1, a large-scale multi-modal high-resolution remote sensing image dataset (including visible light images and thermal infrared images) is constructed, and the image data therein is preprocessed to construct a training set and a test set.
[0095] In this embodiment, the shape of the visible light image and the thermal infrared image is 3xHxW, where H and W represent the height and width of the image.
[0096] In this embodiment, the image preprocessing is specifically set as:
[0097] The resolution of all visible light images and thermal infrared images is adjusted to 2000x2000.
[0098] All visible light images and thermal infrared images are respectively cropped into 16 sub-images with a resolution of 512x512 from top to bottom and from left to right, with an overlap rate of 10%. In addition, the original image with a resolution of 2000x2000 is also adjusted to a resolution of 512x512. The 17 images are considered as a complete modal data sample. The above processed samples are divided into a training set and a test set in a 1:1 ratio, which are respectively used for network training and testing.
[0099] Step S2, a multi-modal remote sensing information intelligent extraction large model is constructed to realize the extraction of multi-modal remote sensing information. The multi-modal remote sensing information intelligent extraction large model is divided into a visible light information encoder and a thermal infrared information encoder, which respectively process visible light images and thermal infrared images. In this example, the dimensions of the visible light and thermal infrared input images are 3x512x512. The multi-modal remote sensing information intelligent extraction large model is composed of a visible light information encoder and a thermal infrared information encoder, and the structures of the two encoders are consistent. Each encoder is composed of 17 global high-order information embedding Transformer branches. The following describes the processing process of the visible light image.
[0100] For each independent global high-order information embedding Transformer branch, the input image first passes through the embedding layer to obtain a tensor of dimension 32×32×768. This tensor is then processed continuously through a four-stage multi-head self-attention mechanism, with the input tensor dimensions remaining consistent before and after each stage. Each stage consists of eight multi-head self-attentions. Finally, each independent global high-order information embedding Transformer branch outputs a tensor of dimension 32×32×768. These 10 tensors are summed pixel by pixel to obtain the final visible light encoding feature tensor of dimension 32×32×768.
[0101] In each global high-order information embedding Transformer module of the global high-order information embedding Transformer branch, this embodiment will query the weight matrix W for the Q, K and V vectors (the input vectors of the module are respectively Q , key weight matrix W K , value weight matrix W V Linear mapping is obtained) to mine and embed high-order information. Taking feature Q as an example to describe the processing process: First, feature Q will reduce the number of channels through 1×1 convolution to obtain a new feature Q / , the dimension is 256×32×32, where 256 is the number of channels after reduction. Then, calculate the feature Q / The correlation between the channels in the matrix M is obtained, and the dimension is 256×256. In this embodiment, it is obtained by calculating the cosine similarity between the feature maps. Next, the feature Q / The transformation dimension is (32×32)×256 and matrix M is multiplied to obtain the feature S Q , the dimension is (32×32)×256. Finally, the feature S Q After a global flat pooling and 1×1 convolution, a new feature with a dimension of 768×1×1 is obtained. This feature is multiplied with the input feature Q to obtain the output Q n Output feature Q n Keep the dimension consistent with the input feature Q.
[0102] In addition, the output of each stage of the 17 parallel branches is subjected to cross-region information interaction and perception through a cross-region information joint intelligent perception module, and the output of the module is used as the input of the next stage of each branch. In the cross-region information joint intelligent perception module, the 32x32x768 dimension tensor from the 17 global high-order information embedding Transformer branches is first subjected to global average pooling and global maximum pooling, respectively, and then pixel-by-pixel addition is performed to obtain 17 1x1x768 dimension vectors. Then, the 17 1x1x768 dimension vectors are pixel-by-pixel added and passed through 3 fully connected layers and ReLU activation functions to obtain a new vector S, which also has a dimension of 1x1x768. The above process is expressed as follows:
[0103]
[0104] where f represents the fully connected layer, GAP represents the global average pooling, F i represents the output feature of the i-th global high-order information embedding Transformer branch.
[0105] Next, the output feature map of each branch at the current stage is point multiplied with the global vector S to obtain a new feature map as the input of the next stage of each branch. The above process is expressed as follows:
[0106] V i = F i ⊙S + F i
[0107] where V i represents the input feature of the i-th global high-order information embedding visual Transformer branch at the next stage, and represents the point multiplication operation.
[0108] Step S3, the encoded features from the visible light information encoder and the encoded features from the thermal infrared information encoder are subjected to multi-modal remote sensing information fusion through a multi-modal remote sensing information fusion module. The multi-modal remote sensing information fusion module is composed of 12 bidirectional interaction attention modules. The dimensions of all tensors remain unchanged before and after passing through the bidirectional interaction attention modules. In the bidirectional interaction attention module, two groups of self-attention calculations are performed. One group uses visible light encoded features as Value, with a dimension of 32x32x768, and uses thermal infrared encoded features as Query and Key, both with a dimension of 32x32x768; the other group uses thermal infrared encoded features as Value, with a dimension of 32x32x768, and uses visible light encoded features as Query and Key, both with a dimension of 32x32x768. After each group of self-attention calculations is completed, the two 32x32x768 dimension tensors are element-by-element added to produce the final output, which has a dimension of 32x32x768.
[0109] In step S4, the high-resolution remote sensing image is interpreted based on the encoder of the BERT model. The tensor dimension after the multi-modal feature fusion is 32x32x768, and the dimension of the tensor is changed to 1024x768 through a dimension change operation. First, the tensor is added element by element with the Embending layer and the position encoding vector to obtain the input of the BERT encoder. Then, the vector passes through 12 consecutive encoding layers of the Transformer, and each layer is composed of 12 multi-head self-attention mechanisms. The dimension of the vector remains 1024x768 when passing through the 12 consecutive encoding layers of the Transformer. Finally, the softmax function is used to predict the belonging word of each Token to generate the final interpretation result.
[0110] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
[0111] The above only describes some embodiments of the present application. Those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A large model driven high-resolution remote sensing image interpretation method, characterized in that, Comprising the following steps: Step 1, constructing a large-scale multi-modal high-resolution remote sensing image dataset comprising visible light images and thermal infrared images; and image pre-processing the image data in the multi-modal high-resolution remote sensing image dataset to obtain a training set for model training; Wherein, the image pre-processing includes image data format unification, image size normalization and image cropping of each modality; wherein the size and number of each image cropped sub-image are consistent, and then the cropped image is scaled to the size of the cropped sub-image; for each modality, all cropped sub-images and scaled cropped images to the size of the cropped sub-image are regarded as a sample; Step 2, constructing a multi-modal remote sensing information extraction intelligent large model for multi-modal remote sensing information coding; Wherein, the multi-modal remote sensing information extraction intelligent large model comprises a visible light information encoder and a thermal infrared information encoder, the visible light information encoder is used to extract visible light coding features of visible light images, and the thermal infrared information encoder is used to extract thermal infrared coding features of thermal infrared images; Step 3, constructing a multi-modal remote sensing information fusion module for multi-modal feature fusion of visible light coding features and thermal infrared coding features, and outputting the fused multi-modal features; Step 4, constructing a multi-modal decoder for high-resolution remote sensing image interpretation based on the network structure of the encoder of the large language model, taking the fused multi-modal features as the input of the multi-modal decoder, and obtaining the interpretation result of the high-resolution remote sensing image based on the output thereof; Step 5, learning and training the model parameters of the remote sensing image interpretation model composed of the multi-modal remote sensing information extraction intelligent large model, the multi-modal remote sensing information fusion module and the multi-modal decoder based on the training set composed of multiple samples, and obtaining the interpretation result of the input target image based on the trained remote sensing image interpretation model when the preset training convergence condition is met; Wherein, in step 2, the visible light information encoder and the thermal infrared information encoder in the multi-modal remote sensing information extraction intelligent large model have the same structure, and each comprises n parallel global high-order information embedding Transformer branches, wherein n is equal to the number of cropped sub-images of each image plus 1; Wherein, each global high-order information embedding Transformer branch comprises m cascaded global high-order information embedding Transformer modules, and m is an integer greater than 2; each global high-order information embedding Transformer module corresponds to a stage; The output feature maps of the global high-order information embedding Transformer modules of the same order in the n branches are subjected to information joint perception through a cross-region information joint intelligent perception module, and the information joint perception features of each branch are output, and then output to the next-order global high-order information embedding Transformer module of the corresponding branch; The global high-order information embedding Transformer module is specifically set as: The input vectors of the global high-order information embedding Transformer module are respectively mapped by a query weight matrix , a key weight matrix , and a value weight matrix to a query vector Q, a key vector K, and a value vector V. Covariance pooling operation is performed on the query vector Q, the key vector K and the value vector V respectively, and then the obtained pooled vectors and the corresponding vectors before pooling are multiplied to obtain new vectors , and ; The vectors , and are input into the multi-head self-attention module to perform self-attention mechanism calculation, and the output of the multi-head self-attention module is obtained , wherein, represents a softmax activation function, represents the dimension of the vector . The output of the multi-head self-attention module is subjected to feature superposition and normalization operation on the input vector of the global high-order information embedding Transformer module through the first superposition & normalization layer, to obtain first self-attention fusion features; The first self-attention fusion features are sent into a fully connected layer, and the output of the fully connected layer is subjected to feature superposition and normalization operation on the first self-attention fusion features through the second superposition & normalization layer, to obtain second self-attention fusion features, which are output features of the global high-order information embedding Transformer module; The specific execution process of the cross-region information joint intelligent perception module includes: (1) The feature maps of each global high-order information embedding Transformer branch are subjected to global average pooling and global maximum pooling respectively, and then the two pooled results are added element by element to obtain the first feature vector of each branch; (2) The first feature vectors of all branches are added element by element, and then a global feature vector is obtained through a plurality of fully connected layers and a ReLU activation function layer; (3) The feature maps of each global high-order information embedding Transformer branch are respectively multiplied with the global feature vector to obtain information joint perception features of each branch, which are used as input features of the global high-order information embedding Transformer module in the next stage of the corresponding branch.
2. The method of claim 1, wherein, In step 1, image cropping includes: based on a specified overlap rate, image cropping is performed on the image with a specified cropping sequence after image size normalization.
3. The method of claim 1, wherein, The covariance pooling operation is specifically: respectively, through a convolution kernel of 1 1, the convolution layer of the vector 、 and The channel number reduction operation is performed to obtain new feature vectors 、 and , define , which represents the reduced channel number; Calculate the correlation between the channels in each feature vector , and respectively, to obtain a correlation matrix with a dimension of , , ; wherein each correlation matrix is obtained by calculating the cosine similarity between the feature maps; The new feature vector 、 and The corresponding correlation matrix 、 、 Perform matrix multiplication to obtain the eigenvector 、 、 ; The feature vectors , , are obtained by a global average pooling layer and a convolution layer with a kernel of 1 1, respectively , , ; The eigenvectors , , are respectively dot-multiplied with the corresponding vectors , and to obtain new vectors , and .
4. The method of claim 1, wherein, In the multi-modal remote sensing information fusion module, the visible light encoded features and the thermal infrared features are fused through the cross multi-head self-attention mechanism.
5. The method of claim 1, wherein, The multi-modal remote sensing information fusion module includes a plurality of multi-head self-attention mechanism layers, and in each multi-head self-attention mechanism layer, two groups of parallel self-attention calculations are performed, one group of which takes the visible light encoded features as the value vector and the thermal infrared encoded features as the query vector and the key vector; the other group takes the thermal infrared encoded features as the value vector and the visible light encoded features as the query vector and the key vector; and after each group of self-attention calculation is completed, the output feature maps of the two groups are added element by element as the input features of the next layer of multi-head self-attention mechanism layer.
6. The method of claim 1, wherein, In step 5, the multi-modal remote sensing information extraction intelligent large model in the remote sensing image interpretation model is first trained through the random size mask self-supervised learning strategy, specifically including: One of the visible light information encoder and the thermal infrared information encoder is selected as a training object, and a general decoder is connected to the training object, which is used to reconstruct the input image of the training object; Among the specified multiple image scales, a random one is selected as the size of the current processing image; On the image of this size, a certain proportion of image blocks are randomly selected for mask processing, and the mask matrix M used for mask processing cannot exceed the size of the current processing image; The loss function of the random size mask self-supervised learning strategy of the training object is set as: where H, W, C represent the height, width and channel number of the input image, respectively; denotes a mask matrix the total number of pixels in the input image that are not masked in the mask matrix, and denote the pixel values of the original image and the reconstructed image at image position , respectively, where the original image is the input image of the training object; When the preset training convergence condition is met, the trained training object is used to obtain a trained visible light information encoder and a trained thermal infrared information encoder in a parameter sharing manner.
7. The method of claim 1, wherein, In step 5, the pre-trained multi-modal remote sensing information extraction intelligent big model is used to train the weight parameters of the multi-modal remote sensing information fusion module and the multi-modal decoder in the remote sensing image interpretation model.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on multi-scale attention fusion
CN113283435A
Land cover remote sensing monitoring method based on multi-source feature fusion
CN115527123A