Decoupling multi-modal remote sensing image element extraction method, system and device based on large model and storage medium
By adopting the parallel training method of twin neural network architecture in the remote sensing large model, a multimodal remote sensing image element extraction model is constructed, which solves the problem of insufficient flexibility and generalization ability of remote sensing image extraction in the existing technology, and achieves efficient and accurate multimodal remote sensing image element extraction.
Patent Information
- Application Number
- CN202510309491.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to achieve flexible extraction of single-modal and multi-modal remote sensing images, and the sample size of multi-modal deep learning algorithms is insufficient and the model generalization ability is poor.
Using a decoupled multimodal remote sensing image feature extraction method based on remote sensing large model, the backbone of the first remote sensing model and the second remote sensing model is trained in parallel through the twin neural network architecture, a feature fusion module is built to perform multimodal feature fusion, and a decoupled feature extraction is achieved using an independent parallel deep learning decoder.
The generalization ability of the model is improved, and the flexible extraction of single-modal and multi-modal remote sensing images is realized, the interference of unimportant feature information is suppressed, and the accuracy and efficiency of multi-modal remote sensing image elements are improved.
Smart Images

Figure CN120164112A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the application in the field of joint interpretation of remote sensing images of multiple payloads, and specifically relates to a decoupled multi-modal remote sensing image element extraction method based on remote sensing large model technology. Background Art
[0002] With the development of human earth observation technology, more and more remote sensing satellites have been launched, constituting a huge earth observation system with multiple modalities and multiple resolutions. Different types of remote sensing satellites have different advantages and limitations. The Jilin-1 sub-meter remote sensing satellite constellation has the advantages of high spatial resolution and high temporal resolution, but lacks rich spectral information, which brings difficulties to the judgment of land type attributes. The Sentinel-2 satellite has rich spectral information, but the resolution is not high and it lacks ground object details, which brings difficulties to the accurate segmentation of ground object boundaries. In short, the current technology is still difficult to mass-produce remote sensing satellites that simultaneously possess multi-spectral and sub-meter high-resolution characteristics. Therefore, combining the advantages of different remote sensing satellite images and performing joint interpretation of multi-modal remote sensing images is of great significance for improving the accuracy of ground object extraction from remote sensing images.
[0003] In recent years, the rapid development of deep learning technology has greatly promoted the application in the field of remote sensing interpretation. A large number of excellent interpretation algorithms based on single remote sensing images have been proposed. Based on this, some scholars fuse multi-modal remote sensing images and then use deep learning algorithms for element extraction, or train different deep learning models for different modal remote sensing images and fuse the model results. Further, some scholars use different deep learning backbones to extract features of different modal remote sensing images respectively, and input these features into a deep learning decoder after fusion to obtain the extraction results of multi-modal remote sensing image elements. However, the above methods first lack flexibility. The reality is that a large number of regions lack multi-modal remote sensing images, resulting in the input not meeting the requirements of multi-modal models. Secondly, the production of multi-modal aligned remote sensing images with labels is very difficult, resulting in insufficient sample size of multi-modal deep learning algorithms and poor generalization ability of the models.
[0004] Currently, a mainstream method for extracting elements from multi-modal remote sensing images comes from the method for extracting elements from single-modal remote sensing images based on deep learning. To meet the input of single remote sensing images, it is necessary to first use a remote sensing image fusion method to fuse multi-modal remote sensing images into single-modal remote sensing images. However, the fusion method itself will cause some loss of remote sensing image information and the process is cumbersome. Another way is to train multiple single-modal remote sensing image element extraction models, and then merge the extraction results using artificial rules. This is essentially still a single-modal model and it is difficult to achieve end-to-end element extraction.
[0005] Another mainstream method for extracting multi-modal remote sensing image elements utilizes a Siamese neural network architecture. After the feature interaction and fusion of multi-modal remote sensing images, decoding is performed to obtain the element results. However, simple feature interaction and fusion may cause unimportant feature information to disrupt the learning of the network.
[0006] In addition, the common problems of the two mainstream methods are the lack of flexibility, poor model generalization due to the lack of large-scale labeled data, and the extraction of elements requires complete multi-modal remote sensing images.
[0007] In the prior art, Chinese patent document CN117372885A discloses a "method and system for multi-modal remote sensing data change detection based on Siamese U-Net neural network". The original data of at least two remote sensing images in different time phases in the target area are respectively preprocessed to obtain multi-modal target data corresponding to the original data. Based on the multi-modal target data and using a pre-trained Siamese neural network model, the surface changes in remote sensing images in different time phases in the target area are detected. Among them, the Siamese neural network model adopts two U-Net neural network structures with the same structure. By fusing the two-dimensional RGB image and three-dimensional geometric features in the remote sensing image, the representation ability of the extracted features is improved. Using the Siamese network model and adding an attention mechanism after the feature map difference operation, the model is more sensitive to real changes, suppresses the interference of irrelevant changes and noise information, and improves the change detection accuracy and efficiency. However, this technical solution is only applicable to the extraction of multi-modal remote sensing images, lacks flexibility, and has poor model generalization due to the lack of large-scale labeled data.
[0008] In summary, the prior art lacks a method for flexibly extracting single-modal and multi-modal remote sensing images. To solve the limitations of the multi-modal remote sensing image ground object element extraction model, the present patent proposes a decoupled multi-modal remote sensing image element extraction method, system, device, and storage medium based on a large model. Summary of the Invention
[0009] The present invention solves the problem that the prior art lacks a method for flexibly extracting single-modal and multi-modal remote sensing images.
[0010] A decoupled multi-modal remote sensing image element extraction method based on a large model according to the present invention includes the following steps:
[0011] Step 1: Using the first remote sensing image and the second remote sensing image, based on the pre-training technology of the remote sensing large model and combined with the contrast learning algorithm, train the first remote sensing model and the second remote sensing model respectively;
[0012] Step 2: Based on the first remote sensing model and the second remote sensing model described in Step 1, construct a multi-modal remote sensing image element extraction model;
[0013] The multi-modal remote sensing image element extraction model mentioned above is specifically as follows:
[0014] Using a siamese neural network architecture, the backbones of the first remote sensing model and the second remote sensing model obtained by training in step 1 are made parallel and used as the first feature extractor and the second feature extractor of the multi-modal remote sensing image element extraction model, and a feature fusion module is constructed to fuse the multi-modal remote sensing image features output by the first feature extractor and the second feature extractor. Finally, parallel encoders are used to decode the prediction results of different remote sensing image elements respectively;
[0015] Step 3: Fine-tune and train the multi-modal remote sensing image element extraction model described in step 2 to obtain a trained multi-modal remote sensing image element extraction model;
[0016] Step 4: Obtain the remote sensing image to be processed, input it into the trained multi-modal remote sensing image element extraction model, and extract the remote sensing image elements.
[0017] Furthermore, in the embodiment of the present invention, the first remote sensing image in step 1 is a sub-meter resolution remote sensing image, and the second remote sensing image is a multi-spectral remote sensing image.
[0018] Furthermore, in the embodiment of the present invention, the remote sensing large model pre-training technology in step 1 is specifically as follows:
[0019] The first remote sensing image after the first data augmentation is input into the teacher branch encoder. After the first remote sensing image after the second data augmentation is randomly masked, it is input into the student branch encoder. Feature calculations are respectively performed on the first remote sensing image after the first data augmentation passing through the teacher branch encoder and the first remote sensing image after the second data augmentation passing through the student branch encoder, and the teacher branch feature map and the student branch feature map corresponding to the first remote sensing image are respectively output;
[0020] Calculate the loss value of the remote sensing large model based on the teacher branch feature map and the student branch feature map corresponding to the first remote sensing image;
[0021] L1 = ∑ K p t logp s ;
[0022] L2 = -∑ n ∑ i p ti logp si ;
[0023] L = L1 + L2;
[0024] In the formula, L represents the loss value, L1 represents the image-level contrast loss, L2 represents the patch-level contrast loss, K represents the number of student branch feature maps, pt Denote the classification code of the teacher branch feature map, p s Denote the classification code of the student branch feature map, n denotes the number of teacher branch feature maps, i denotes the patch index of the mask mark, p ti Denote the code of the patch at index i of the teacher branch feature map, p si Denote the code of the patch at index i of the student branch feature map;
[0025] Update the weights of the remote sensing large model based on the loss value of the remote sensing large model to obtain the first remote sensing model;
[0026] The second remote sensing image repeats the operation of the first remote sensing image to obtain the second remote sensing model.
[0027] Furthermore, in the embodiment of the present invention, the fine-tuning training in step 3 is specifically as follows:
[0028] Step 301, update the weights of the first remote sensing model and the second remote sensing model in step 1 respectively, and fine-tune the first feature extractor and the second feature extractor based on the updated weights;
[0029] Step 302, after the first remote sensing image passes through the first feature extraction module described in step 301, obtain multi-layer encoded features, input the multi-layer encoded features into the multi-scale feature calculation module for dimension transformation, and output a multi-scale feature map; the second remote sensing image passes through the second feature extraction module described in step 301, performs the same operation as the first remote sensing image, and outputs a multi-scale feature map corresponding to its scale;
[0030] Step 303, input the multi-scale feature map corresponding to the first remote sensing image and the multi-scale feature map corresponding to the second remote sensing image and having the same scale as the multi-scale feature map corresponding to the first remote sensing image into the feature fusion module for multi-modal feature fusion, and output a multi-scale fusion feature map;
[0031] Step 304, input the multi-scale feature maps respectively output by the first feature extraction module and the second feature extraction module and the multi-scale fusion feature map output by the feature fusion module into their corresponding decoders respectively, and output the remote sensing image element prediction results corresponding to them;
[0032] Step 305, calculate the loss value corresponding to it based on the remote sensing image element prediction result corresponding to the corresponding decoder, and fine-tune the weights of the multi-modal remote sensing image element extraction model based on the loss value corresponding to it.
[0033] Furthermore, in the embodiment of the present invention, the feature fusion module in step 303 is specifically as follows:
[0034] After subtracting the second remote sensing image feature map corresponding to its scale from the first remote sensing image feature map, perform max pooling to obtain a feature vector. Then, use a fully connected neural network for mapping to obtain the weights of the features, and perform attention weighting on the first remote sensing image feature map to obtain the weighted first remote sensing image feature map. After performing the same operation on the second remote sensing image feature map corresponding to its scale, obtain the weighted second remote sensing image feature map;
[0035] Subsequently, concatenate the weighted first remote sensing image feature map and the weighted second remote sensing image feature map in the channel dimension to obtain dimension features, and use convolution operations for channel fusion to obtain the fused dimension features. Finally, perform average pooling to obtain a feature vector, and use parallel fully connected neural networks for mapping respectively to obtain the weights of the two features, which are respectively used for attention weighting with the weighted first remote sensing image feature map and the weighted second remote sensing image feature map to obtain the first feature map. This module repeats this operation to output the second feature map, and add the first feature map and the second feature map to obtain the fused remote sensing image element prediction result.
[0036] Furthermore, in the embodiment of the present invention, in step 305, based on the corresponding remote sensing image element prediction result output by the corresponding decoder, calculate the corresponding loss value, specifically:
[0037] Perform element annotation on the multi-modal remote sensing image formed by fusing the first remote sensing image and the second remote sensing image, use the multi-modal remote sensing image with element annotation for supervised learning, and calculate the corresponding loss values by respectively comparing the remote sensing image element prediction results obtained by the corresponding decoder with the true element annotation values.
[0038] Furthermore, in the embodiment of the present invention, for the multi-modal remote sensing image element extraction model in step 4, when there is only the first remote sensing image, use the corresponding first feature extractor, multi-scale feature calculation module, and decoder for remote sensing image element extraction;
[0039] Or when there is only the second remote sensing image, use the corresponding second feature extractor, multi-scale feature calculation module, and decoder for remote sensing image element extraction;
[0040] When there are both the first remote sensing image and the second remote sensing image, use the multi-modal remote sensing image element extraction model for remote sensing image element extraction.
[0041] A decoupled multi-modal remote sensing image element extraction system based on a large model according to the present invention includes the following modules:
[0042] Module 1: Using the first remote sensing image and the second remote sensing image, based on the pre-training technology of the remote sensing large model and combined with the contrast learning algorithm, train the first remote sensing model and the second remote sensing model respectively;
[0043] Module 2: Based on the first remote sensing model and the second remote sensing model described in Module 1, construct a multi-modal remote sensing image element extraction model;
[0044] The multi-modal remote sensing image element extraction model described above is specifically:
[0045] Using the siamese neural network architecture, parallelize the backbones of the first remote sensing model and the second remote sensing model separately trained in Module 1, and use them as the first feature extractor and the second feature extractor of the multi-modal remote sensing image element extraction model. Then construct a feature fusion module to fuse the multi-modal remote sensing image features output by the first feature extractor and the second feature extractor. Finally, use parallel encoders to decode the prediction results of different remote sensing image elements respectively;
[0046] Module 3: Fine-tune and train the multi-modal remote sensing image element extraction model described in Module 2 to obtain a trained multi-modal remote sensing image element extraction model;
[0047] Module 4: Obtain the remote sensing image to be processed, input it into the trained multi-modal remote sensing image element extraction model, and extract the remote sensing image elements.
[0048] An electronic device according to the present invention includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0049] The memory is used to store computer programs;
[0050] The processor is used to implement the method steps described in any of the above when executing the programs stored on the memory.
[0051] A computer-readable storage medium according to the present invention stores a computer program therein, and when the computer program is executed by a processor, it implements the method steps described in any of the above.
[0052] The present invention solves the problem that the prior art lacks a method for flexibly extracting single-modal and multi-modal remote sensing images. The specific beneficial effects include:
[0053] A decoupled multi-modal remote sensing image element extraction method based on a large model according to the present invention uses remote sensing large model technology to improve the generalization ability of the model, and uses independent and parallel deep learning decoders to flexibly extract single-modal and multi-modal remote sensing images. At the same time, an attention feature fusion mechanism is introduced to suppress the interference of unimportant feature information, so as to achieve efficient and accurate extraction of multi-modal remote sensing image elements. Description of the Drawings
[0054] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, wherein:
[0055] Figure 1 is the framework of the decoupled multi-modal remote sensing image element extraction method based on a large model described in Embodiment 1;
[0056] Figure 2 is the contrastive learning algorithm for remote sensing large models described in Embodiment 1;
[0057] Figure 3 is the fine-tuning model architecture diagram described in Embodiment 1;
[0058] Figure 4 is the partitioned parallel multi-GPU accelerated inference architecture diagram described in Embodiment 1;
[0059] Figure 5 is the crop extraction result map of Chuanying District, Jilin City described in Embodiment 1;
[0060] Figure 6 is the crop extraction result map of Zuojia Town, Changyi District, Jilin City described in Embodiment 1;
[0061] Figure 7 is the crop extraction result map of Wangben Town, Shuangliao City described in Embodiment 1. Specific Embodiments
[0062] The following will clearly and completely describe various embodiments of the present invention in conjunction with the accompanying drawings. The embodiments described by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0063] Embodiment 1. A decoupled multi-modal remote sensing image element extraction method, system, device, and storage medium based on a large model described in this embodiment include the following steps:
[0064] Step 1: Using the first remote sensing image and the second remote sensing image, based on the remote sensing large model pre-training technology, and in combination with the contrastive learning algorithm, respectively train to obtain the first remote sensing model and the second remote sensing model;
[0065] Step 2: Based on the first remote sensing model and the second remote sensing model described in Step 1, construct a multi-modal remote sensing image element extraction model;
[0066] The multi-modal remote sensing image element extraction model is specifically:
[0067] Using a twin neural network architecture, the backbones of the first remote sensing model and the second remote sensing model obtained by separately training in step 1 are made parallel and serve as the first feature extractor and the second feature extractor of the multi-modal remote sensing image element extraction model. A feature fusion module is constructed to fuse the multi-modal remote sensing image features output by the first feature extractor and the second feature extractor. Finally, parallel encoders are used to separately decode the prediction results of different remote sensing image elements;
[0068] Step 3: Fine-tune and train the multi-modal remote sensing image element extraction model described in step 2 to obtain a trained multi-modal remote sensing image element extraction model;
[0069] Step 4: Obtain the remote sensing image to be processed, input it into the trained multi-modal remote sensing image element extraction model, and extract the remote sensing image elements.
[0070] In this embodiment, the first remote sensing image in step 1 is a sub-meter resolution remote sensing image, and the second remote sensing image is a multi-spectral remote sensing image.
[0071] In this embodiment, the remote sensing large model pre-training technology in step 1 is specifically as follows:
[0072] The first remote sensing image after the first data augmentation is input into the teacher branch encoder. After the first remote sensing image after the second data augmentation is randomly masked, it is input into the student branch encoder. Feature calculations are respectively performed on the first remote sensing image after the first data augmentation passing through the teacher branch encoder and the first remote sensing image after the second data augmentation passing through the student branch encoder, and the teacher branch feature map and the student branch feature map corresponding to the first remote sensing image are respectively output;
[0073] Calculate the loss value of the remote sensing large model based on the teacher branch feature map and the student branch feature map corresponding to the first remote sensing image;
[0074] L1 = ∑ K p t logp s ;
[0075] L2 = -∑ n ∑ i p ti logp si ;
[0076] L = L1 + L2;
[0077] In the formula, L represents the loss value, L1 represents the image-level contrast loss, L2 represents the patch-level contrast loss, K represents the number of student branch feature maps, p t represents the classification code of the teacher branch feature map, p sRepresents the classification code of the student branch feature map, n represents the number of teacher branch feature maps, i represents the patch index of the mask mark, and p ti Represents the code of the patch at index i of the teacher branch feature map, and p si Represents the code of the patch at index i of the student branch feature map;
[0078] Update the weights of the remote sensing large model based on the loss value of the remote sensing large model to obtain the first remote sensing model;
[0079] The second remote sensing image repeats the operation of the first remote sensing image to obtain the second remote sensing model.
[0080] In this embodiment, the fine-tuning training in step 3 is specifically as follows:
[0081] Step 301: Update the weights of the first remote sensing model and the second remote sensing model in step 1 respectively, and fine-tune the first feature extractor and the second feature extractor based on the updated weights;
[0082] Step 302: After the first remote sensing image passes through the first feature extraction module described in step 301, multi-layer encoded features are obtained. The multi-layer encoded features are input into the multi-scale feature calculation module for dimension transformation, and a multi-scale feature map is output; the second remote sensing image passes through the second feature extraction module described in step 301, performs the same operation as the first remote sensing image, and outputs a multi-scale feature map corresponding to its scale;
[0083] Step 303: Input the multi-scale feature map corresponding to the first remote sensing image and the multi-scale feature map corresponding to the second remote sensing image with the same scale as the multi-scale feature map corresponding to the first remote sensing image into the feature fusion module for multi-modal feature fusion, and output a multi-scale fusion feature map;
[0084] Step 304: The multi-scale feature maps respectively output by the first feature extraction module and the second feature extraction module, and the multi-scale fusion feature map output by the feature fusion module are respectively input into the corresponding decoders, and the remote sensing image element prediction results corresponding to them are output;
[0085] Step 305: Calculate the loss value corresponding to it based on the remote sensing image element prediction result corresponding to the corresponding decoder, and fine-tune the weights of the multi-modal remote sensing image element extraction model based on the loss value corresponding to it.
[0086] In this embodiment, the feature fusion module in step 303 is specifically as follows:
[0087] After subtracting the second remote sensing image feature map corresponding to its scale from the first remote sensing image feature map, perform max pooling to obtain a feature vector. Then, use a fully connected neural network for mapping to obtain the weights of the features, and perform attention weighting on the first remote sensing image feature map to obtain the weighted first remote sensing image feature map. After performing the same operation on the second remote sensing image feature map corresponding to its scale, obtain the weighted second remote sensing image feature map;
[0088] Subsequently, concatenate the weighted first remote sensing image feature map and the weighted second remote sensing image feature map in the channel dimension to obtain a dimensional feature, and use a convolution operation for channel fusion to obtain the fused dimensional feature. Finally, perform average pooling to obtain a feature vector, and use parallel fully connected neural networks for mapping respectively to obtain the weights of the two features, which are respectively used for attention weighting with the weighted first remote sensing image feature map and the weighted second remote sensing image feature map to obtain the first feature map. This module repeats this operation to output the second feature map, and add the first feature map and the second feature map to obtain the fused remote sensing image element prediction result.
[0089] In this embodiment, for the remote sensing image element prediction result corresponding to the output of the corresponding decoder in step 305, calculate the corresponding loss value, specifically:
[0090] Perform element annotation on the multi-modal remote sensing image formed by fusing the first remote sensing image and the second remote sensing image, use the multi-modal remote sensing image with element annotation for supervised learning, and calculate the corresponding loss values by respectively comparing the remote sensing image element prediction results obtained by the corresponding decoder with the true element annotation values.
[0091] In this embodiment, for the multi-modal remote sensing image element extraction model in step 4, when there is only the first remote sensing image, use the corresponding first feature extractor, multi-scale feature calculation module and decoder for remote sensing image element extraction;
[0092] Or when there is only the second remote sensing image, use the corresponding second feature extractor, multi-scale feature calculation module and decoder for remote sensing image element extraction;
[0093] When there are both the first remote sensing image and the second remote sensing image, use the multi-modal remote sensing image element extraction model for remote sensing image element extraction.
[0094] Such as Figure 1As shown in the figure, this embodiment proposes a decoupled multi-modal remote sensing image element extraction method based on a large model. First, using remote sensing large model pre-training technology, a feature extraction model for Jilin-1 high-resolution sub-meter remote sensing images and a feature extraction model for Sentinel-2 multi-spectral remote sensing images are trained. Then, on the basis of pre-training, supervised learning is carried out. The twin neural network architecture is used to extract and fuse the features of multi-modal remote sensing images, and three independent and parallel deep learning decoders are used to achieve decoupled high-resolution remote sensing image element extraction, Sentinel-2 multi-spectral remote sensing image element extraction, and multi-modal remote sensing image element extraction. A method framework realizes flexible element extraction of single-modal and multi-modal remote sensing images. Finally, using partitioned parallel multi-GPU acceleration, the element extraction of large-scale remote sensing images is realized.
[0095] The method framework mainly includes three stages: ① Pre-training stage: Collect millions of Jilin-1 high-resolution remote sensing image samples and Sentinel-2 multi-spectral remote sensing image samples, and use contrastive learning technology to train a large remote sensing image feature extraction model; ② Fine-tuning stage: Use multi-modal remote sensing images with element annotations for supervised learning. The fine-tuning architecture uses a twin neural network. The initial parameters of the backbone of feature extraction come from the pre-trained large remote sensing image feature extraction model. The fusion module combines the attention mechanism to fuse the features of different modal remote sensing images. The decoding modules are independent of each other, and the sum of the losses of different decoding modules is used as the overall loss to optimize the network parameters; ③ Inference stage: Use the fine-tuned model and combine the partitioned parallel multi-GPU acceleration method to realize the extraction of multi-modal remote sensing image elements in a large area. At the same time, this method can use some branches alone to realize the extraction of single-modal remote sensing image elements, which is applicable to areas lacking multi-modal remote sensing images.
[0096] Specifically, a decoupled multi-modal remote sensing image element extraction method based on a large model proposed in this embodiment includes the following steps:
[0097] Step 1: Pre-training of the remote sensing large model
[0098] Using a large amount of Jilin-1 high-resolution remote sensing images and Sentinel-2 multi-spectral remote sensing images in combination with the contrastive learning algorithm, a large remote sensing model is constructed. The algorithm architecture is as Figure 2 shown, specifically:
[0099] (1) Data augmentation
[0100] For the input single-modal remote sensing image, the following data augmentation methods are used: ① Random scale scaling, with the scaling ratio being randomly selected from 0.8 to 1.2 of the original input; ② Random cropping, with the cropping area being randomly selected from 0.6 to 1.2 of the input, and the height-width ratio being randomly selected from 0.75 to 1.33; ③ Random rotation, with the rotation angle being randomly selected from 0 to 90 degrees; ④ Random color perturbation, with a 50% probability of adjusting the input brightness and contrast. The four data augmentation methods are applied for two enhancement transformations. After the transformation, one copy is used as the input for the teacher branch, and the other copy is used as the input for the student branch after random masking. The width and height of the mask block are 14 pixels, the masking ratio is 75%, and the mask positions are randomly distributed.
[0101] (2) Feature calculation
[0102] The feature calculation module adopts a parallel structure, enabling the simultaneous acquisition of the feature encoding outputs of the teacher branch and the student branch. The remotely sensed image after the data augmentation in (1) is input into the feature calculation module. Among them, the encoder structures of the teacher branch and the student branch are the same, both adopting the Vision Image Transformer (VIT) structure. The input remotely sensed image is split into pixel blocks with a size of 14 pixels in both height and width. Each pixel block is encoded into a 768-length vector using linear mapping, and then a position encoding of 768-length vector is added to each encoded pixel block, and a classification encoding block of 768-length vector is additionally added. Finally, the image encoding and the classification encoding are concatenated and input into the VIT model for feature encoding to obtain a feature output with the same shape as the input. Eventually, after the VIT structure, 2 image features of the teacher branch and 10 image features of the student branch are obtained.
[0103] (3) Loss calculation
[0104] Using the results obtained from (2) feature calculation, the model loss is calculated. The loss L includes the image-level contrast loss L1 and the patch-level contrast loss L2, which are specifically calculated as shown in the following equations (1), (2), and (3):
[0105] L1 = ∑ K p t log p s ; (1)
[0106] L2 = -∑ n ∑ i p ti log p si ; (2)
[0107] L = L1 + L2; (3)
[0108] In the formula, K represents the number of feature maps of the student branch, p t represents the classification encoding of the feature map of the teacher branch, ps represents the classification code of the student branch feature map, n represents the number of teacher branch feature maps, i represents the patch index of the mask mark, and p ti represents the code of the patch at index i of the teacher branch feature map, and p si represents the code of the patch at index i of the student branch feature map.
[0109] (4) Weight update
[0110] Perform backpropagation of the gradient from the loss L, and update the parameters of the student branch network model in combination with the stochastic gradient descent algorithm. After the parameters of the student branch network model are updated, use the momentum update strategy to update the parameters of the teacher branch network model. The momentum update strategy is shown in formula (4).
[0111] θ t = mθ t-1 + (1 - m)θ s ; (4)
[0112] In the formula, θ s , θ t both represent the parameters of the student branch network and the teacher branch network updated at the current moment, θ t-1 represents the parameter of the teacher branch network that has not been updated at the current moment, and m represents the weight.
[0113] Step 2: Fine-tuning:
[0114] Use the backbone of the large remote sensing model trained in Step 1 as the feature extractor for multi-modal remote sensing images, construct a feature fusion module with a fusion attention mechanism to fuse multi-modal features, and finally use parallel encoders to decode different feature outputs respectively to adapt to application inferences of various modalities. The specific fine-tuning algorithm architecture is as Figure 3 shown, specifically:
[0115] (1) Weight initialization
[0116] Use the first remote sensing model and the second remote sensing model in Step 1 to perform weight initialization fine-tuning on modules ① and ④ in the model respectively.
[0117] (2) Multi-scale feature calculation
[0118] After the Jilin-1 sample image (image size: 518×518×3) passes through Module ①, the encoded features of the 3rd, 6th, 9th, and 12th layers of the VIT model are obtained (feature dimension: 1369×768), and then input into Module ②. Using the transposed convolution structure, the encoded feature dimensions of the 3rd, 6th, 9th, and 12th layers are respectively transformed into 148×148×768, 74×74×768, 37×37×768, and 19×19×768. For the Sentinel-2 sample image (image size: 518×518×10), Modules ④ and ⑤ are used, and the calculation process is the same as above. Modules ② and ⑤ finally output four feature maps respectively, constituting a multi-scale feature map.
[0119] (3) Multi-modal Feature Fusion
[0120] The multi-scale features of the Jilin-1 sample image and the multi-scale features of the Sentinel-2 sample image are fused using Module ⑦. First, subtract the feature map of the Sentinel-2 sample image of the corresponding size from the feature map of the Jilin-1 sample image with a size of 148×148×768, and then perform max pooling on the result to obtain a feature vector of 1×1×768. Then, use a fully connected neural network for mapping to obtain the weights of the features, and perform attention weighting on the feature map of the Jilin-1 sample image. After performing the same operation on the feature map of the Sentinel-2 sample image with a size of 148×148×768, a weighted feature map is obtained. Subsequently, the two feature maps are concatenated in the channel dimension to obtain a feature of 148×148×1536 dimensions, and channel fusion is performed using a convolution operation with a kernel size of 1×1 to obtain a feature of 148×148×768 dimensions. Finally, average pooling is performed on the result to obtain a feature vector of 1×1×768, and two parallel fully connected neural networks are used for mapping respectively to obtain the weights of the two features, and attention weighting is performed on the feature maps of the Jilin-1 sample image and the Sentinel-2 sample image respectively to obtain new features. After repeating this module calculation twice, the feature maps are added to obtain a fused feature map with a size of 148×148×768. Similarly, after fusing the multi-modal feature maps with sizes of 74×74×768, 37×37×768, and 19×19×768 using Module ⑦, fused feature maps with sizes of 74×74×768, 37×37×768, and 19×19×768 are obtained. Module ⑦ finally outputs four feature maps, constituting a multi-scale fused feature map.
[0121] (4) Feature Decoding
[0122] The multi-scale feature maps output by Modules ②, ⑤, and ⑦ are respectively input into three parallel UperNet decoders to obtain three model prediction results.
[0123] (5) Model Loss
[0124] The prediction results obtained by the three decoders are respectively calculated with the true label values using the cross-entropy loss function to obtain three loss values. After adding the three loss values, backpropagation is performed, and the weights of the fine-tuning network architecture are updated using the stochastic gradient descent algorithm.
[0125] Step 3: Inference
[0126] The multi-modal remote sensing image element extraction model in Step 4 can flexibly use two methods, single-modal inference and multi-modal inference, to extract remote sensing image elements.
[0127] (1) Single-modal inference
[0128] For areas lacking Sentinel-2 multi-spectral remote sensing images, the Jilin-1 high-resolution remote sensing images can be input into modules ①②③ to obtain the interpretation results. For areas lacking Jilin-1 high-resolution remote sensing images, the Sentinel-2 multi-spectral remote sensing images can be input into modules ④⑤⑥ to obtain the interpretation results.
[0129] (2) Multi-modal inference
[0130] For areas that have both Jilin-1 high-resolution remote sensing images and Sentinel-2 multi-spectral remote sensing images, the multi-modal remote sensing images are input into modules ①②④⑤⑦⑧ to obtain the interpretation results.
[0131] (3) Partitioned parallel multi-GPU accelerated inference
[0132] The multi-modal inference mentioned above uses the method of partitioned parallel multi-GPU acceleration to achieve the extraction of large-area multi-modal remote sensing image elements. The specific algorithm architecture is as Figure 4 shown. On a 4-GPU server, first, the main process divides the remote sensing image into windows of 518×518, with 10% pixel overlap between adjacent windows. Then, according to the number of windows, all windows are evenly divided into four parts. Next, the main process creates four child processes and assigns 1 part of the windows to each of them. Subsequently, each child process schedules a single GPU to perform inference window by window according to the assigned windows. At the same time, the main process continuously writes the inference results received from the child processes to the disk and waits for all child processes to finish inference. Finally, after all child processes have finished inference, the main process writes the inference results that have not been written to the disk to the disk.
[0133] To better illustrate a decoupled multi-modal remote sensing image element extraction method based on a large model described in this embodiment, it is described in detail through the following examples:
[0134] A decouplable multi-modal remote sensing image element extraction method based on a large model according to this embodiment randomly collected 12,000 points in Jilin City, Baicheng City, and Shuangliao City. For each point, a pair of remote sensing image samples of Jilin-1 with a resolution of 0.75 meters and Sentinel-2 with a resolution of 10 meters was cropped, and the planted crops were manually labeled into five types: corn, rice, soybeans, other crops, and others. Finally, 10,000 pairs of samples were used for fine-tuning, and 2,000 pairs of samples were used for testing. The test results are shown in Table 1, and the overall extraction accuracy of the planted crops reached 98.37%.
[0135] Table 1
[0136]
[0137]
[0138] Using the fine-tuned model, the extraction of the planted crop type elements was carried out in more regions using partitioned parallel multi-GPU acceleration. The extraction results are as shown in Figure 5 、 Figure 6 and Figure 7 . This method can give full play to the advantages of multi-modal remote sensing images, has a high extraction accuracy, and has strong robustness, with high application value.
[0139] Embodiment 2. A decouplable multi-modal remote sensing image element extraction system according to this embodiment includes the following modules:
[0140] Module 1: Using the first remote sensing image and the second remote sensing image, based on the pre-training technology of the remote sensing large model and combined with the contrast learning algorithm, the first remote sensing model and the second remote sensing model are respectively trained;
[0141] Module 2: Based on the first remote sensing model and the second remote sensing model described in Module 1, a multi-modal remote sensing image element extraction model is constructed;
[0142] The multi-modal remote sensing image element extraction model described is specifically:
[0143] Using the siamese neural network architecture, the backbones of the first remote sensing model and the second remote sensing model respectively trained in Module 1 are parallelized and used as the first feature extractor and the second feature extractor of the multi-modal remote sensing image element extraction model. A feature fusion module is constructed to fuse the multi-modal remote sensing image features output by the first feature extractor and the second feature extractor, and finally the parallel encoders are used to decode the prediction results of different remote sensing image elements;
[0144] Module 3: Fine-tuning and training the multi-modal remote sensing image element extraction model described in Module 2 to obtain the trained multi-modal remote sensing image element extraction model;
[0145] Module 4: Obtain the remotely sensed image to be processed, input the trained multi-modal remotely sensed image feature extraction model, and extract the remotely sensed image features.
[0146] Embodiment 3. An electronic device according to this embodiment includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0147] The memory is used to store computer programs;
[0148] The processor is used to implement the method steps described in Embodiment 1 when executing the programs stored on the memory.
[0149] Embodiment 4. A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in Embodiment 1 are implemented.
[0150] The above has introduced in detail a method, system, device, and storage medium for extracting multi-modal remotely sensed image features based on a large model and decoupling proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A decoupled multi-modal remote sensing image element extraction method based on a large model, characterized in that: The following steps are involved: Step 1: Using the first remote sensing image and the second remote sensing image, based on the remote sensing large model pre-training technology and combined with the contrast learning algorithm, respectively train the first remote sensing model and the second remote sensing model; Step 2: Based on the first remote sensing model and the second remote sensing model described in step 1, a multimodal remote sensing image element extraction model is constructed; The multimodal remote sensing image element extraction model is specifically: Using the twin neural network architecture, the backbone of the first remote sensing model and the backbone of the second remote sensing model respectively trained in step 1 are parallelized and used as the first feature extractor and the second feature extractor of the multimodal remote sensing image element extraction model, and a feature fusion module is constructed to fuse the multimodal remote sensing image features output by the first feature extractor and the second feature extractor, and finally, parallel encoders are used to decode different remote sensing image element prediction results respectively; Step 3: fine-tune the multimodal remote sensing image feature extraction model described in step 2 to obtain a trained multimodal remote sensing image feature extraction model; Step 4: Obtain the remote sensing image to be processed, input the trained multimodal remote sensing image feature extraction model, and extract the remote sensing image features.
2. The decoupling multimodal remote sensing image element extraction method based on a large model according to claim 1 is characterized in that: The first remote sensing image in step 1 is a sub-meter resolution remote sensing image, and the second remote sensing image is a multispectral remote sensing image.
3. The decoupling multimodal remote sensing image element extraction method based on a large model according to claim 1 is characterized in that: The remote sensing large model pre-training technology in step 1 is specifically: The first remote sensing image after the first data enhancement is input into the teacher branch encoder, and the first remote sensing image after the second data enhancement is randomly masked and then input into the student branch encoder, and the features of the first remote sensing image after the first data enhancement of the teacher branch encoder and the first remote sensing image after the second data enhancement of the student branch encoder are respectively calculated, and the teacher branch feature map and the student branch feature map corresponding to the first remote sensing image are respectively output; Calculate the loss value of the remote sensing large model based on the teacher branch feature map and the student branch feature map corresponding to the first remote sensing image; L1=∑ K p t logp s ; L2=-∑ n ∑ i p ti logp si ; L = L1 + L2; Where L represents the loss value, L1 represents the image-level contrast loss, L2 represents the patch-level contrast loss, K represents the number of student branch feature maps, and p t represents the classification encoding of the teacher branch feature map, p s represents the classification encoding of the student branch feature map, n represents the number of teacher branch feature maps, i represents the patch index of the mask mark, and p ti represents the encoding of the patch at index i in the teacher branch feature map, p si Represents the encoding of the patch at index i in the student branch feature map; The weight of the remote sensing large model is updated based on the loss value of the remote sensing large model to obtain a first remote sensing model; The second remote sensing image repeats the operation of the first remote sensing image to obtain a second remote sensing model.
4. The decoupling multimodal remote sensing image element extraction method based on a large model according to claim 1 is characterized in that: The fine-tuning training in step 3 is specifically as follows: Step 301, respectively updating the weight of the first remote sensing model and the weight of the second remote sensing model in step 1, and respectively fine-tuning the first feature extractor and the second feature extractor based on the updated weights; Step 302: After the first remote sensing image passes through the first feature extraction module described in step 301, multi-layer coding features are obtained, and the multi-layer coding features are input into the multi-scale feature calculation module for dimensional transformation, and a multi-scale feature map is output; the second remote sensing image passes through the second feature extraction module described in step 301, performs the same operation as the first remote sensing image, and outputs a multi-scale feature map corresponding to its scale; Step 303, inputting the multi-scale feature map corresponding to the first remote sensing image and the multi-scale feature map corresponding to the second remote sensing image and having the same scale as the first remote sensing image into a feature fusion module for multimodal feature fusion, and outputting a multi-scale fusion feature map; Step 304, the multi-scale feature maps outputted by the first feature extraction module and the second feature extraction module, and the multi-scale fusion feature map outputted by the feature fusion module are respectively inputted into the corresponding decoders, and the corresponding remote sensing image element prediction results are outputted; Step 305, based on the corresponding remote sensing image element prediction result output by the corresponding decoder, calculate the corresponding loss value, and fine-tune the weight of the multimodal remote sensing image element extraction model based on the corresponding loss value.
5. The decoupling multimodal remote sensing image element extraction method based on a large model according to claim 4 is characterized in that: The feature fusion module in step 303 is specifically: After the first remote sensing image feature map is subtracted from the second remote sensing image feature map corresponding to its scale, the maximum pooling is performed to obtain the feature vector, and then the fully connected neural network is used for mapping to obtain the feature weight, and the first remote sensing image feature map is weighted by attention to obtain the weighted first remote sensing image feature map. The same operation is performed on the second remote sensing image feature map corresponding to its scale to obtain the weighted second remote sensing image feature map; Subsequently, the weighted first remote sensing image feature map and the weighted second remote sensing image feature map are spliced in the channel dimension to obtain dimensional features, and the convolution operation is used to perform channel fusion to obtain fused dimensional features. Finally, average pooling is performed to obtain a feature vector, and the parallel fully connected neural network is used to map them separately to obtain the weights of the two features, which are respectively weighted with the weighted first remote sensing image feature map and the weighted second remote sensing image feature map to obtain the first feature map. This module repeats this operation, outputs the second feature map, and adds the first feature map and the second feature map to obtain the fused remote sensing image element prediction result.
6. The decoupling multimodal remote sensing image element extraction method based on a large model according to claim 4 is characterized in that: In step 305, based on the corresponding remote sensing image element prediction result output by the corresponding decoder, the corresponding loss value is calculated, specifically: The multimodal remote sensing image formed by the fusion of the first remote sensing image and the second remote sensing image is annotated with elements, and supervised learning is performed using the multimodal remote sensing image with element annotations. The remote sensing image element prediction results obtained by the corresponding decoder are respectively compared with the real element annotation values to obtain the corresponding loss values.
7. The decoupling multimodal remote sensing image element extraction method based on a large model according to claim 1 is characterized in that: The multimodal remote sensing image element extraction model in step 4 uses the first feature extractor, multi-scale feature calculation module and decoder corresponding to the first remote sensing image to extract remote sensing image elements; or for only the second remote sensing image, using the second feature extractor, multi-scale feature calculation module and decoder corresponding thereto to extract remote sensing image elements; When there are both a first remote sensing image and a second remote sensing image, a multimodal remote sensing image feature extraction model is used to extract remote sensing image features.
8. A decoupled multi-modal remote sensing image element extraction system based on a large model, characterized in that: Includes the following modules: Module 1: Using the first remote sensing image and the second remote sensing image, based on the remote sensing large model pre-training technology and combined with the contrast learning algorithm, the first remote sensing model and the second remote sensing model are trained respectively; Module 2: Based on the first remote sensing model and the second remote sensing model described in Module 1, a multimodal remote sensing image feature extraction model is constructed; The multimodal remote sensing image element extraction model is specifically: Using the twin neural network architecture, the backbone of the first remote sensing model and the backbone of the second remote sensing model respectively trained by module 1 are parallelized and used as the first feature extractor and the second feature extractor of the multimodal remote sensing image element extraction model, and a feature fusion module is constructed to fuse the multimodal remote sensing image features output by the first feature extractor and the second feature extractor. Finally, parallel encoders are used to decode different remote sensing image element prediction results respectively. Module 3: Fine-tune the multimodal remote sensing image feature extraction model described in Module 2 to obtain a trained multimodal remote sensing image feature extraction model; Module 4: Obtain the remote sensing image to be processed, input the trained multimodal remote sensing image feature extraction model, and extract the remote sensing image features.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-modal remote sensing data change detection method and system based on twin U-Net neural network
CN117372885A
Cited By
Remote sensing image edge detection system and algorithm based on large model
CN121213599A
Federal remote sensing large model training method and system based on low-rank self-adaption
CN121438126A
A federated remote sensing large model training method and system based on low-rank adaptation
CN121438126B