Image text fusion-based multi-mode dual-branch twin network remote sensing change detection method and image text fusion-based multi-mode dual-branch twin network remote sensing change detection system

Through a multimodal dual-branch twin network with image text fusion, the remote sensing image features are extracted using the CLIP model and SwinTransformer, which solves the problems of low accuracy and poor robustness caused by illumination changes and land complexity in remote sensing change detection, and improves the recognition accuracy of changing areas.

CN120472214APending Publication Date: 2025-08-12NORTHWEST A & F UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510557005.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing remote sensing change detection methods have low accuracy and poor robustness under factors such as lighting changes and land objects complexity. The traditional Siamese network does not fully utilize text information, resulting in insufficient change detection accuracy.

Method used

A multimodal dual-branch twin network with image text fusion is used to obtain text information through the CLIP model, image features are extracted in combination with SwinTransformer, feature interaction and difference information calculation is used to obtain the changed image segmentation mask.

Benefits of technology

Effectively reduce the problem of missed detection of single-modal data, improve the accuracy of identifying changing areas, reduce false alarm rates, and adapt to the influence of environmental noise in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472214A_ABST
    Figure CN120472214A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and artificial intelligence, in particular to a multi-mode double-branch twin network remote sensing change detection method and system for image and text fusion, and the method comprises the steps: carrying out the interactive fusion of text features and a high-level semantic feature map, enabling the text features to find the most relevant visual clues, and enabling the visual clues to be more accurate; performing residual connection on the refined text features, and updating the text features; the updated text features are combined with the high-level semantic feature map to generate a high-level semantic feature map with higher discrimination; the method comprises the following steps of: extracting difference information among multi-level features according to multi-level feature maps of front and back time phases and a high-level semantic feature map with higher distinction, and then acquiring a changed image segmentation mask based on a cross-shaped Transform multi-mode decoder guided by a U-shaped visual language. By fusing the remote sensing image and the associated text, the problems of false detection and missing detection of single-mode data are effectively reduced, and the recognition accuracy of the change area is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and artificial intelligence technology, and in particular to a multimodal dual-branch twin network remote sensing change detection method and system for image-text fusion. Background Art

[0002] Change detection (CD) in remote sensing images aims to identify areas of land cover change (such as building expansion, deforestation, and disaster damage) by analyzing remote sensing images at different times. This approach has important applications in environmental monitoring, urban planning, and disaster assessment. However, due to the complexity of remote sensing data (such as illumination variations, seasonal differences, and sensor noise), traditional methods still face challenges in accuracy and robustness. The main issues with existing technologies are as follows:

[0003] ① Limitations of single-modality methods: Most existing methods rely solely on optical imagery (e.g., Landsat and Sentinel-2) and lack the use of auxiliary information (e.g., text metadata and geotags). This results in insufficient ability to distinguish semantically ambiguous changes (e.g., temporary shadows vs. real ground object changes). For example, pixel-level difference analysis (e.g., CVA) or deep learning models (e.g., U-Net) are susceptible to interference from lighting and noise.

[0004] ② Shortcomings of multimodal fusion: A few studies have attempted to combine multi-source data (e.g., optical and SAR imagery), but cross-modal feature alignment is difficult and the semantic guidance of textual information is not fully utilized. For example, early fusion (directly concatenating image and text features) can lead to inconsistent feature spaces, hindering network convergence.

[0005] ③ Optimization bottleneck of Siamese networks: Traditional Siamese networks (such as the dual-branch structure based on ResNet) usually only process image pairs, are not extended to multimodal inputs, and lack interaction between branches, resulting in redundant feature expression. Summary of the Invention

[0006] The technical problem solved by the present invention is to provide a multimodal fusion dual-branch Siamese network that, through the collaborative enhancement of images and text, solves the problems of low change detection accuracy and poor robustness in remote sensing scenes caused by illumination differences and the complexity of ground objects. The present invention proposes a multimodal dual-branch Siamese network remote sensing change detection method and system with image and text fusion. By fusing remote sensing images with associated text, this method effectively reduces the problem of false detection and missed detection in single-modal data and improves the recognition accuracy of changed areas.

[0007] The first object of the present invention is to provide a multimodal dual-branch twin network remote sensing change detection method based on image and text fusion, which is characterized by comprising:

[0008] Obtain text information of remote sensing categories and obtain remote sensing images before and after the event;

[0009] According to the text information and remote sensing images of remote sensing categories, the text information corresponding to the remote sensing images before and after the time phase is obtained based on the CLIP model;

[0010] The corresponding text features are obtained through Transformer encoding based on the text information corresponding to the remote sensing images of the previous and next phases;

[0011] Based on the remote sensing images of the previous and next phases, the image encoder SwinTransformer, which has no correlation with text encoding, is used to extract the corresponding image features;

[0012] Interactively fuse text features with image features and use the Transformer decoder to obtain text features with visual clues;

[0013] The text features with visual clues are updated through residual connection to obtain updated text features;

[0014] Calculate the pixel-text score map based on the updated text features and image features;

[0015] The pixel-text score map is jointly corrected with the high-level feature map to obtain the multi-level image features corresponding to the temporal remote sensing image before and after the correction, as well as the more discriminative high-level semantic feature map;

[0016] Calculate the optimal distance between the multi-level image features corresponding to the remote sensing images before and after the change and the more discriminative high-level semantic feature map, and fuse the difference information of the image features before and after the change of absolute difference and splicing feature construction;

[0017] According to the difference information of image features before and after the change, the cross-shaped Transformer multimodal decoder guided by U-shaped visual language is used to obtain the changed image segmentation mask.

[0018] Preferably, the CLIP model pre-training method uses the CLIP model to perform initial inference on original image slices extracted from a remote sensing dataset to generate text prompts and language prior knowledge corresponding to the remote sensing images.

[0019] Preferably, the visual features are obtained by fusing text features with image features through a multimodal Transformer encoder and decoder.

[0020] Preferably, when extracting image features, the method includes: visually embedding the remote sensing image into a high-level feature map using a global average pooling module in a Swin Transformer encoder to obtain image features;

[0021] The image features include rich semantic category information of the original input image and a language-compatible feature map.

[0022] Preferably, the acquisition of the discriminative high-level semantic feature map includes:

[0023] The pixel-text score map is connected to the last feature map of the remote sensing image in the Swin Transformer encoder, so that the semantic information related to the text is supplemented in the high-level semantic features, and a more discriminative high-level semantic feature map is obtained.

[0024] Preferably, the difference information of the image features before and after the change is obtained according to the following steps:

[0025] According to the remote sensing images before and after the change, the features of the images before and after the change are obtained in each layer during the layered encoding stage of the Transformer encoder;

[0026] Obtain the optimal distance between the features of each layer before and after the change according to the features of the images before and after the change of each layer;

[0027] Obtain the initial difference between the features of each layer before and after the change according to the features of the images before and after the change of each layer;

[0028] Connect the features of the images before and after each layer to obtain complete bi-temporal feature information;

[0029] The optimal distance between the images before and after the change, the initial difference between the features of each layer before and after the change, and the complete bi-temporal feature information are integrated to obtain the difference information of the image features before and after the change based on the DFE module.

[0030] Preferably, based on the multi-level feature maps and more discriminative high-level semantic feature maps obtained before and after the change, the DFE module is used to extract the difference information between the multi-level features, and the cross-shaped Transformer multimodal decoder guided by the U-shaped visual language is used to extract the changed image mask.

[0031] The second object of the present invention is to provide a computer program product, comprising a computer program, which, when executed by a processor, implements a multimodal two-branch twin network remote sensing change detection method for image-text fusion.

[0032] A third object of the present invention is to provide an electronic device comprising:

[0033] processor; and

[0034] a memory for storing executable instructions of the processor;

[0035] The processor is configured to perform a multimodal two-branch twin network remote sensing change detection method for image-text fusion by executing the executable instructions.

[0036] A fourth object of the present invention is to provide a multimodal dual-branch twin network remote sensing change detection system for image and text fusion, comprising:

[0037] The data acquisition module is used to obtain text information of remote sensing categories and obtain remote sensing images of the preceding and following time phases; according to the text information of remote sensing categories and the remote sensing images, the text information corresponding to the preceding and following time phases of remote sensing images is obtained based on the CLIP model;

[0038] The feature extraction module is used to obtain the corresponding text features through Transformer encoding based on the text information corresponding to the previous and next phase remote sensing images; based on the previous and next phase remote sensing images, the image encoder Swin Transformer, which is not correlated with the text encoding, is used to extract the corresponding image features; the text features are interactively fused with the image features, and the text features with visual clues are obtained using the Transformer decoder; the text features with visual clues are updated through residual connection to obtain the updated text features; the pixel-text score map is calculated based on the updated text features and image features; the pixel-text score map is used to jointly correct the high-level feature map to obtain the multi-level image features corresponding to the previous and next phase remote sensing images and the more discriminative high-level semantic feature map;

[0039] The difference module is used to calculate the optimal distance between the multi-level image features corresponding to the remote sensing images before and after the time phase and the more discriminative high-level semantic feature map, and fuse the difference information of the image features before and after the absolute difference and splicing feature construction;

[0040] The segmentation mask module is used to obtain the segmentation mask of the changed image based on the difference information of the image features before and after the change, using a cross-shaped Transformer multimodal decoder guided by a U-shaped visual language.

[0041] The present invention has at least the following beneficial effects:

[0042] The present invention provides a multimodal dual-branch twin network remote sensing change detection method and system for image-text fusion. The present invention effectively reduces the false detection problem (such as shadow, illumination change, seasonal change interference) of single-modal data (image only) by fusing remote sensing images with associated text (such as time, coordinates, and description of land features), and improves the recognition accuracy of changed areas. The present invention adopts a cross-modal attention mechanism to enable the network to adaptively adjust the weights of image and text features, adapt to different scenes (such as urban buildings, farmland, and forests), and reduce the impact of environmental noise. The present invention can still maintain stable detection performance and reduce false alarm rate under complex conditions such as illumination changes and partial occlusion.

[0043] This invention is not only applicable to optical remote sensing imagery but can also be extended to combinations such as SAR (Synthetic Aperture Radar) imagery + text, multi-temporal remote sensing data + meteorological text, and other combinations, offering even wider applicability. The network architecture provided by this invention can be adapted to different tasks, such as object classification and target detection, by reusing the core multimodal fusion module simply by adjusting the output layer. This invention can be applied to disaster monitoring (e.g., floods and earthquakes) or urban expansion tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is the overall structure diagram of MDS-Net;

[0045] Figure 2 A multimodal fusion Transformer decoder for image and text integration;

[0046] Figure 3 Detailed architecture of the Transformer multimodal decoder for U-shaped vision-language guidance. DETAILED DESCRIPTION

[0047] In order to illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following is a detailed description with reference to the embodiments.

[0048] The present invention mainly addresses the difficulties of traditional remote sensing change detection methods (such as pure image comparison and single-modal CNN) in effectively fusing multimodal data (such as optical images and text descriptions), resulting in a high false detection rate for complex scenes (such as shadows and seasonal changes); the existing Siamese network lacks the ability to model text information and cannot use auxiliary text (such as satellite shooting parameters and geographic tags) to improve change detection accuracy; the dual-branch network structure has insufficient feature interaction, resulting in insufficient semantic alignment between image and text modalities, affecting the quality of binary change mask generation. The core problem solved by the present invention is to provide a multimodal fusion dual-branch Siamese network that, through the collaborative enhancement of images and text, solves the problems of low change detection accuracy and poor robustness caused by illumination differences and the complexity of ground objects in remote sensing scenes.

[0049] The purpose of this invention is to provide a multimodal dual-branch twin network remote sensing change detection method and system for image and text fusion. The present invention proposes the MDS-Net model, see Figure 1As shown, the input of MDS-Net includes two sets of parallel image encoders and text encoders. The feature maps generated by the new backbone do not have strong correlation with the text features generated by the CLIP text encoder, compared with the original image encoders (ResNet and ViT). The present invention utilizes the unsupervised prediction ability of CLIP to formulate the original text hints of bi-temporal images, while adopting a hierarchical Swin Transformer to generate multi-scale features. The architecture consists of three main modules: a multimodal encoder, a difference feature enhancement module, and a U-shaped visual language guided Transformer decoder.

[0050] The overall structure of the MDS-Net proposed in this paper consists of several steps: first, the CLIP model is used to generate text representations of bi-temporal images, then a text encoder and a Transformer decoder are used to refine text features and identify the most relevant visual cues, and finally a pixel-text score map is calculated to guide the training of the learning process. The method uses the DFE module to calculate the differential information and then uses the UVLGT multimodal decoder to enhance the interaction between visual and text information.

[0051] To achieve the above objectives, the present invention provides a multimodal dual-branch Siamese network remote sensing change detection method for image-text fusion, comprising:

[0052] S1. Obtain text information of remote sensing categories; obtain remote sensing images before and after the time phase;

[0053] According to the text information and remote sensing images of remote sensing categories, the text information corresponding to the remote sensing images before and after the time phase is obtained based on the CLIP model;

[0054] The CLIP model pre-training method uses the CLIP model to perform initial reasoning on original image slices extracted from a remote sensing dataset to generate text prompts and language prior knowledge corresponding to the remote sensing images.

[0055] Among them, the before-after phase remote sensing images refer to the remote sensing images before and after the change;

[0056] In this embodiment, due to the lack of pixel-level change detection datasets that provide textual hint labels, we generate textual hints for bi-temporal images based on 56 common land cover categories, leveraging the unsupervised prediction capabilities of CLIP. CLIP exhibits different confidence levels for the predicted categories for each image; in this invention, the textual information corresponding to these different confidence levels is utilized. A higher confidence level indicates a higher probability that the image contains the corresponding category. Changing targets can be directly identified using bi-temporal textual hints, which also provides the necessary prior knowledge for remote sensing. This textual prior knowledge is combined with pre-trained image features and input into our model for training and fine-tuning.

[0057] S2. Obtain the corresponding text features through Transformer encoding based on the text information corresponding to the previous and next phase remote sensing images; and extract the corresponding image features using SwinTransformer, an image encoder that has no correlation with text encoding, based on the previous and next phase remote sensing images.

[0058] In this embodiment, in order to obtain text features, a template "a photo of a[CLS]" with K class names is used to construct a text prompt, and CLIP is used to extract text sequence features T∈R N×2×C' See also Figure 2 As shown in Figure 2, a transformer-based text encoder is used to refine the text feature T.

[0059] Obtaining image features through feature extraction based on remote sensing images, including: visually embedding high-level feature maps of remote sensing images in a SwinTransformer encoder using a global average pooling module to obtain image features;

[0060] The image features include rich semantic category information of the original input image and a language-compatible feature map.

[0061] In this embodiment, see Figure 1 As shown, given an input remote sensing image with a resolution of H×W×3, the SwinTransformer encoder output resolution is The feature map X i , where i = {1, 2, 3, 4} and C i +1>C i , then, the high-level semantic feature map is combined with the text sequence features. Among them, the global average pooling module is used on the Swin Transformer encoder side to perform visual embedding on the high-level feature map to extract image features. The specific calculation method is as follows:

[0062] X'=Concat(Mean(Reshape(X4)),X4) (1)

[0063] X'=Reshape(X'W)(2)

[0064]

[0065] Z=Reshape(X”[:,:,1:]) (4)

[0066] Where, Represents the high-level feature map extracted by Swin Transformer in the encoding stage, which contains high-level semantic information. N, C, H4, and W4 represent the four dimensions of X4, namely batch size, channel, height, and width. C×C' Matrix is a learnable parameter matrix, and Mean refers to global average pooling. represents the global representation, which contains rich semantic category information of the original input image. It is a language-compatible feature map that can be integrated with subsequent text sequence features.

[0067] S3. Interactively fuse text features with image features and use the Transformer decoder to obtain text features with visual clues; update the text features with visual clues through residual connections to obtain updated text features;

[0068] Calculate the pixel-text score map based on the updated text features and image features;

[0069] The pixel-text score map is jointly corrected with the high-level feature map to obtain the multi-level image features corresponding to the temporal remote sensing image before and after the correction, as well as the more discriminative high-level semantic feature map;

[0070] That is, the multi-level image features and more discriminative high-level semantic feature maps corresponding to the remote sensing images before and after the change are obtained respectively;

[0071] The visual features are obtained by fusing text features with image features through a multimodal Transformer encoder and decoder.

[0072] The acquisition of the discriminative high-level semantic feature map includes:

[0073] The pixel-text score map is connected to the last feature map of the remote sensing image in the Swin Transformer encoder, so that the semantic information related to the text is supplemented in the high-level semantic features to obtain a discriminative high-level semantic feature map.

[0074] It should be noted that when obtaining a more discriminative high-level semantic feature map, the remote sensing image before the change is used. In step 1, the text information of the remote sensing category and the remote sensing image before the change are pre-trained using the CLIP model to obtain the text information corresponding to the remote sensing image before the change; after going through step 2 to this step, a more discriminative high-level semantic feature map corresponding to the remote sensing image before the change is obtained;

[0075] Similarly, obtain a more discriminative high-level semantic feature map corresponding to the changed remote sensing image;

[0076] In this embodiment, text features and image features are fused through the query of the Transformer decoder through multimodal fusion:

[0077]

[0078] The attention is calculated using the following formula:

[0079]

[0080] Where Q, K, and V represent query, key, and value, respectively.

[0081] For self-attention, the input consists of textual features only, while for cross-attention, textual features act as queries and visual features act as keys and values.

[0082] In order to enable the text features to identify the most relevant visual clues, the text features are updated using residual connections as follows:

[0083] T=T+γV (7)

[0084] Here, γ is a learnable parameter that adjusts the scaling of the residual.

[0085] T is then combined with the language-compatible feature map Z to calculate the pixel-text score map, which is used to guide the learning process of the change detection model. The calculation formula is as follows:

[0086]

[0087] Where S represents the score map; and is the L2-normalized version of Z and T along the channel dimension;

[0088] It will be integrated into the decoder to supplement semantic features in the decoding stage and enhance the overall feature representation;

[0089] The score map S describes the result of pixel-text matching and constitutes one of the most critical components in the multimodal encoder. The score map is connected to the last feature map to explicitly include the language prior as follows:

[0090]

[0091] This supplements the text-related semantic information in the high-level semantic features, making the high-level semantic feature maps more discriminative.

[0092] S4, calculating the optimal distance between the multi-level image features corresponding to the remote sensing images before and after the change and the more discriminative high-level semantic feature map, and fusing the difference information of the image features before and after the change with the absolute difference and the splicing feature construction;

[0093] The difference information of the image features before and after the change is obtained according to the following steps:

[0094] According to the remote sensing images before and after the change, the features of the images before and after the change are obtained in each layer during the layered encoding stage of the Transformer encoder;

[0095] Obtain the optimal distance between the features of each layer before and after the change according to the features of the images before and after the change of each layer;

[0096] Obtain the initial difference between the features of each layer before and after the change according to the features of the images before and after the change of each layer;

[0097] Connect the features of the images before and after each layer to obtain complete bi-temporal feature information;

[0098] The optimal distance between the images before and after the change, the initial difference between the features of each layer before and after the change, and the complete bi-temporal feature information are integrated to obtain the difference information of the image features before and after the change based on the DFE module.

[0099] In this embodiment, four difference modules are used to calculate the differences between the multi-level features extracted from paired images by the layered Transformer encoder, aiming to learn the optimal distance (Optdist) metric at each scale. The calculation formula is as follows:

[0100]

[0101] where X i1 and X i2 represents the features of the pre- and post-change images from the i-th layer of the remote sensing images in the encoding stage, i∈{1,2,3,4}; BN represents BatchNorm, Conv 2D refers to 3×3 depthwise convolution, and Cat represents cascade.

[0102] See also Figure 3 As shown in the figure, the detailed structure of the U-shaped visual language guided Transformer multimodal decoder is given. and The spatial resolutions are 8×8, 16×16, 32×32, 64×64 and 128×128 respectively.

[0103] The initial difference is obtained by using the features in the remote sensing images before and after the change, which is defined as follows:

[0104]

[0105] The feature maps are concatenated to preserve the complete bi-temporal feature information, which is defined as follows:

[0106]

[0107] The calculation of DFE can be described as follows:

[0108]

[0109] where |·| represents the absolute value operation, Conv is the convolution block, [·] represents tensor concatenation, and CA is the channel attention module. The DFE module optimizes the differential feature learning strategy and significantly improves the accuracy and robustness of the RSCD task.

[0110] S5. Based on the difference information of the image features before and after the change, a cross-shaped Transformer multimodal decoder guided by the U-shaped visual language is used to obtain the segmentation mask of the changed image.

[0111] Based on the multi-level feature maps and more discriminative high-level semantic feature maps obtained before and after the change, the DFE module is used to extract the difference information between the multi-level features. Based on the U-shaped visual language-guided cross-shaped Transformer multimodal decoder, according to the rich image difference information and text sequence features obtained in the encoding stage, the cross-attention mechanism is used to capture the mutual dependence and matching degree between image features and text features. Based on the multi-level feature fusion structure of the cross-shaped cross-shaped Transformer module, the text semantic information is fused into the image feature map layer by layer, so that the model can better identify the changed area and extract the image mask after the change.

[0112] In this example, a new decoder, called the U-shaped visual-linguistic guided Transformer multimodal decoder, is used. It enhances the interaction between visual and textual information and establishes a global attention relationship. The visual-linguistic features obtained from the encoding stage contain rich semantic information from both the image and the text. By integrating this semantic information into the decoding stage, the model can more effectively identify changed regions, thereby enhancing the model's feature representation capabilities. The calculation steps are as follows:

[0113]

[0114] in, represents the text sequence features extracted from the bi-temporal image during the encoding stage, W'∈R C '×d is a linear transformation matrix. A cross-attention mechanism can be used to capture the interaction between vision and language as follows:

[0115] Q=W D ·D',K=W T ·T',V=W V ·T' (16)

[0116]

[0117] in represents the DFE feature map, Represents the visual language feature sequence W D , and are three linear transformation matrices. Then, X is evenly divided into non-overlapping horizontal and vertical stripes. The width of each head is sw. The K heads are evenly divided into two parallel groups. One group uses horizontal strip self-attention and the other group uses vertical strip self-attention. The outputs of these two groups are then connected. The detailed calculation is as follows:

[0118]

[0119] CSWAtt(X)=Concat(head1,...head k )W O (19)

[0120] Among them, H-Attention is the horizontal strip self-attention, V-Attention is the vertical strip self-attention, and W O ∈R d×d To project the self-attention result to the projection matrix of the target output dimension, the strip width sw is set to 1, 4, 4, 4 by default. Equipped with the above self-attention mechanism, the multi-mode Transformer decoder is formally defined as:

[0121] X l =MLP(LN(CSWAtt(LN(X l-1 ))+X l-1 )) (20)

[0122] Among them, X l Represents the output of the lth Transformer decoder.

[0123] In computer vision, it is generally believed that high-level image features effectively capture the semantic information of an image, while low-level image features are good at capturing finer details. Inspired by the U-Net series, we integrate multiple skip connections between the DFE and the Transformer decoder path to preserve the transmission of local details and enhance multi-scale representation. In addition, we utilize a multi-level feature fusion module to perform multimodal decoding processing on different features of dual-temporal remote sensing images, see Figure 3 As shown, the calculation steps are as follows:

[0124] F4=Conv(Concat(TransDecoder(D2,text),D1)) (21)

[0125] F3=Conv(Concat(TransDecoder(F4,text),D3)) (22)

[0126] F2=Conv(Concat(TransDecoder(F3,text),D4)) (23)

[0127] F1=Conv(Concat(TransDecoder(F2,text),D5)) (24)

[0128] Among them, D1, D2, D3, D4 and D5 are different feature maps implemented by the DFE module, and text is a sequence of visual language features.

[0129] The present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements a multimodal dual-branch twin network remote sensing change detection method for image-text fusion.

[0130] The present invention provides an electronic device, comprising:

[0131] processor; and

[0132] a memory for storing executable instructions of the processor;

[0133] The processor is configured to perform a multimodal two-branch twin network remote sensing change detection method for image-text fusion by executing the executable instructions.

[0134] The present invention provides a multimodal dual-branch twin network remote sensing change detection system for image and text fusion, comprising:

[0135] The data acquisition module is used to obtain text information of remote sensing categories; and obtain remote sensing images before and after the change; and obtain the text information corresponding to the remote sensing image through pre-training of the CLIP model based on the text information of the remote sensing category and the remote sensing image;

[0136] A feature extraction module is configured to obtain text features by encoding text information corresponding to the remote sensing image; obtain image features by feature extraction based on the remote sensing image; interactively fuse the text features with the image features to obtain visual features; update the text features with the visual features by residual connection to obtain updated text features; calculate a pixel-text score map based on the updated text features and image features; jointly correct the image features with the pixel-text score map to obtain a more discriminative high-level semantic feature map; and thereby obtain more discriminative high-level semantic feature maps corresponding to the remote sensing images before and after the change, respectively;

[0137] The difference module is used to hierarchically encode the remote sensing images before and after the change through the Transformer encoder to obtain the difference information between the multi-level features extracted from the images;

[0138] The segmentation mask module is used to obtain the segmentation mask of the changed image based on the more discriminative high-level semantic feature map before the change, the more discriminative high-level semantic feature map corresponding to the remote sensing image after the change, and the difference information between the multi-level features extracted from the image, based on the Transformer multimodal decoder guided by the U-shaped visual language.

[0139] In summary, the method provided by the present invention uses a dual-branch twin network of text and image encoding without correlation to obtain multi-level feature maps corresponding to the remote sensing images of the previous and next phases; and interactively fuses the text features with the high-level semantic feature maps, so that the text features find the most relevant visual clues, and updates the text features after the refined text features through residual connection; uses the updated text features to jointly generate a more discriminative high-level semantic feature map with the high-level semantic feature map; extracts the difference information between the multi-level features based on the multi-level feature maps of the previous and next phases and the more discriminative high-level semantic feature maps, and then obtains the changed image segmentation mask based on the cross-shaped Transformer multimodal decoder guided by the U-shaped visual language. The present invention effectively reduces the problem of false detection and missed detection of single-modal data by fusing remote sensing images with associated texts, and improves the recognition accuracy of changed areas.

[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multimodal dual-branch Siamese network remote sensing change detection method for image and text fusion, characterized by: include: Obtain text information of remote sensing categories and obtain remote sensing images before and after the event; According to the text information and remote sensing images of remote sensing categories, the text information corresponding to the remote sensing images before and after the time phase is obtained based on the CLIP model; The corresponding text features are obtained through Transformer encoding based on the text information corresponding to the remote sensing images of the previous and next phases; Based on the remote sensing images of the previous and next phases, the corresponding image features are extracted using the Swin Transformer, an image encoder that has no correlation with text encoding. Interactively fuse text features with image features and use the Transformer decoder to obtain text features with visual clues; The text features with visual clues are updated through residual connection to obtain updated text features; Calculate the pixel-text score map based on the updated text features and image features; The pixel-text score map is jointly corrected with the high-level feature map to obtain the multi-level image features corresponding to the temporal remote sensing image before and after the correction, as well as the more discriminative high-level semantic feature map; Calculate the optimal distance between the multi-level image features corresponding to the remote sensing images before and after the change and the more discriminative high-level semantic feature map, and fuse the difference information of the image features before and after the change of absolute difference and splicing feature construction; According to the difference information of image features before and after the change, the cross-shaped Transformer multimodal decoder guided by U-shaped visual language is used to obtain the changed image segmentation mask.

2. The multimodal dual-branch Siamese network remote sensing change detection method for image-text fusion according to claim 1 is characterized in that: The CLIP model pre-training method uses the CLIP model to perform initial reasoning on original image slices extracted from a remote sensing dataset to generate text prompts and language prior knowledge corresponding to the remote sensing images.

3. The multimodal dual-branch Siamese network remote sensing change detection method for image-text fusion according to claim 1 is characterized in that: The visual features are obtained by fusing text features with image features through a multimodal Transformer encoder and decoder.

4. The multimodal dual-branch Siamese network remote sensing change detection method for image-text fusion according to claim 1 is characterized in that: When extracting image features, it includes: using the global average pooling module to visually embed the high-level feature map of the remote sensing image in the Swin Transformer encoder to obtain image features; The image features include rich semantic category information of the original input image and a language-compatible feature map.

5. The multimodal dual-branch Siamese network remote sensing change detection method for image-text fusion according to claim 4 is characterized in that: The acquisition of the discriminative high-level semantic feature map includes: The pixel-text score map is connected to the last feature map of the remote sensing image in the Swin Transformer encoder, so that the semantic information related to the text is supplemented in the high-level semantic features, and a more discriminative high-level semantic feature map is obtained.

6. The multimodal dual-branch Siamese network remote sensing change detection method for image-text fusion according to claim 1 is characterized in that: The difference information of the image features before and after the change is obtained according to the following steps: According to the remote sensing images before and after the change, the features of the images before and after the change are obtained in each layer during the layered encoding stage of the Transformer encoder; Obtain the optimal distance between the features of each layer before and after the change according to the features of the images before and after the change of each layer; Obtain the initial difference between the features of each layer before and after the change according to the features of the images before and after the change of each layer; Connect the features of the images before and after each layer to obtain complete bi-temporal feature information; The optimal distance between the images before and after the change, the initial difference between the features of each layer before and after the change, and the complete bi-temporal feature information are integrated to obtain the difference information of the image features before and after the change based on the DFE module.

7. The multimodal dual-branch Siamese network remote sensing change detection method for image-text fusion according to claim 1 is characterized in that: Based on the multi-level feature maps and more discriminative high-level semantic feature maps obtained before and after the change, the DFE module is used to extract the difference information between the multi-level features, and the cross-shaped Transformer multimodal decoder guided by the U-shaped visual language is used to extract the changed image mask.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the multimodal dual-branch twin network remote sensing change detection method for image-text fusion according to any one of claims 1 to 7.

9. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the multimodal two-branch twin network remote sensing change detection method for image-text fusion according to any one of claims 1 to 7 by executing the executable instructions.

10. A multimodal dual-branch twin network remote sensing change detection system for image and text fusion, characterized by: include: The data acquisition module is used to obtain text information of remote sensing categories and obtain remote sensing images of previous and next phases; According to the text information of remote sensing categories and remote sensing images, the text information corresponding to the remote sensing images before and after the CLIP model is obtained; The feature extraction module is used to obtain the corresponding text features through Transformer encoding based on the text information corresponding to the previous and next phase remote sensing images; based on the previous and next phase remote sensing images, the image encoder SwinTransformer, which is not correlated with the text encoding, is used to extract the corresponding image features; the text features are interactively fused with the image features, and the text features with visual clues are obtained using the Transformer decoder; the text features with visual clues are updated through residual connection to obtain the updated text features; the pixel-text score map is calculated based on the updated text features and image features; the pixel-text score map is used to jointly correct the high-level feature map to obtain the multi-level image features corresponding to the previous and next phase remote sensing images and the more discriminative high-level semantic feature map; The difference module is used to calculate the optimal distance between the multi-level image features corresponding to the remote sensing images before and after the time phase and the more discriminative high-level semantic feature map, and fuse the difference information of the image features before and after the absolute difference and splicing feature construction; The segmentation mask module is used to obtain the segmentation mask of the changed image based on the difference information of the image features before and after the change, using a cross-shaped Transformer multimodal decoder guided by a U-shaped visual language.

Citation Information

Cited By

  • Remote sensing image change detection method

    CN121811266A

  • A change detection method of remote sensing image

    CN121811266B