A remote sensing image change detection and change description method based on multi-task learning
The remote sensing image change detection and description method based on multi-task learning utilizes the SegformerB1 network and feature fusion module to generate a unified change feature representation, which solves the problem of insufficient feature information sharing in remote sensing image change detection and description models, and improves the accuracy and generalization of the model.
Patent Information
- Application Number
- CN202510427094.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Existing remote sensing image change detection and change description models process the two sub-tasks independently, resulting in the inability to share feature information, which affects the accuracy and generalization of the model, especially in complex scenarios.
We adopt a multi-task learning approach, extracting shared features through a weight-sharing SegformerB1 network, and combining deformable attention modules and dynamic convolution modules for feature fusion to construct a unified representation of change features. We also optimize change detection and description tasks through multi-task learning.
It improves the pixel-level localization accuracy and semantic-level description quality of remote sensing image change detection, enhances the robustness and collaborative efficiency of the model, and improves the accuracy and generalization of change detection and description.
Smart Images

Figure CN120411725B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image change detection and change description method based on multi-task learning. Background Art
[0002] Change interpretation in remote sensing imagery is a key research area in remote sensing image analysis, primarily encompassing two subtasks: change detection and change description. Change detection primarily focuses on pixel-level spatial localization, accurately identifying the spatial locations of surface changes in remote sensing images. However, while change detection can provide precise location information for the changed region, it lacks a deep semantic understanding of the underlying meaning of the change. For example, it cannot explicitly identify the characteristics of ground objects within the changed region (e.g., color, shape), their spatial relationships (e.g., "next to," "around"), or the dynamics of the change (e.g., "appear," "remove"). In contrast, change description focuses on expressing the semantics of change, using natural language to describe the attributes and meaning of the change. While change description can provide rich semantic information, it is limited in accurately localizing changes at the pixel level.
[0003] In the existing technology, image change description tasks can be mainly divided into three categories, including template-based methods, retrieval-based methods and deep learning-based methods.
[0004] Traditional image captioning tasks often use template-based and retrieval-based methods. Template-based methods are limited by the quality of the templates, and the generated descriptions are often simple and lack innovation. Retrieval-based methods rely on prior databases, resulting in a lack of diversity and poor adaptability. Both methods are unable to generate novel descriptions and their descriptions of image changes are relatively limited and rigid.
[0005] Deep learning-based methods typically use convolutional neural networks (CNNs) as feature extractors and combine them with recurrent neural networks (RNNs) or long short-term memory networks (LSTMs) to generate descriptions. These methods still face challenges in capturing long-distance semantic relationships between images and descriptions, resulting in the descriptions' overall consistency and semantic accuracy remaining to be improved.
[0006] To address the above issues, recent research has mainly focused on the following three aspects: Transformer-based methods, methods based on improved attention mechanisms, and methods based on multi-task learning.
[0007] Among Transformer-based image captioning methods, Liu et al. used a Transformer encoder to process visual features extracted by CNN, while Shen et al. proposed a two-stage multi-task learning model that combined a variational autoencoder (VAE) with reinforcement learning to optimize the processing of spatial and semantic features. Chen et al. designed TypeFormer, which captures multi-scale object information through a multi-scale visual Transformer and incorporates a user controller to control the type of generated sentences. Furthermore, Zia et al. utilized a multi-scale visual feature encoder and an adaptive multi-head attention decoder to refine description generation. To further explore short-term spatial semantic relationships and long-term transformation dependencies, the paper proposed the Long Short-Term Relation Transformer (LSRT), aiming to more comprehensively understand the relationships between objects and generate captions. Furthermore, the paper proposed the Intra- and Inter-Relation Embedding Transformer (I2Transformer), which effectively understands the semantics of video content and subtitles, and their relationships, through cross-modal information enhancement.
[0008] Among image description methods based on improved attention mechanisms, Anderson et al. proposed a method that combines bottom-up and top-down attention mechanisms. They used Faster R-CNN to extract bottom-up image attention features and then filtered features through top-down attention, thereby improving image understanding and description generation. Zhang et al. proposed a visual alignment attention model, while Ma et al. used multi-head attention to obtain contextual features from different layers and selected object detection as an auxiliary task to obtain object-level features. However, over-reliance on object detection can affect description performance when detection errors are large. To address the semantic gap between low-level features and high-level semantics, Zhang et al. used a fully convolutional network (FCN) to generate image features and obtained intermediate vectors through an attention mechanism as input to the LSTM decoder. However, this process can result in information loss.
[0009] Image captioning methods based on multi-task learning have also made some progress in recent years. By managing multiple tasks within a single model, multi-task learning not only improves data utilization efficiency, but also mitigates overfitting by learning shared representations and enhances model robustness by integrating complementary information from multiple tasks. In natural language processing, Liu et al. proposed three RNN-based shared information mechanisms to model specific tasks, while Luong et al. explored three multi-task learning models for sequence-to-sequence models. Meanwhile, applications of multi-task learning have also emerged in the field of computer vision. For example, Chen et al. proposed a Transformer-based large-scale multi-task learning network that fine-tunes pre-trained models to tackle various visual tasks. Mohamed et al. proposed jointly learning object detection and semantic segmentation tasks, further demonstrating the advantages of the multi-task learning paradigm. In remote sensing, multi-task learning has also been applied to multimodal tasks such as remote sensing image description, text-to-image generation, and visual question answering. These tasks require models to simultaneously understand remote sensing images and associated textual information, which poses significant challenges. To address the problem of unified processing of change detection and description tasks, Liu et al. proposed the LEVIR-MCI dataset and developed the MCINet model, which employs a dual-branch structure for change detection and description. However, since the two branches are independent in feature extraction and optimization, they struggle to share useful information. To address this, Wang et al. proposed an optimization method based on a multi-task predictor. This method utilizes a unified change decoder and a classifier for both change detection and description tasks, achieving simultaneous learning within an end-to-end multi-task learning framework.
[0010] In summary, although existing research has made significant progress in remote sensing image change detection and change description, the existing models process the two subtasks of change detection and change description independently, resulting in the inability to share feature information between the two. Relying solely on the separate features of the two subtasks cannot fully explain the change information in remote sensing images, especially in the complex scenarios of remote sensing image interpretation, which affects the accuracy and generalization of remote sensing image change detection and change description models. Summary of the Invention
[0011] In view of the above-mentioned deficiencies in the prior art, the present invention provides a remote sensing image change detection and change description method based on multi-task learning.
[0012] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0013] A remote sensing image change detection and change description method based on multi-task learning includes the following steps:
[0014] Step S1, collecting a set of bi-temporal remote sensing images including a change mask and a change description;
[0015] Step S2: constructing a remote sensing image change detection and change description model, including a multi-level shared feature extraction network, a feature fusion layer, and a decoder layer; wherein the feature fusion layer integrates a deformable attention module and a dynamic convolution module;
[0016] Step S3: Input the collected image set into a multi-level shared feature extraction network, and obtain the shared features of the image set through the SegformerB1 network based on weight sharing;
[0017] Step S4: Input the acquired shared features into the feature fusion layer; enhance the shared features through the deformable attention module to obtain attention-weighted features; perform multi-scale convolution on the acquired attention-weighted features through the dynamic convolution module to extract features at different scales; perform feature splicing and feature fusion on the features at different scales to generate general features for image changes and change descriptions;
[0018] Step S5: Input the generated universal features into the change detection decoder and change description decoder in the decoder layer, perform multi-task learning on the change detection task and change description task based on the preset loss function, and output the change detection and change description results through the Softmax function;
[0019] Step S6: input the bi-temporal remote sensing image to be tested into the trained remote sensing image change detection and change description model to obtain the location information of the changed area of the bi-temporal remote sensing image to be tested and the semantic description of the change.
[0020] Furthermore, the decoder layer in step S2 includes a change detection decoder and a change description decoder.
[0021] Furthermore, the change detection decoder includes a deconvolution layer, a convolution layer and an output layer; the change detection decoder adjusts the input feature map to a uniform scale through the deconvolution layer, and inputs it into the 1×1 convolution layer, and combines the activation function of the output layer to obtain the predicted change binary map.
[0022] Furthermore, the change description decoder consists of multiple layers of Transformer, each of which contains a mask-based multi-head attention sublayer, a cross-attention layer and a feed-forward neural network.
[0023] Furthermore, in step S4, the shared features are enhanced by the deformable attention module to obtain the attention-weighted features, which includes the following steps:
[0024] (1) Based on the shared features of the input, the horizontal and vertical offsets are generated through the convolution layer, and the Tanh function is used to limit the offsets to a reasonable range;
[0025] (2) Perform bilinear interpolation sampling on the input shared features according to the generated offset to obtain the sampled features;
[0026] (3) Through the multi-head attention mechanism, the attention-weighted features of the sampled features are obtained.
[0027] Furthermore, the step S4 generates universal features by the dynamic convolution module, including the following steps:
[0028] For the attention-weighted features, convolution kernels of different scales are used to extract features of different scales. The formula is as follows:
[0029] out1=Conv 3x3 (out_mid)
[0030] out2=Conv 1x5 (out mid )
[0031] out3=Conv 5x1 (out_mid)
[0032] Among them, out_mid is the attention weighted feature, Conv 3x3 Represents a 3x3 convolution operation, used to extract local features; Conv 1X5 Represents a 1x5 convolution operation, which is used to extract global features in the horizontal direction; Conv 5x1 Represents a 5x1 convolution operation, which is used to extract global features in the vertical direction; out1, out2, and out3 represent features of different scales obtained through three convolution operations;
[0033] Use the Concat function to concatenate features of different scales to obtain the concatenated features out cat , the formula is as follows:
[0034] out cat =Concat(out1, out2, out3)
[0035] The concatenated features out are processed through the BatchNorm layer and the activation function ReLU cat Perform normalization and nonlinear transformation, and then pass 1x1 convolution Conv 1x1 Map the features back to the original dimension to generate the general features out of the image changes and the description of the changes. The formula is as follows:
[0036] out=Conv 1x1 (ReLU(BatchNorm(out cat ))).
[0037] Furthermore, the loss function preset in step S5 is as follows:
[0038] Loss function for the change detection task:
[0039]
[0040] in, is the loss of the change detection task, C is the number of categories, H and W are the height and width of the ground truth respectively, represents the true label at position (h, w), Represents the predicted change probability at position (h, w);
[0041] Change the loss function of the description task:
[0042]
[0043] in, is the loss of the change description task, N represents the length of the predicted change description, V represents the vocabulary size, represents the nth word in the corresponding descriptive sentence, represents the predicted probability that the nth word is classified into vocabulary index v;
[0044] Total loss of change detection and change description models:
[0045]
[0046] in, is the total loss of the change detection and change description models, and detach(·) represents the operation of separating the loss gradient flow.
[0047] The present invention has the following beneficial effects:
[0048] This method proposes a remote sensing image change detection and change description method based on multi-task learning. The shared features of bi-temporal remote sensing images are obtained through the SegformerB1 network based on weight sharing. A deformable attention mechanism is introduced to focus on key areas. A dynamic convolution mechanism is introduced to perform multi-scale convolution operations and feature splicing on the shared features, thereby obtaining universal features that can fully explain the change information of remote sensing images. A unified change feature representation is constructed, which enhances the expressiveness and robustness of the features, and improves the accuracy and generalization of the remote sensing image change detection and change description models. Using universal features, multi-task learning is performed on the change detection task and the change description task based on the loss function preset by the model, and the change detection and change description results are output through the Softmax function, which improves the collaborative efficiency between tasks and improves the accuracy of pixel-level change positioning of remote sensing images and the quality of semantic-level change description. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is the overall framework diagram of this method.
[0050] Figure 2 This is an example of a bi-temporal remote sensing image.
[0051] Figure 3 The process of building a change detection and change description model for remote sensing images.
[0052] Figure 4 This is the structural framework diagram of the feature fusion layer.
[0053] Figure 5 This is the structural framework diagram of the change detection decoder.
[0054] Figure 6 The structural framework diagram of the decoder is described for the changes. DETAILED DESCRIPTION
[0055] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0056] like Figure 1 As shown, a remote sensing image change detection and change description method based on multi-task learning includes the following steps:
[0057] Step S1, collecting a bi-temporal remote sensing dataset including a change mask and a change description;
[0058] The training used the open source dataset LEVIR-MCI, which contains a large number of change masks and change descriptions. The LEVIR-MCI dataset contains 10,077 dual-temporal images, each with a spatial size of 256×256 pixels and a high resolution of 0.5m / pixel, and has corresponding annotated masks and five annotated descriptions; the LEVIR-MCI dataset is an extension of the dataset LEVIR-CC, and further provides additional change detection masks for each pair of dual-temporal images, highlighting the changed roads and changed buildings. Unlike its predecessor, the LEVIR-MCI dataset provides each pair of images with diverse annotations from different interpretation perspectives, further enhancing its practicality for comprehensive remote sensing change interpretation tasks. Figure 2As shown in Figure 2, some examples of LEVIRMCI are shown. Each pair of bi-temporal images has a change detection mask and five sentences describing the changes. In the change detection mask, the changed buildings are marked in red, and the changed roads are marked in yellow.
[0059] Step S2: constructing a remote sensing image change detection and change description model, including a multi-level shared feature extraction network, a feature fusion layer, and a decoder layer; wherein the feature fusion layer integrates a deformable attention module and a dynamic convolution module;
[0060] like Figure 3 As shown in the figure, it is the construction process of remote sensing image change detection and change description model;
[0061] The decoder layer includes a change detection decoder and a change description decoder.
[0062] Step S3: Input the collected image set into a multi-level shared feature extraction network, and obtain the shared features of the image set through the SegformerB1 network based on weight sharing;
[0063] The weight-sharing SegformerB1 serves as the backbone network for shared feature extraction to extract features from bi-temporal remote sensing imagery. SegformerB1 is a Transformer-based, efficient feature extraction network that is lightweight and high-performance, making it particularly suitable for remote sensing image processing tasks. This network has been pre-trained on the ImageNet dataset. By using weight sharing, the feature extraction process for bi-temporal imagery can fully utilize the common information between the two temporal images, thereby achieving more efficient feature representation.
[0064] The dual-temporal remote sensing images I1 and I2 are input into the SegformerB1 network respectively to obtain the shared feature representations F1 and F2 of the two temporal images. The formula is as follows:
[0065] F1=SegformerB1(I1)
[0066] F2=SegformerB1(I2)
[0067] Through this shared feature extraction layer, the model can ensure that the deep shared features extracted from the input image can not only support pixel-level change localization, but also provide strong support for semantic-level change description.
[0068] Step S4: Input the acquired shared features into the feature fusion layer; enhance the shared features through the deformable attention module to obtain attention-weighted features; perform multi-scale convolution on the acquired attention-weighted features through the dynamic convolution module to extract features at different scales; perform feature splicing and feature fusion on the features at different scales to generate general features for image changes and change descriptions;
[0069] like Figure 4 As shown in Figure 1, it is a structural diagram of the feature fusion layer;
[0070] The deformable attention module generates offsets through deformable convolution and dynamically adjusts the attention area, enabling the model to focus on key areas in the image. Combined with relative position offsets, it enhances the spatial context information of features, effectively captures local and global features in the image, and improves the model's ability to model complex spatial structures in the image.
[0071] First, the horizontal and vertical offsets are generated through the convolution layer, and the formula is as follows:
[0072] offsets = Conv(Query)
[0073] Where Query represents the input query feature, and its shape is (b, c, h, w), where b is the batch size, c is the number of channels, h and w are the spatial dimensions, Conv represents the convolution operation, and offsets represents the generated offset, and its shape is (b, 2, h′, w′), where 2 represents the horizontal and vertical offsets, and h′ and w′ are the spatial dimensions after downsampling.
[0074] Use the Tanh function to limit the offset to a reasonable range and adjust the initial grid according to the offset. The formula is as follows:
[0075] vgrid=grid+offsets
[0076] grid represents the initial regular grid with a shape of (h′, w′, 2), and vgrid represents the adjusted grid with a shape of (b, h′, w′, 2);
[0077] Next, based on the generated offset, the input shared features are sampled using bilinear interpolation to obtain the adjusted sampled features. The formula is as follows:
[0078] kv_feats=GridSample(group(x),vgrid sca1ed )
[0079] Among them, group(x) is the group representation of the input feature, the shape is (b×g,c / g,h,w), where g is the number of groups, vgrid scaledRepresents the normalized adjustment grid, with a shape of (b, h′, w′, 2); GridSample represents the bilinear interpolation sampling operation, and kv_feats represents the sampled features, with a shape of (b, c, h′, w′);
[0080] Afterwards, the multi-head attention mechanism is used to calculate the similarity between the query, key, and value, and combined with the relative position bias to generate the attention weight. The formula is as follows:
[0081]
[0082] Among them, Query represents the query feature and Key key feature, both of which have the shape of (b, h, n, d); h is the number of attention heads, n is the number of spatial positions, d is the dimension of each head, and d k Represents the dimension of each head, rel_pos_bias is the relative position bias, sim is the similarity matrix, attn is the attention weight, and the shape is (b,h,n,n);
[0083] Finally, the value is weighted and summed according to the attention weight, and the attention weighted feature is output. The formula is as follows:
[0084] out_mid=attn·Value#(7)
[0085] Among them, Value is the value feature, the shape is (b, h, n, d), out_mid is the output attention weighted feature, the shape is (b, h, n, d);
[0086] The dynamic convolution module is used to enhance the feature expression capability. The module uses three different scale convolution kernels (3x3, 1x5, 5x1) to convolve the input features and extract multi-scale features. The formula is as follows:
[0087] out1=Conv 3x3 (out_mid)
[0088] out2=Conv 1x5 (out_mid)
[0089] out3=Conv 5x1 (out_mid)
[0090] Among them, Conv 3x3 Represents a 3x3 convolution operation, used to extract local features; Conv 1x5 Represents a 1x5 convolution operation, which is used to extract global features in the horizontal direction; Conv 5x1Represents a 5x1 convolution operation, which is used to extract global features in the vertical direction; out1, out2, and out3 represent features of different scales output by the three convolution operations, all with the shape of (b, c, h, w);
[0091] Use the Concat function to concatenate features of three different scales to form a richer feature representation. The formula is as follows:
[0092] out cat =Concat(out1, out2, out3)
[0093] Among them, out cat Represents the concatenated features, the shape is (b, 3c, h, w);
[0094] Next, the concatenated features are normalized and nonlinearly transformed through the BatchNorm layer and activation function, and then the features are mapped back to the original dimension through 1x1 convolution. The formula is as follows:
[0095] out=Conv 1x1 (ReLU(BatchNorm(out cat )))
[0096] Among them, BatchNorm represents the batch normalization operation, which normalizes the concatenated features; ReLU is the activation function, which enhances the nonlinear ability; out represents the common features after fusion, and the shape is (b, c, h, w);
[0097] Through this fusion mechanism, the network can generate a universal change representation under a unified multi-task learning framework, thereby simultaneously providing pixel-level change localization and semantic-level change description, ensuring the synergy between change detection and change description tasks, enabling the network to share knowledge between different tasks, and improving the accuracy and comprehensiveness of change interpretation.
[0098] Step S5: Input the generated universal features into the change detection decoder and change description decoder in the decoder layer, perform multi-task learning on the change detection task and change description task based on the preset loss function, and output the change detection and change description results through the Softmax function;
[0099] like Figure 5 As shown, a simplified multi-scale change detection decoder is proposed;
[0100] The change detection decoder consists of a deconvolution layer, a convolution layer, and an output layer. To facilitate decoding of the change mask, the common feature F extracted by the multi-level shared feature representation module and fused after feature fusion is used. out, these feature maps are adjusted to a uniform scale through the deconvolution operation and finally input into a 1 x 1 convolution layer Conv 1x1 , combined with the Softmax function of the output layer, we get the predicted change binary map P CD , the formula is as follows:
[0101] P CD =Softmax(Conv 1x1 (F out ))
[0102] like Figure 6 As shown in the figure, the change description decoder consists of multiple layers of Transformers. Each layer has a mask-based multi-head attention sublayer, a cross-attention layer, and a feedforward neural network. To ensure information continuity and enhance the robustness of the model, these sublayers are enhanced through residual connections and layer normalization techniques. Finally, the change description output is generated by a linear layer, and then the Softmax activation function is applied to output the change description result.
[0103] During the training phase, the model performs multi-task learning on change detection and change description. The training parameters and model network parameters are shown in Tables 1 and 2.
[0104] Table 1 Training parameters
[0105]
[0106] Table 2 Model network parameters
[0107]
[0108]
[0109] For the change detection task, the widely adopted cross-entropy loss function is used to measure the difference between the predicted change mask and the true annotation. The loss function is defined as:
[0110]
[0111] in, is the loss of the change detection task, C is the number of categories, H and W are the height and width of the ground truth respectively, represents the true label at position (h, w), Represents the predicted change probability at position (h, w);
[0112] Similarly, for the change description task, the cross entropy function is used to calculate the loss between the predicted sentence and its corresponding true sentence. The loss function of the change description task is defined as:
[0113]
[0114] in, is the loss of the change description task, N represents the length of the predicted change description, V represents the vocabulary size, represents the nth word in the corresponding descriptive sentence, represents the predicted probability that the nth word is classified into vocabulary index v;
[0115] In order to balance the losses of the two tasks during model training and ensure that the two tasks contribute equally to model optimization, the total model loss is defined as:
[0116]
[0117] in, is the total loss of the change detection and change description models, and detach(·) represents the operation of separating the loss gradient flow.
[0118] Due to the lack of methods that simultaneously address change detection and change description, in order to evaluate the effectiveness of ChangeFusionNet in handling these two tasks, the model in this method is compared with established methods in their respective fields;
[0119] Tables 3 and 4 list the performance comparisons of the models in this method and the comparison methods on the LEVIR-MCI dataset. To facilitate the presentation of the results, B-1, B-2, B-3, B-4, M, R, and C are used in the tables to represent BLEU-1, BLEU-2, BLEU-3, BLEU-4, METEOR, ROUGE-L, and CIDEr-D, respectively.
[0120] Table 3 Experimental comparison of change detection methods on the LEVIR-MCI dataset
[0121]
[0122] Table 4 Experimental comparison of change description methods on the LEVIR-MCI dataset
[0123]
[0124] Comparative analysis reveals that our method outperforms other methods across most metrics. Specifically, compared to the state-of-the-art MCINet, our method improves mIoU by 0.14%, BLEU-1 by 0.13%, BLEU-2 by 0.32%, and BLEU-3 by 0.08%. Compared to nine other change detection methods, our method achieves significant improvements in mIoU ranging from 0.14% to 3.87%. Furthermore, compared to Chg2Cap, our method improves BLEU-4 by 1.36%, METEOR by 0.4%, and CIDEr-D by 2.06%. These quantitative results demonstrate that the use of a multi-level shared feature representation based on weight sharing and a feature fusion module with a deformable attention mechanism can effectively improve the accuracy of change detection and the quality of change description. Furthermore, our method significantly outperforms other single-task methods, demonstrating the superiority of multi-task learning strategies and effectively improving the accuracy and generalization of change detection and change description models for remote sensing imagery.
[0125] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0126] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0128] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
[0129] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A remote sensing image change detection and change description method based on multi-task learning, characterized in that: The following steps are involved: Step S1, collecting a set of bi-temporal remote sensing images including a change mask and a change description; Step S2: constructing a remote sensing image change detection and change description model, including a multi-level shared feature extraction network, a feature fusion layer, and a decoder layer; wherein the feature fusion layer integrates a deformable attention module and a dynamic convolution module; Step S3: Input the collected image set into a multi-level shared feature extraction network, and obtain the shared features of the image set through the SegformerB1 network based on weight sharing; Step S4: Input the acquired shared features into the feature fusion layer; enhance the shared features through the deformable attention module to obtain attention-weighted features; perform multi-scale convolution on the acquired attention-weighted features through the dynamic convolution module to extract features at different scales; perform feature splicing and feature fusion on the features at different scales to generate general features for image changes and change descriptions; Step S5: Input the generated universal features into the change detection decoder and change description decoder in the decoder layer, perform multi-task learning on the change detection task and change description task based on the preset loss function, and output the change detection and change description results through the Softmax function; Step S6: input the bi-temporal remote sensing image to be tested into the trained remote sensing image change detection and change description model to obtain the location information of the changed area of the bi-temporal remote sensing image to be tested and the semantic description of the change.
2. The method for remote sensing image change detection and change description based on multi-task learning according to claim 1 is characterized in that: The decoder layer in step S2 includes a change detection decoder and a change description decoder.
3. The method for remote sensing image change detection and change description based on multi-task learning according to claim 2 is characterized in that: The change detection decoder includes a deconvolution layer, a convolution layer, and an output layer. The change detection decoder adjusts the input feature map to a uniform scale through the deconvolution layer, inputs it to the 1×1 convolution layer, and combines it with the activation function of the output layer to obtain a predicted change binary map.
4. The method for remote sensing image change detection and change description based on multi-task learning according to claim 2, characterized in that: The change description decoder consists of multiple layers of Transformer, each of which contains a mask-based multi-head attention sublayer, a cross-attention layer and a feed-forward neural network.
5. The method for remote sensing image change detection and change description based on multi-task learning according to claim 1, characterized in that: In step S4, the shared features are enhanced by the deformable attention module to obtain the attention-weighted features, which includes the following steps: (1) Based on the shared features of the input, the horizontal and vertical offsets are generated through the convolution layer, and the Tanh function is used to limit the offsets to a reasonable range; (2) Perform bilinear interpolation sampling on the input shared features according to the generated offset to obtain the sampled features; (3) Through the multi-head attention mechanism, the attention-weighted features of the sampled features are obtained.
6. The method for remote sensing image change detection and change description based on multi-task learning according to claim 5, characterized in that: The step S4 generates universal features by using a dynamic convolution module, including the following steps: For the attention-weighted features, convolution kernels of different scales are used to extract features of different scales. The formula is as follows: out1=Conv 3x3 (out_mid) out2=Conv 1x5 (out mid ) out3=Conv 5x1 (out_mid) Among them, out_mid is the attention weighted feature, Conv 3x3 Represents a 3x3 convolution operation, used to extract local features; Conv 1x5 Represents a 1x5 convolution operation, which is used to extract global features in the horizontal direction; Conv 5x1 Represents a 5x1 convolution operation, which is used to extract global features in the vertical direction; out1, out2, and out3 represent features of different scales obtained through three convolution operations; Use the Concat function to concatenate features of different scales to obtain the concatenated features out cat , the formula is as follows: out cat =Concat(out1,out2,out3) The concatenated features out are processed through the BatchNorm layer and the activation function ReLU cat Perform normalization and nonlinear transformation, and then pass 1x1 convolution Conv 1x1 Map the features back to the original dimension to generate the general features out of the image changes and the description of the changes. The formula is as follows: out=Conv 1x1 (ReLU(BatchNorm(out cat )))。 7. The method for remote sensing image change detection and change description based on multi-task learning according to claim 1, characterized in that: The loss function preset in step S5 is as follows: Loss function for the change detection task: in, is the loss of the change detection task, C is the number of categories, H and W are the height and width of the ground truth respectively, represents the true label at position (h, w), Represents the predicted change probability at position (h, w); Change the loss function of the description task: in, is the loss of the change description task, N represents the length of the predicted change description, V represents the vocabulary size, represents the nth word in the corresponding descriptive sentence, represents the predicted probability that the nth word is classified into vocabulary index v; Total loss of change detection and change description models: in, is the total loss of the change detection and change description models, and detach(·) represents the operation of separating the loss gradient flow.
Citation Information
Patent Citations
Strong supervision change detection method based on convolutional neural network and visual attention model
CN119206487A
Remote sensing image semantic segmentation method based on double-branch multi-scale fusion network
CN119579891A