A Multimodal Remote Sensing Image Change Detection Method Based on Self-Supervised Learning
Through the multimodal remote sensing image change detection method based on self-supervised learning, multimodal image features are extracted and unified, and the problem of large dependence on labeled data in the prior art is solved, and efficient change detection is achieved.
Patent Information
- Application Number
- CN202311060952.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-08-22
AI Technical Summary
Existing remote sensing image change detection methods are difficult to process multimodal dual-time images from different sensors, and rely heavily on labeled data, and the image domain gap leads to direct comparison and analysis of images before and after changes.
A multimodal remote sensing image change detection method based on self-supervised learning is adopted to extract feature maps through ternary feature extraction networks, and a unified mapping unit is used to map the feature maps to comparable feature spaces. Self-supervised training is performed by combining cross entropy loss and comparison loss, and finally a change map is generated through the threshold segmentation algorithm.
Without labeling data, unifying multimodal remote sensing image features from the depth feature space reduces the consumption of human and material resources in the change detection task, solving the problem of image domain gap in multimodal image change detection, and achieving efficient change detection.
Smart Images

Figure CN117237801B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal remote sensing image change detection method based on self-supervised learning, belonging to the field of computer vision. Background Art
[0002] Change detection is the process of identifying differences in the state of an object or phenomenon by observing it at different times. Change detection based on remote sensing images is an important method for detecting changes on the Earth's surface and has a wide range of applications in urban planning, environmental monitoring, agricultural surveys, disaster assessment, and map revision.
[0003] Existing remote sensing image change detection methods mainly target pre-change and post-change images from the same sensor (i.e., the pre-change and post-change images are of the same modality). However, in the real world, for some specific applications such as disaster management, there is a strong sense of timeliness and urgency, and the immediately available post-change images may be of a different modality from the pre-change images, which poses a major challenge to the remote sensing image change detection task. There will be different image domain gaps in multi-modal dual-temporal image pairs from different sensors, which makes it impossible to directly compare and analyze the pre-change and post-change images to obtain the change map. In addition, since multi-modal dual-temporal image pairs require the collaboration of experts from different image fields to perform pixel-level annotation on the image pairs, this requirement makes the cost of obtaining labeled samples extremely high, resulting in a very small number of labeled samples.
[0004] Utilizing the self-supervised learning paradigm to reduce the dependence of change detection methods on labeled data, and inspired by the excellent performance of deep learning in various industries, the present invention designs a change detection framework based on self-supervised learning for multi-modal remote sensing images. Summary of the Invention
[0005] The technical problem to be solved by the present invention is:
[0006] In order to avoid the deficiencies of the prior art, the present invention provides a multi-modal remote sensing image change detection method based on self-supervised learning.
[0007] In order to solve the above technical problem, the technical solution adopted by the present invention is:
[0008] A multi-modal remote sensing image change detection method based on self-supervised learning, characterized by the following steps:
[0009] Step 1: Feature map extraction
[0010] Taking the pre-change image of modality one, the post-change image of modality two, and the spliced image as three independent inputs, and inputting them into a ternary feature extraction network to respectively obtain feature maps F m1 , F m2 and Fd ; The spliced image is obtained by stacking the pre - change image of modality one and the post - change image of modality two along the channel dimension;
[0011] Step 2: Unify the spatial feature maps
[0012] The feature maps F m1 , F m2 and F d are mapped to a comparable feature space by the unified mapping unit UMU to obtain the feature maps F′ m1 , F′ m2 and F d ′;
[0013] Step 3: Network self - supervised training
[0014] During the training stage of the network, cross - entropy loss is used to supervise the effectiveness of the pre - change and post - change image feature maps, and contrastive loss is used to ensure the label - free self - supervised training of the entire network system;
[0015] Step 4: Network inference and generation of the change map
[0016] The obtained features are analyzed using a threshold segmentation algorithm to generate the final change map.
[0017] A further technical solution of the present invention: The three - element feature extraction network system is composed of a pseudo - twin network and a differential information network;
[0018] Each branch of the pseudo - twin network consists of 5 convolutional layers with a convolutional kernel size of 3×3. After each convolutional layer, a batch normalization layer and a rectified linear unit activation function are used to maintain gradient stability, prevent network overfitting, and enhance the network's ability to learn non - linear features; The pre - change image of modality one and the post - change image of modality two are input into the pseudo - twin network to extract features, obtaining the feature maps F m1 and F m2 ;
[0019] The differential information network includes four stages. The first stage contains 4 residual blocks and a 3×3 kernel convolutional layer; In the second stage, each branch processes the feature maps at different scales; These branches operate independently and are composed of multiple consecutive residual blocks; The third and fourth stages mimic the structure of the second stage; When implementing the fusion strategy for feature maps of different resolutions, the upsampling part uses bilinear upsampling operation followed by a 1×1 convolution, and the downsampling uses a convolutional layer with a kernel size of 3×3 and a stride of 2; The spliced image is input into the differential information network structure to extract features, obtaining the feature map F d .
[0020] Further technical solution of the present invention: The unified mapping unit consists of a token encoder and a token decoder.
[0021] The input of the token decoder is three independent features F obtained by the triple feature extraction network. m1 ,F m2 and F d ; The input feature map is represented as F ∈ R b×c×h×w , which is converted into a three-dimensional token embedding vector of a specific size, with the size of b×l×c; b, c, h, and w represent the batch size, the number of channels, and the height and width of the input features respectively, and l represents the token length.
[0022] Encoding process of the token encoder: The three-dimensional token embedding is in the encoder to capture the context information in the global; in this process, a set of trainable parameters are added to the token for position embedding PE; the encoder follows the standard transformer structure, including a multi-head attention MHA module and a feed-forward neural network module; in addition, layer normalization LN is applied before each block; thus, the token embedding vector is obtained, denoted as T ∈ R b×l×c ;
[0023] Input of the token decoder: The token decoder receives two different inputs; one is the feature map F obtained by the convolutional network, which can also be considered as the feature map extracted by the triple feature extraction network; the other input is the token embedding vector T, which contains the global context information generated by the token encoder.
[0024] Decoding process of the token decoder: The token decoder is similar to the structure of the token encoder, using PE to endow the original convolutional feature F with position information; it consists of multiple layers, and each layer contains a combination of self-attention and feed-forward neural network; the following are two key subroutines:
[0025] Layer normalization LN: Before each decoder layer, layer normalization is applied to normalize the features, thereby enhancing the training stability.
[0026] Multi-head attention MHA: The decoder adopts the multi-head attention mechanism, aiming to understand the relationship between different tokens, thereby enriching the context understanding; there are differences between this MHA and the MHA used in the token encoder; among them, Query is derived from the convolutional feature F, while Key and Value are derived from the token embedding vector T.
[0027] Further technical solution of the present invention: The function of the cross-entropy loss is expressed as:
[0028] L1 = crossentropy(F′ m1 ,C m1 )
[0029] L2 = crossentropy(F′ m2 , C m2 )
[0030]
[0031]
[0032] where C m1 is the pseudo-label of F′ m1 , and C m2 is the pseudo-label of F′ m2 .
[0033] A further technical solution of the present invention: The function of the contrastive loss is expressed as:
[0034]
[0035] where d i,j represents the distance between the pixels corresponding to the feature maps F′ m1 and F′ m2 at the coordinates (i, j), y i,j represents the value of F′ d at the coordinates (i, j), and Margin represents a manually set threshold, and setting this threshold is to strengthen the distance of the feature map pair.
[0036] A further technical solution of the present invention: The threshold segmentation algorithm is the OSTU threshold algorithm.
[0037] A computer system, comprising: one or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0038] A computer-readable storage medium, storing computer-executable instructions, which are used to implement the above method when executed.
[0039] The beneficial effects of the present invention are as follows:
[0040] Based on self-supervised learning, the present invention unifies multi-modal remote sensing image features from the deep feature space without the need for any labels, and directly analyzes the deep feature map by integrating the traditional threshold segmentation method to obtain the required change map. The present invention overcomes the dependence on labeled data in previous remote sensing image change detection methods, reduces the consumption of human and material resources in the change detection task, and at the same time solves the problem that there is an image domain gap between two temporal image pairs in multi-modal image change detection and they cannot be directly compared. The overall learning framework is easy to implement, the algorithm is simple, and the execution efficiency is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings are only for the purpose of illustrating specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components.
[0042] Figure 1 Self-supervised learning framework.
[0043] Figure 2 Pseudo-twin network structure.
[0044] Figure 3 Differential information network structure.
[0045] Figure 4 Unified mapping unit encoder structure. Where Q, K, and V represent Query, Key, and Value, coming from feature F. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0047] The present invention reduces the dependence of the change detection method on data based on the self-supervised learning paradigm, and makes use of the characteristics of the change detection task to ingeniously design a differential contrast auxiliary task, so that the network can obtain a feature map that can represent multi-modal two-temporal images through iterative training without labels. Then, considering from the global perspective of the image, the distance of the feature map in the dimension space caused by different image domains is reduced, so that the image features are comparable in the dimension of the feature space.
[0048] A multi-modal remote sensing image change detection method based on self-supervised learning provided by the present invention, as Figure 1 shown, includes the following steps:
[0049] Step 1: Feature map extraction. The image before the change in modality one, the image after the change in modality two, and the spliced image (obtained by stacking the image before the change in modality one and the image after the change in modality two along the channel dimension) are used as three independent inputs to train a ternary feature extraction network system (there is no shared parameter between the three branch networks), and the feature maps F m1 , F m2 and F d are obtained respectively.
[0050] Step 2: Unify the feature map space. The features F m1 , F m2 and F d are mapped to a comparable feature space through the proposed Unified Mapping Unit (UMU). The feature maps F' m1 , F' m2 and F d ' are obtained, which is convenient for comparison and learning between feature maps.
[0051] Step 3: Network self-supervised training. In the training stage of the network, cross-entropy loss is used to supervise the effectiveness of the feature maps of the images before and after the change. In addition, contrastive loss is used to ensure the self-supervised training of the entire network system without labels.
[0052] Step 4: Network inference and generation of the change map. Through self-supervised training, comparable dual-temporal image feature pairs F' m1 and F' m2 in the feature space are obtained, effectively retaining the information of the dual-temporal images. Then, an appropriate threshold segmentation algorithm is used to analyze the obtained feature pairs F' m1 and F' m2 to generate the final change map.
[0053] Example:
[0054] Step 1: Feature map extraction.
[0055] In the present invention, the image before the change in modality one used for training the network is a multi-spectral image (including four spectral bands of red, blue, green, and infrared) captured by the Sentinel-2 sensor, the image after the change in modality two is a SAR image captured by the Sentinel-1 sensor at the same location before the change, the spliced image is obtained by stacking the multi-spectral image and the SAR image along the channel, the acquisition location of the dataset is in Hong Kong, and the image size is 695×540. We input the three images into the ternary feature extraction network system, and the ternary feature extraction network consists of a pseudo-twin network ( Figure 2 ) and a differential information network ( Figure 3 ).
[0056] Structure of the pseudo-twin network:
[0057] Each branch of the network consists of 5 convolutional layers with a convolutional kernel size of 3×3. After each convolutional layer, a batch normalization layer and a rectified linear unit (ReLU) activation function are used to maintain gradient stability, prevent network overfitting, and enhance the network's ability to learn non-linear features. Note that the two branches share the same structure but have independent weights. Compared with existing deep models, the proposed pseudo-twin network is simpler and more efficient. The pseudo-twin network does not contain any downsampling layers, thus eliminating the possible loss of image information during the downsampling process. The pre-change image of modality one and the post-change image of modality two are input into the pseudo-twin network to extract features, obtaining feature maps F m1 and F m2 .
[0058] Structure of the differential information network:
[0059] The network performs four stages of calculations. The first stage contains 4 residual blocks and a convolutional layer with a 3×3 kernel. In the second stage, each branch processes the feature map at different scales. These branches run independently and consist of multiple consecutive residual blocks. The third and fourth stages mimic the structure of the second stage: based on the two branches in the original second stage processing the feature map at two different scales, a branch is added in the third stage; a branch is added again in the fourth stage based on the third stage. That is, the second stage uses two branches to process at two different scales, and the third and fourth stages use three and four branches respectively for processing. Each branch runs independently and consists of multiple consecutive residual blocks. The key motivation for this design is that the features learned by the network can both maintain a high-resolution representation and learn semantic information. In addition, when implementing the fusion strategy for feature maps of different resolutions, the upsampling part uses bilinear upsampling operation followed by a 1×1 convolution, and the downsampling uses a convolutional layer with a kernel size of 3×3 and a stride of 2. The spliced image is input into the differential information network structure to extract features, obtaining feature map F d .
[0060] Step 2: Unify the feature map space.
[0061] This invention studies the multi-modal remote sensing image change detection task. Since there is usually a large image domain gap between different modal images, mapping multi-modal bi-temporal images to a comparable feature space is still a major obstacle in multi-modal remote sensing image CD. To solve this problem, this invention designs a unified mapping unit (UnifiedMapping Unit, UMU) to map the three independent features obtained from the ternary feature extraction network (which are F m1 , F m2 and F d)Projected into a comparable feature space.
[0062] The unified mapping unit consists of a token encoder ( Figure 4 ) and a token decoder, and its structure is as follows:
[0063] Token encoder:
[0064] Input: The input of the token decoder is three independent features F obtained by the triple feature extraction network m1 , F m2 and F d . To adapt to the limitations of computing and storage, the input feature map is represented as F ∈ R b×c×h×w , and before further processing, it is converted into a three-dimensional token embedding vector of a specific size, with dimensions b×l×c. Here, b, c, h, and w represent the batch size, the number of channels, and the height and width of the input features respectively, and l represents the token length (empirically set to 4 in the present invention).
[0065] Encoding process: The three-dimensional token embedding is in the encoder to capture the context information in the global. During this process, a set of trainable parameters is added to the token for position embedding (Position Embedding, PE). The encoder follows the standard transformer structure, including a multi-head attention (Multi-Head Attention, MHA) module and a feedforward neural network block. In addition, layer normalisation (Layer Normalisation, LN) is applied before each block. Thus, the token embedding vector is obtained, denoted as T ∈ R b×l×c .
[0066] Token decoder:
[0067] Input: The token decoder receives two different inputs. One is the feature map F obtained by the convolutional network, which can also be considered as the feature map extracted by the triple feature extraction network. The other input is the token embedding vector T, which contains the global context information generated by the token encoder.
[0068] Decoding process: The token decoder is similar to the structure of the token encoder. The PE is used to endow the original convolutional feature F with position information. It consists of multiple layers, and each layer contains a combination of self-attention and a feedforward neural network. The following are two key subroutines:
[0069] a) Layer Normalization (LN): Before each decoder layer, layer normalization is applied to normalize the features, thereby enhancing training stability.
[0070] b) Multi-Head Attention (MHA): The decoder adopts the multi-head attention mechanism, aiming to understand the relationships between different tokens, thereby enriching context understanding. Note that there are differences between this MHA and the MHA used in the token encoder. Among them, Query is derived from the convolutional feature F, while Key and Value are derived from the token embedding vector T. This arrangement enables the decoder to focus on the relevant token information based on the convolutional feature representation.
[0071] Step 3: Network self-supervised training.
[0072] The present invention is trained based on self-supervised learning. It is carried out under the Linux operating system, the design of the change detection network is implemented under the open-source PyTorch deep learning framework, and the network is trained under a single Nvidia GeForce GTX 1080Ti GPU. The Adam optimization method is adopted in the backpropagation process of the network. The training process of the network is described as follows:
[0073] F m1 = f m1 (Image before change)
[0074] F m2 = f m2 (Image after change)
[0075] where f m1 (·) and f m2 (·) represent two different modal mapping branches of the pseudo-twin network, and F m1 and F m2 are respectively the representative features learned through the pseudo-twin network. In addition, a difference information learning network f hd (·) that can maintain high-resolution features learns the difference information contained in the stitched image, and the process of extracting the difference information can be expressed as:
[0076] F d = f hd (Stitched image)
[0077] F d represents the difference information feature. To ensure that the three independent features F m1 , F m2 and F d are all in the same comparable space, these three features are simultaneously input into the UMU to obtain a comparable feature map, and the process can be expressed as:
[0078] F′ m1 , F′m2 , F' d = U(F m1 , F m2 , F d )
[0079] In F' m1 , F' m2 and F' d ∈ R N×N×K , belonging to the same similar space, where U represents the unified mapping unit. During the entire training phase, the cross-entropy function is used to evaluate whether the pseudo-twin network has sufficiently obtained the feature maps that effectively represent the images.
[0080] Considering that the training dataset is unlabeled and the network parameters cannot be adjusted according to the labels, pseudo-labels need to be introduced to ensure that the pseudo-twin network can capture the features of the dual-temporal image pairs. The K dimensions of F' m1 are converted into a one-dimensional label C through the argmax function m1 as the pseudo-label of F' m1 . In the experiment, the loss function of the pseudo-twin network can be expressed as:
[0081] L1 = crossentropy(F', m1 C m1 )
[0082] L2 = crossentropy(F', m2 C m2 )
[0083] where
[0084]
[0085]
[0086] In addition to requiring the features extracted by the pseudo-twin network to be representative, it is also expected that the obtained dual-temporal feature map pairs have sufficient specificity (distinguishability). For this purpose, differential information is used to supervise the feature maps output by the pseudo-twin network. The loss function adopted by differential supervision can be expressed as:
[0087]
[0088] d i,j represents the distance between the pixels corresponding to the feature maps F' m1 and F' m2 at the coordinate (i, j), y i,j represents the value of F' d at the coordinate (i, j), and Margin represents a manually set threshold, which is set to strengthen the distance of the feature map pairs.
[0089] Table 1 Algorithm Flow of the Change Detection Framework Based on Self-Supervised Learning
[0090]
[0091] Step 4: Network Inference and Generation of the Change Map
[0092] After the training phase, in an ideal situation, the feature maps extracted from the multi-modal dual-temporal images are directly applicable to the subsequent inference phase. In the inference phase, the traditional threshold segmentation method (specifically, the OSTU threshold algorithm is adopted in the present invention) is applied to the difference feature map for threshold segmentation, thereby obtaining the final change map under unsupervised conditions. It should be noted that in this phase, the adopted threshold segmentation method can be replaced by any other change detection algorithm based on traditional methods because the feature image pairs have already obtained robustness for subsequent inference.
[0093] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A multi-modal remote sensing image change detection method based on self-supervised learning, characterized in that The steps are as follows: Step 1: Feature map extraction Take the image before the change in modality one, the image after the change in modality two, and the spliced image as three independent inputs, and input them into the ternary feature extraction network to obtain the feature maps F m1 , F m2 and F d ; The spliced image is obtained by stacking the image before the change in modality one and the image after the change in modality two along the channel dimension; The ternary feature extraction network system consists of two parts: a pseudo-twin network and a differential information network; Each branch of the pseudo-twin network consists of 5 convolutional layers with a convolutional kernel size of 3×3. After each convolutional layer, a batch normalization layer and a rectified linear unit activation function are used to maintain gradient stability, prevent network overfitting, and enhance the network's ability to learn non-linear features. The pre-change image of modality one and the post-change image of modality two are input into the pseudo-twin network to extract features, obtaining feature maps F m1 and F m2 ; The differential information network includes four stages. The first stage contains 4 residual blocks and a 3×3 kernel convolutional layer. In the second stage, each branch processes the feature map at different scales. These branches operate independently and are composed of multiple consecutive residual blocks. The third and fourth stages mimic the structure of the second stage. When implementing the fusion strategy for feature maps of different resolutions, the upsampling part uses bilinear upsampling operation followed by a 1×1 convolution, and the downsampling uses a convolutional layer with a kernel size of 3×3 and a stride of 2. The spliced image is input into the differential information network structure to extract features, obtaining the feature map F d ; Step 2: Feature map unified space The feature maps F m1 , F m2 and F d are mapped to a comparable feature space through the unified mapping unit UMU to obtain the feature maps F' m1 , F' m2 and F' d ; Step 3: Network self-supervised training In the training stage of the network, cross-entropy loss is used to supervise the effectiveness of the image feature maps before and after the change, and contrastive loss is used to ensure the self-supervised training of the entire network without labels; Step 4: Network inference and generation of the change map The obtained features are analyzed using a threshold segmentation algorithm to generate the final change map.
2. The multi-modal remote sensing image change detection method based on self-supervised learning according to claim 1, wherein: The unified mapping unit consists of a token encoder and a token decoder, The input of the token decoder is three independent features F obtained by the triple feature extraction network m1 , F m2 and F d ; the input feature map is represented as F ∈ R b×c×h×w , which is converted into a three-dimensional token embedding vector of a specific size, with dimensions b × l × c; b, c, h, and w represent the batch size, the number of channels, and the height and width of the input features respectively, and l represents the token length; Encoding process of the token encoder: 3D token embedding is in the encoder to capture context information in the global; during this process, a set of trainable parameters are added to the token for positional embedding PE; the encoder follows the standard transformer structure, including a multi-head attention MHA module and a feed-forward neural network module; In addition, layer normalization LN is applied before each module; thereby, a token embedding vector is obtained, denoted as T ∈ R b×l×c ; Inputs to the token decoder: The token decoder receives two different inputs; one is the feature map F obtained from the convolutional network, which can also be considered as the feature map extracted by the ternary feature extraction network; the other input is the token embedding vector T, which contains the global context information generated by the token encoder; Decoding process of the token decoder: The token decoder is similar to the structure of the token encoder, using PE to endow the original convolutional feature F with positional information; It consists of multiple layers, each layer containing a combination of self-attention and a feed-forward neural network; two key subroutines are given below: Layer normalization LN: Before each decoder layer, layer normalization is applied to normalize the features, thereby enhancing training stability; Multi-head attention MHA: The decoder adopts a multi-head attention mechanism, aiming to understand the relationships between different tokens, thereby enriching context understanding; there are differences between this MHA and the MHA used in the token encoder; among them, Query is derived from the convolutional feature F, while Key and Value are derived from the token embedding vector T.
3. The multi-modal remote sensing image change detection method based on self-supervised learning according to claim 2, wherein: The function of the cross-entropy loss is expressed as: l1 = crossentropy(F' m1 , C m1 ) L2 = crossentropy(F' m2 , C m2 ) Among them, C m1 is the pseudo-label of F' m1 , and C m2 is the pseudo-label of F' m2 .
4. The multi-modal remote sensing image change detection method based on self-supervised learning according to claim 3, wherein: The function of the contrastive loss is expressed as: Among them, d i,j represents the distance between the pixels corresponding to the feature maps F' m1 and F' m2 at the coordinates (i, j), and y i,j represents the value of F' d at the coordinates (i, j). Margin represents a manually set threshold, and setting this threshold is to strengthen the distance between the feature map pairs.
5. The multi-modal remote sensing image change detection method based on self-supervised learning according to claim 1, wherein: The threshold segmentation algorithm mentioned is the OSTU threshold algorithm.
6. A computer system, characterized in that Including: One or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
7. A computer-readable storage medium, characterized in that Stored with computer-executable instructions, the instructions are used to implement the method according to any one of claims 1-5 when executed.
Citation Information
Patent Citations
Image similarity detection method based on text fusion under label-free sample
CN114298159A
Model training method and device, equipment and storage medium
CN114528762A