A weakly supervised change detection method for remote sensing images based on deep learning

By combining the attention refinement module and change prior constraints in the remote sensing image change detection, the problem of complex detection steps and relying on high-quality labeled data in the prior art is solved, and end-to-end, simple and efficient remote sensing image weak supervision change detection is realized, which significantly improves detection accuracy and efficiency.

CN119399647BActive Publication Date: 2025-05-13INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411602578.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-05-13
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The prior art cannot provide a method for weakly supervised change detection of remote sensing images that can achieve end-to-end, sufficient supervision and simple and efficient detection steps. Especially in high-resolution images, the changing lands are usually small and numerous, and it takes time and effort to mark the changing areas by pixel.

Method used

By combining the attention refinement module and change prior constraints, a weakly supervised change detection method for remote sensing images based on deep learning is proposed. The method includes obtaining the target dual-time image, extracting the multi-head self-attention and feature map, fusing the feature map and multi-head self-attention, generating the initial pseudo-label, and generating the final pseudo-label through random walk propagation and change prior constraints, which is finally decoded by the decoder.

Benefits of technology

This method significantly improves the accuracy and efficiency of weakly supervised change detection of remote sensing images, optimizes the overall target loss function, improves the performance of the end-to-end decoding framework, and surpasses the current most advanced weakly supervised change detection method of image-level labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399647B_ABST
    Figure CN119399647B_ABST
Patent Text Reader

Abstract

This application relates to the field of deep learning technology, and particularly to a weakly supervised change detection method for remote sensing images based on deep learning, including obtaining target dual-temporal images T1 and T2; extracting multi-head self-attention and feature maps from the target dual-temporal images T1 and T2 to obtain multi-head self-attention A A , A B and feature map F A , F B ; perform fusion on the multi-head self-attention A A , A B , perform fusion on the feature map F A , F B to obtain multi-head self-attention A f and feature map F f ; based on the feature map F f , obtain the class activation mapping CAM, perform fixed-threshold discrimination on the class activation mapping CAM to obtain the initial pseudo-label; perform generation of change attention degree C on the multi-head self-attention A f ; perform random walk propagation on the class activation mapping CAM based on the change attention degree C to obtain the propagated class activation mapping CAM; generate the final pseudo-label based on the change prior constraint; perform decoding based on the decoder. This application uses an attention refinement module, combines prior knowledge to design a change prior constraint, and significantly outperforms the current weakly supervised change detection methods at the image-level label
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition, computing or deep learning technology, and in particular to a remote sensing image weakly supervised change detection method based on deep learning. Background Art

[0002] Change detection is a key remote sensing image analysis technology that aims to identify changes between remote sensing images of the same area acquired at different times. This technology is widely used in many fields such as urban planning, environmental monitoring, and disaster assessment, helping experts and decision makers obtain intuitive information about surface changes so as to make timely and effective responses. With the rapid development of deep learning technology, its application in change detection tasks has significantly improved detection accuracy and efficiency. In the field of remote sensing change detection (CD), convolutional neural networks and visual transformer networks that rely on fully supervised learning can automatically identify and locate change areas from remote sensing images by learning complex features from a large amount of annotated data, thereby achieving highly automated and intelligent change detection.

[0003] However, the success of most deep learning models depends on a large amount of accurately annotated training data. In the field of change detection, annotating the changed areas pixel by pixel is not only time-consuming and labor-intensive, but also requires domain expertise, which makes the acquisition of high-quality annotated data a major challenge. In addition, the changed objects in high-resolution images are usually small and numerous, and their pixel-level annotations are more complex than those in medium and low-resolution images. Summary of the invention

[0004] The inventors found through research that in recent years, with the diversification of high-resolution remote sensing image acquisition methods, researchers have begun to explore weakly supervised change detection methods using cheap labels. Cheap labels include image-level labels, point annotations, bounding boxes and other labels that are relatively easy to obtain and low-cost. By using cheap labels, weakly supervised change detection aims to reduce the reliance on precise pixel-level annotations. Among cheap labels, image-level labels are the most convenient form of labels to obtain. They are expressed in binary form and do not require professional knowledge to obtain. However, due to the lack of detailed information on the specific location of the change, weakly supervised change detection using image-level labels is the most challenging task.

[0005] The purpose of this application is to provide a weakly supervised change detection method for remote sensing images based on deep learning. By combining an attention refinement module with a change prior constraint, the technical problem that the prior art cannot provide a detection method that can be implemented end-to-end, has sufficient supervision, and has simple and efficient detection steps is solved.

[0006] According to one aspect of the present application, a remote sensing image weakly supervised change detection method based on deep learning is provided, the method is executed by a processor, and includes: obtaining target dual-time images T1 and T2; extracting at least multi-head self-attention and feature maps from the target dual-time images T1 and T2 to obtain a multi-head self-attention A A , A B And the feature map F A 、F B ; For multi-head self-attention A A , A B Perform fusion on the feature map F A 、F B Perform fusion to obtain multi-head self-attention A f And the feature map F f ; Based on the feature map F f , obtain the category activation map CAM, perform fixed threshold differentiation on the category activation map CAM, and obtain the initial pseudo label; for the multi-head self-attention A f Execute and generate the change attention C; perform random walk propagation on the category activation map CAM based on the change attention C to obtain the propagated category activation map CAM; generate the final pseudo label based on the change prior constraint;

[0007] Perform decoding based on the decoder for the final pseudo-label.

[0008] In some embodiments, the target dual-time images T1 and T2 are at least extracted with multi-head self-attention and feature maps, specifically: the target dual-time images T1 and T2 are uploaded to a layered encoder to obtain the multi-head self-attention and feature maps corresponding to the last layer in the layered encoder.

[0009] In some embodiments, the multi-head self-attention A A , A B Perform fusion to obtain multi-head self-attention A f Specifically, use the absolute difference operation to perform multi-head self-attention A A , A B The fusion of multi-head self-attention A f .

[0010] In some embodiments, the multi-head self-attention A f It is expressed as:

[0011] A f =A A -A B

[0012] In some embodiments, the feature map F A 、F B Perform fusion to obtain feature map F fSpecifically, the absolute difference operation is used to perform the feature map F A 、F B The fusion of f .

[0013] In some embodiments, the feature map F f It is expressed as:

[0014] F f =F A -F B

[0015] In some embodiments, the multi-head self-attention A f Execute and generate a change attention degree C, wherein the attention degree C is:

[0016]

[0017] Among them, σ() is the sigmoid activation function; is the transposed matrix.

[0018] In some embodiments, in the random walk propagation performed on the class activation map CAM based on the change attention C, the random walk process is expressed as:

[0019]

[0020] Among them, C is the normalized propagation probability matrix.

[0021] Compared with the prior art, the present application has the following advantages and beneficial effects: the method of the present application uses an attention refinement module and designs a priori constraints on changes in combination with prior knowledge, which not only improves the performance of the attention refinement module, but also optimizes the overall target loss function and improves the performance of the entire end-to-end decoding framework; at the same time, the method of the present application significantly surpasses the most advanced weakly supervised change detection method for image-level labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 is a flow chart of the detection method of the present application;

[0024] Figure 2 It is a schematic diagram of the detection method of the present application;

[0025] Figure 3 It is a schematic diagram of the framework corresponding to the detection method of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the attached examples of the present application Figure 1-3 The technical solutions in the embodiments of the present application are clearly and completely described together. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0027] In order to better understand this application, some terms in this application are explained:

[0028] WSCD: Weakly Supervised Change Detection;

[0029] CAM: Category Activation Map;

[0030] AR: Attention Refinement Module;

[0031] CP: Variation prior constraints.

[0032] Application Overview

[0033] There are several different methods for WSCD. Khan et al and Andermatt et al used conditional random field (CRF) techniques to process the change features extracted by CNNs. Kalita et al used Kmeans clustering algorithm and classic principal component analysis (PCA) to segment the change feature maps generated by CNNs. Wu et al introduced generative adversarial networks (GAN) into WSCD. Huang et al used a data augmentation method to increase generalization. Although these techniques provide strategies to solve the WSCD task, their complexity masks the potential of utilizing the inherent information of neural networks. Generating class activation maps can visualize the degree of attention paid by neural networks to objects. Using the inherent information in neural networks, some advanced WSCD methods based on CNNs began to generate pseudo labels from class activation maps as pixel-level supervision signals to train CD networks. However, recent work has shown that class activation maps are often flawed. They can only focus on the most discriminative change areas, which greatly impairs the accuracy of change detection. One of the reasons is that the class activation maps generated by CNNs only perceive local features, while the multi-head self-attention in Transformers can achieve global feature interaction, thereby improving the problem that CNNs cannot activate the overall change area. In addition, the self-attention mechanism is essentially a directed graph model, which is also inherent information in the network, from which change information can be derived to refine the initial pseudo-labels generated by the class activation map. In addition, influenced by informed learning, we incorporate the underlying logic in WSCD into the refinement of the initial pseudo-labels by multi-head self-attention, that is, unchanged dual-temporal image pairs should not have changed pixel-level ground predictions and changed dual-temporal image pairs should not have unchanged pixel-level ground predictions.

[0034] Based on the above analysis of WSCD, in order to make full use of the inherent information of the neural network, this application extracts the multi-head self-attention and feature maps of each pair of dual-time images based on the Transformers encoder, and fuses them using absolute difference operations respectively. For the fused feature map, the class activation map is used to generate the initial pseudo-label, and then the attention refinement module is used to generate the change attention from the fused multi-head self-attention. Under the control of the change prior constraint, the attention refinement module uses the change attention to perform random walk propagation on the class activation map to obtain the refined class activation map after propagation and generate the final pseudo-label. In order to stabilize the training, the decoder first uses the initial pseudo-label as pixel-level supervision, and then uses the final pseudo-label as pixel-level supervision.

[0035] Exemplary Methods

[0036] A remote sensing image weakly supervised change detection method based on deep learning is performed by a processor, which may be a central processing unit (CPU), a graphics processing unit (GPU), other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0037] Specifically, it includes: obtaining target dual-time images T1 and T2; extracting at least multi-head self-attention and feature maps from the target dual-time images T1 and T2, obtaining multi-head self-attention A A , A B And the feature map F A 、F B ; For multi-head self-attention A A , A B Perform fusion on the feature map F A 、F B Perform fusion to obtain multi-head self-attention A f And the feature map F f ; Based on the feature map F f , obtain the category activation map CAM, perform fixed threshold differentiation on the category activation map CAM, and obtain the initial pseudo label; for the multi-head self-attention A f Execute and generate a change attention C; perform random walk propagation on the class activation map CAM based on the change attention C to obtain the propagated class activation map CAM; generate a final pseudo label based on the change prior constraint; and perform decoding on the final pseudo label based on the decoder. It should be noted that the change prior constraint at least includes prior knowledge and a variable threshold.

[0038] Below, each step is described in detail:

[0039] In some embodiments, reference Figure 3, since the method of this application focuses on proposing a weakly supervised change detection method framework with universal and cheap labels, only a minimalist segmentation head is used as a decoder in this embodiment. Then, following the mature change detection process, a dual-stream architecture is constructed to extract and fuse the features of the dual-time image pair. That is, at least multi-head self-attention and feature maps are extracted for the target dual-time images T1 and T2. Specifically: the target dual-time images T1 and T2 are uploaded to the layered encoder to obtain the multi-head self-attention and feature maps corresponding to the last layer in the layered encoder.

[0040] In some embodiments, for the multi-head self-attention A A , A B Perform fusion to obtain multi-head self-attention A f Specifically, use the absolute difference operation to perform multi-head self-attention A A , A B The fusion of multi-head self-attention A f .

[0041] Among them, multi-head self-attention A f It is expressed as:

[0042] A f =A A -A B

[0043] It should be noted that in order to obtain the change attention matrix, as shown in the attached Figure 3 First, we obtain multi-head self-attention from the target dual-time images T1 and T2. In order to highlight the change information as much as possible, we still use the absolute difference operation to fuse the multi-head self-attention between the dual-time images. The multi-head self-attention of the change area can be expressed as A f =A A -A B

[0044] Since multi-head self-attention is a directed graph model, nodes that share the same change information should be equal, so for multi-head self-attention A f Execute to generate a change in attention C, which is:

[0045]

[0046] Among them, σ() is the sigmoid activation function; is the transposed matrix.

[0047] In some embodiments, the feature map F A 、F B Perform fusion to obtain feature map F f Specifically, the absolute difference operation is used to perform the feature map F A 、FB The fusion of f .

[0048] Among them, the feature map F f It is expressed as:

[0049] F f =F A -F B

[0050] In some embodiments, in order to generate higher quality pseudo labels, the initial class activation map is refined using a random walk algorithm to obtain a finer activation area. The algorithm performs random walks based on the connection strength between nodes, and performs random walk propagation on the class activation map CAM based on the change attention C. The random walk process is expressed as:

[0051]

[0052] Among them, C is the normalized propagation probability matrix.

[0053] The above shows and describes the basic principles and main features of the present application and the advantages of the present application. For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or basic features of the present application. Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present application is defined by the attached claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved.

[0054] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A remote sensing image weakly supervised change detection method based on deep learning, the method is executed by a processor, characterized in that: include: Acquire target dual-time images T1 and T2; Extract at least the multi-head self-attention and feature map of the target dual-time images T1 and T2 to obtain the multi-head self-attention A A , A B And the feature map F A 、F B ; For multi-head self-attention A A , A B Perform fusion on the feature map F A 、F B Perform fusion to obtain multi-head self-attention A f And the feature map F f ; Based on the feature map F f , obtain the category activation map CAM, perform fixed threshold differentiation on the category activation map CAM, and obtain the initial pseudo label; For multi-head self-attention A f Execute to generate change attention C; Perform random walk propagation on the category activation map CAM based on the change attention C to obtain the propagated category activation map CAM; Generate the final pseudo label based on the change prior constraints; Decoding is performed based on the decoder for the final pseudo-labels.

2. The method according to claim 1, characterized in that The method of extracting at least multi-head self-attention and feature maps from the target dual-time images T1 and T2 is specifically as follows: uploading the target dual-time images T1 and T2 to a layered encoder to obtain the multi-head self-attention and feature maps corresponding to the last layer in the layered encoder.

3. The method according to claim 1, characterized in that The multi-head self-attention A A , A B Perform fusion to obtain multi-head self-attention A f Specifically, use the absolute difference operation to perform multi-head self-attention A A , A B The fusion of multi-head self-attention A f .

4. The method according to claim 3, characterized in that The multi-head self-attention A f It is expressed as: A f =|A A -A B |。 5. The method according to claim 1, characterized in that The pair feature map F A 、F B Perform fusion to obtain feature map F f Specifically, the absolute difference operation is used to perform the feature map F A 、F B The fusion of f .

6. The method according to claim 4, characterized in that The feature map F f It is expressed as: F f =|F A -F B |。 7. The method according to claim 1, characterized in that The multi-head self-attention A f Execute and generate a change attention degree C, wherein the attention degree C is: Among them, σ() is the sigmoid activation function; is the transposed matrix.

8. The method according to claim 1, characterized in that In the random walk propagation performed on the category activation map CAM based on the change attention C, the random walk process is expressed as: Among them, C is the normalized propagation probability matrix.

Citation Information

Patent Citations

  • Weakly supervised building segmentation method taking reliable region as attention mechanism supervision

    CN114820655A

  • SAR image change detection method based on multi-scale differential feature attention mechanism

    CN114926746A