Remote sensing target change detection method based on hybrid self-crossover attention mechanism
Through the remote sensing target change detection method with a hybrid self-cross attention mechanism, the problem of insufficient detection capability of tiny changes in remote sensing images is solved, and efficient and accurate change detection is achieved, which is suitable for multi-resolution remote sensing image analysis.
Patent Information
- Application Number
- CN202510282787.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has limited ability to detect small changes in the responsible context in remote sensing target change detection, making it difficult to effectively process images under different sensors or imaging conditions, and the processing efficiency of large-scale data is not high.
The remote sensing target change detection method based on a hybrid self-cross attention mechanism is adopted, and multi-scale features are extracted through a hierarchical convolutional structure, combined with the self-attention module and the cross-attention module, efficient fusion of space-time information is achieved, key features are strengthened and redundant information is suppressed.
It significantly improves the accuracy and robustness of change detection, especially in terms of weak change and small object detection, which can fully capture multi-scale, spatial and temporal changes characteristics and adapt to multi-resolution remote sensing image analysis.
Smart Images

Figure CN120388299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular, to a remote sensing target change detection method based on a hybrid self-cross attention mechanism. Background Art
[0002] With the rapid development of satellite remote sensing technology, it has become increasingly convenient to obtain a large amount of remote sensing image data in multiple time periods. This also provides opportunities for change detection in environmental monitoring (Literature 1: C. Song, B. Huang, L. Ke, and K. S. Richards, "Remote sensing of alpine lake water environment changes on the Tibetan Plateau and surroundings: A view," ISPRS J. Photogramm. Remote Sens., vol. 92, pp. 26-37, Jun. 2014, doi: 10.1016 / j.isprsjprs.2014.03.001.), urban planning (Literature 2: G. Xian and C. Homer, "Updating the 2001 National Land Cover Database impervious surf aceproducts to 2006 unsing Landsat imagery change detection methods,” Remote Sens. Environ., vol. 114, no. 8, pp. 1676 - 1686, Aug. 2010, doi: 10.1016 / j.rse.2010.02.018.)、urban land use dynamic monitoring (Literature 3 P.K. Mishra, A. Rai, and S.C. Rai, “Land use and land cover change detection using geospatial techniques in the Sikkim Himalaya, India,” Egypt. J. Remote Sens. Space Sci., vol. 23, no. 2, pp. 133 - 143, Aug. 2020, doi: 10 / 1016 / j.ejrs.2019.02.001.)、disaster assessment (Literature 4 P. Lu, Y. Qin, Z. Li, A.C. Mondini, and N. Casagli, “Landslide mapping from multi - sensor data through improved change detection - based Markov random field,” Remote Sens. Environ., vol. 231, Sep. 2019, Art. No. 111235, doi: 10.1016 / j.rse.2019.111235.) and other fields have been widely used. As one of the key technologies in remote sensing image analysis, the remote sensing target change detection task aims to identify and analyze the change information between remotely sensed images acquired at different times. Specifically, remote sensing image change detection is a key technology for identifying the differences between remotely sensed images acquired at different times in the same area (Literature 5 A. Singh, “Review article digital change detection techniques using remotely - sensed data,” ,Int.J. Remote Sens., vol. 10, no. 6, pp. 989 - 1003, jun. 1989, doi: 10.1080 / 01431168908903939.). For example, the increase in buildings caused by urban expansion, the change in vegetation cover in farmland due to drought or irrigation, the delineation of burned areas after forest fires, and the dynamic monitoring of flood inundation areas, etc.
[0003] Early traditional change detection methods were mainly based on simple operations at the pixel level. For example, the difference method directly calculates the gray difference of corresponding points in two temporal images and determines whether the pixel has changed by setting a threshold; the ratio method is similar, using the pixel gray ratio for judgment. However, these methods are extremely sensitive to factors such as lighting conditions, atmospheric scattering, and sensor noise. Slight differences in lighting angles and intensities at different times, or slight changes in atmospheric conditions, may lead to a large number of misjudgments, making the change detection results full of noise and greatly reducing the accuracy. To overcome the defects of pixel-level methods, feature-based change detection methods emerged. These methods first manually design and extract features such as texture, shape, and spectrum, and then compare the feature differences between different temporal phases to determine the change area. Although it enhances the robustness to noise to a certain extent, manually extracting features heavily relies on professional knowledge and experience, and appropriate features need to be redesigned for different scenarios, with poor generalization ability. For example, the texture features tried in urban areas are difficult to effectively capture change information in mountainous areas with dense vegetation and cannot meet the diverse remote sensing application requirements. In recent years, deep learning technology has achieved obvious advantages in the field of remote sensing image processing with its powerful automatic feature learning ability. Among many methods, convolutional neural networks (CNNs) are the main force, such as the classic U-Net (Reference 6 Da udtRC, Saux B L, Boulch A, 2018. Fully Convolutional Siamese Networks for Change Detection [A]. arXiv.) and its variants. With the encoder-decoder structure, they have shone in the field of remote sensing images. The encoder gradually extracts the deep features of the image, and the decoder restores the resolution and accurately locates the changes, which is extremely accurate for target change detection in complex scenarios such as urban expansion and forest degradation and can effectively handle the interference of lighting and scale changes. The participation of generative adversarial networks (GANs) further improves the performance. The generator generates change samples, and the discriminator determines the authenticity. The adversarial game between the two enables the model to learn more discriminative features, which has obvious advantages in detecting small changes and progressive changes, reducing the problems of missed detection and false detection caused by insufficient feature extraction in conventional methods. However, the existing deep learning-based change detection methods have the following problems: 1) Limited ability to detect small changes in complex backgrounds; 2) Difficulty in effectively processing images under different sensors or imaging conditions; 3) Low processing efficiency for large-scale data. Summary of the Invention
[0004] The present invention provides a remote sensing target change detection method based on a hybrid self-cross attention mechanism, which can solve the technical problems in the prior art, such as limited ability to detect small changes in complex backgrounds, difficulty in effectively processing images under different sensors or imaging conditions, and low processing efficiency for large-scale data.
[0005] According to one aspect of the present invention, a remote sensing target change detection method based on a hybrid self-cross attention mechanism is provided. The remote sensing target change detection method based on the hybrid self-cross attention mechanism includes: Step 1, input the first-phase remote sensing image T1 and the second-phase remote sensing image T2, and use the first convolutional structure F1, the second convolutional structure F2, the third convolutional structure F3, and the fourth convolutional structure F4 to extract features from the first and second-phase remote sensing images T1 and the second-phase remote sensing image T2 respectively; Step 2, the features of the first-phase remote sensing image T1 and the features of the second-phase remote sensing image T2 extracted in Step 1 enter the first hybrid self-cross attention fusion module HSC-AFM at the same time. The first hybrid self-cross attention fusion module HSC-AFM includes a self-attention module, a cross-attention module, and a fusion module. First, use the self-attention module to extract the global context information in the single-phase image and capture the correlation between different regions in the image; Secondly, input the output features of the self-attention module and the output feature map of the fourth convolutional structure F4 into the cross-attention module together, and use the cross-attention module to focus on the time-varying features and learn the spatio-temporal interaction information between images at different times; Finally, input the output features of the cross-attention module and the output feature map of the third convolutional structure F3 into the fusion module together. The fusion module integrates the two attention mechanisms of the self-attention module and the cross-attention module to achieve efficient fusion of spatio-temporal information; Step 3, input the result extracted in Step 2 and the output feature map of the second convolutional structure F2 into the second hybrid self-cross attention fusion module HSC-AFM together to strengthen the features of the change information; Step 4, input the result extracted in Step 3 and the output feature map of the first convolutional structure F1 into the third hybrid self-cross attention fusion module HSC-AFM together. The third hybrid self-cross attention fusion module HSC-AFM perceives the differences between multi-scale features, strengthens the key features, and suppresses redundant information at the same time; Step 5, output the features in Step 4 to the classifier, and after passing through the classifier, output a binary prediction map to complete the remote sensing target change detection task.
[0006] Further, the self-attention module and the cross-attention module constitute the hybrid self-cross attention module. A part of the output feature map of the fourth convolutional structure F4 is input to the self-attention module, and another part of the output feature map of the fourth convolutional structure F4 is input to the cross-attention module.
[0007] Further, the self-attention module is used to perform convolution operations, self-attention calculations, and linear transformations; the convolution operation specifically includes: performing a convolution operation on a part of the output feature map of the fourth convolution structure F4 to obtain feature maps Q1, K1, and V1, where the sizes of feature maps Q1 and V1 are both (h×w)×c, and the size of feature map K1 is c×(h×w); the self-attention calculation specifically includes: calculating the dot product of the obtained feature maps Q1 and K1 to obtain an attention weight matrix X1 with a size of c×(h×w); performing a Softmax operation on the attention weight matrix X1 to obtain a normalized attention weight; multiplying the normalized attention weight by the feature map V1 to obtain a feature map B with a size of (h×w)×c; the linear transformation specifically includes: performing a linear transformation on the feature map B to obtain a feature map Q2 with a size of (h×w)×c, and the linear transformation integrates features and provides decision information for the change detection task.
[0008] Further, the cross-attention mechanism calculation of the cross-attention module specifically includes: performing a linear operation on another part of the output feature map of the fourth convolution structure F4 to obtain feature maps K2 and V2, calculating the dot product of the feature maps Q2 and K2 to obtain an attention weight matrix X2 with a size of c×(h×w); performing a softmax operation on the attention weight matrix X2 to obtain a normalized attention weight; multiplying the normalized attention weight by the feature map V2 to obtain a feature map M with a size of h×w×C.
[0009] Further, the fusion module includes an upsampling unit and a convolution block. The upsampling unit is used to perform an upsampling operation on the feature map M to obtain a feature map with a size of 2h×2w×C, and the convolution block is used to perform a convolution operation on the upsampled feature map and the output feature map of the third convolution structure F3 to obtain a final output feature map O with a size of 2h×2w×C.
[0010] Further, the remote sensing target change detection method based on the hybrid self-cross attention mechanism further includes: using a loss function to verify the model effect.
[0011] Further, the loss function uses the Dice loss function, and the Dice loss function is where y i is the predicted value of the true label, is the predicted value of pixel i, and N represents the total number of pixels.
[0012] According to another aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps of the remote sensing target change detection method based on the hybrid self-cross attention mechanism as described above.
[0013] According to another aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the remote sensing target change detection method based on the hybrid self-cross attention mechanism as described above are implemented.
[0014] Applying the technical solution of the present invention, there is provided a remote sensing target change detection method based on a hybrid self-cross attention mechanism. The method uses a hierarchical convolutional structure to extract features from the input multi-temporal remote sensing images (T1 and T2). By extracting multi-scale features layer by layer, the model can effectively capture spatial information from local to global, laying a solid foundation for subsequent spatio-temporal fusion. Compared with traditional single-scale methods, the present invention has higher expressive power in multi-scale modeling. Secondly, the network adopts a hybrid self-cross attention mechanism (HybridSelf-CrossAttentionFusionModule, HSC-AFM), including a self-attention module and a cross-attention module. The self-attention module focuses on extracting global context information in a single-time image and capturing the correlation between different regions in the image. The cross-attention module focuses on time-varying features and learns spatio-temporal interaction information between different-time images. The fusion module integrates these two attention mechanisms together to achieve efficient fusion of spatio-temporal information and significantly improve the model's ability to understand complex change patterns. The HSC-AFM module can perceive the differences between multi-scale features, strengthen key features, and at the same time suppress redundant information. This mechanism not only effectively enhances the feature representation of the change region but also retains the target edge details, greatly improving the accuracy of change detection. Through these innovative designs, the present invention can comprehensively capture multi-scale and spatio-temporal change features, showing excellent change detection capabilities in complex scenes. Especially in the detection of weak changes and small targets, the network shows stronger robustness, providing an efficient and accurate solution for high-resolution remote sensing image analysis. Therefore, compared with the prior art, the change detection network based on the hybrid self-cross attention mechanism provided by the present invention fuses global and cross-scale information of the target region through the hybrid self-attention mechanism and the cross-attention mechanism; proposes a hierarchical scale-aware feature module, which strengthens the key features of target change information by perceiving the differences between multi-scale features and suppresses redundant information at the same time; constructs an efficient target change detection network structure that adapts to multi-resolution remote sensing images, and verifies its superior performance on two public datasets. Brief Description of the Drawings
[0015] The accompanying drawings included are used to provide a further understanding of the embodiments of the present invention, which form a part of the specification, illustrate the embodiments of the present invention, and explain the principles of the present invention together with the written description. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 Shows a change detection network model diagram based on a hybrid self-cross attention mechanism provided according to a specific embodiment of the present invention;
[0017] Figure 2 Shows a schematic diagram of a hybrid self-cross attention mechanism module and a fusion module provided according to a specific embodiment of the present invention;
[0018] Figure 3 Shows the visualization effect diagrams on the LEVIR-CD and WHU-CD data sets provided according to a specific embodiment of the present invention. Detailed implementation manners
[0019] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0020] It should be noted that the terms used here are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present application. As used here, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0021] Unless otherwise specifically noted, the relative arrangement of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the authorized specification. In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as limiting. Thus, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.
[0022] As Figures 1 to 3As shown, according to a specific embodiment of the present invention, a remote sensing target change detection method based on a hybrid self-cross attention mechanism is provided. The remote sensing target change detection method based on the hybrid self-cross attention mechanism includes: Step 1, input the first-phase remote sensing image T1 and the second-phase remote sensing image T2, and use the first convolutional structure F1, the second convolutional structure F2, the third convolutional structure F3, and the fourth convolutional structure F4 to perform feature extraction on the first and second-phase remote sensing images T1 and the second-phase remote sensing image T2 respectively; Step 2, the features of the first-phase remote sensing image T1 and the features of the second-phase remote sensing image T2 extracted in Step 1 enter the first hybrid self-cross attention fusion module HSC-AFM at the same time. The first hybrid self-cross attention fusion module HSC-AFM includes a self-attention module, a cross-attention module, and a fusion module. First, use the self-attention module to extract the global context information in the single-phase image and capture the correlation between different regions in the image; Secondly, input the output features of the self-attention module and the output feature map of the fourth convolutional structure F4 into the cross-attention module together, and use the cross-attention module to focus on the time change features and learn the spatio-temporal interaction information between images at different times; Finally, input the output features of the cross-attention module and the output feature map of the third convolutional structure F3 into the fusion module together. The fusion module integrates the two attention mechanisms of the self-attention module and the cross-attention module to achieve efficient fusion of spatio-temporal information; Step 3, input the result extracted in Step 2 and the output feature map of the second convolutional structure F2 into the second hybrid self-cross attention fusion module HSC-AFM together to strengthen the features of the change information; Step 4, input the result extracted in Step 3 and the output feature map of the first convolutional structure F1 into the third hybrid self-cross attention fusion module HSC-AFM together. The third hybrid self-cross attention fusion module HSC-AFM perceives the differences between multi-scale features, strengthens the key features, and suppresses redundant information at the same time; Step 5, output the features in Step 4 to the classifier, and after passing through the classifier, output a binary prediction map to complete the remote sensing target change detection task.
[0023] Applying this configuration method, a remote sensing target change detection method based on a hybrid self-cross attention mechanism is provided. This method uses a hierarchical convolutional structure to extract features from the input multi-temporal remote sensing images (T1 and T2). By extracting multi-scale features layer by layer, the model can effectively capture spatial information from local to global, laying a solid foundation for subsequent spatio-temporal fusion. Compared with traditional single-scale methods, the present invention has higher expressive power in multi-scale modeling. Secondly, the network adopts a hybrid self-cross attention mechanism (Hybrid Self-Cross Attention Fusion Module, HSC-AFM), including a self-attention module and a cross-attention module. The self-attention module focuses on extracting global context information in a single-time image and capturing the correlation between different regions in the image. The cross-attention module focuses on time-varying features and learns spatio-temporal interaction information between different-time images. The fusion module integrates these two attention mechanisms to achieve efficient fusion of spatio-temporal information and significantly improves the model's ability to understand complex change patterns. The HSC-AFM module can perceive the differences between multi-scale features, strengthen key features, and at the same time suppress redundant information. This mechanism not only effectively enhances the feature representation of the change region but also retains the target edge details, greatly improving the accuracy of change detection. Through these innovative designs, the present invention can comprehensively capture multi-scale and spatio-temporal change features, showing excellent change detection capabilities in complex scenarios. Especially in the detection of weak changes and small targets, the network shows stronger robustness, providing an efficient and accurate solution for high-resolution remote sensing image analysis. Therefore, compared with the prior art, the change detection network based on the hybrid self-cross attention mechanism provided by the present invention fuses the global and cross-scale information of the target region through the hybrid self-attention mechanism and the cross-attention mechanism; proposes a hierarchical scale-aware feature module, which strengthens the key features of the target change information by perceiving the differences between multi-scale features and suppresses redundant information at the same time; constructs an efficient target change detection network structure that adapts to multi-resolution remote sensing images, and verifies its superior performance on two public datasets.
[0024] Furthermore, in the present invention, the self-attention module and the cross-attention module constitute the hybrid self-cross attention module. A part of the output feature map of the fourth convolutional structure F4 is input to the self-attention module, and another part of the output feature map of the fourth convolutional structure F4 is input to the cross-attention module.
[0025] Among them, the self-attention module is used to perform convolution operations, self-attention calculations, and linear transformations. The convolution operation specifically includes: performing a convolution operation on a part of the output feature map of the fourth convolution structure F4 to obtain feature maps Q1, K1, and V1. The sizes of feature maps Q1 and V1 are both (h×w)×c, and the size of feature map K1 is c×(h×w). The self-attention calculation specifically includes: calculating the dot product of the obtained feature maps Q1 and K1 to obtain an attention weight matrix X1 with a size of c×(h×w); performing a Softmax operation on the attention weight matrix X1 to obtain a normalized attention weight; multiplying the normalized attention weight by feature map V1 to obtain feature map B with a size of (h×w)×c. The linear transformation specifically includes: performing a linear transformation on feature map B to obtain feature map Q2 with a size of (h×w)×c. The linear transformation integrates features and provides decision-making information for the change detection task.
[0026] Further, the cross-attention mechanism calculation of the cross-attention module specifically includes: performing a linear operation on another part of the output feature map of the fourth convolution structure F4 to obtain feature maps K2 and V2, calculating the dot product of feature maps Q2 and K2 to obtain an attention weight matrix X2 with a size of c×(h×w); performing a softmax operation on the attention weight matrix X2 to obtain a normalized attention weight; multiplying the normalized attention weight by feature map V2 to obtain feature map M with a size of h×w×C.
[0027] The fusion module includes an upsampling unit and a convolution block. The upsampling unit is used to perform an upsampling operation on feature map M to obtain a feature map with a size of 2h×2w×C. The convolution block is used to perform a convolution operation on the upsampled feature map and the output feature map of the third convolution structure F3 to obtain a final output feature map O with a size of 2h×2w×C.
[0028] Further, in the present invention, the remote sensing target change detection method based on the hybrid self-cross attention mechanism further includes: using a loss function to verify the model effect. Among them, the loss function uses the Dice loss function, and the Dice loss function is where y i is the predicted value of the true label, is the predicted value of pixel i, and N represents the total number of pixels.
[0029] According to another aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the remote sensing target change detection method based on the hybrid self-cross attention mechanism as described above.
[0030] According to another aspect of the present invention, there is provided a computer-readable storage medium storing a computer program which, when executed by a processor, implements the steps of the remote sensing target change detection method based on the hybrid self-cross attention mechanism as described above.
[0031] To further understand the present invention, the following combines Figures 1 to 3 to detail the remote sensing target change detection method based on the hybrid self-cross attention mechanism provided by the present invention.
[0032] As Figures 1 to 3 shown, according to a specific embodiment of the present invention, there is provided a remote sensing target change detection method based on a hybrid self-cross attention mechanism. The main innovations are as follows:
[0033] 1) A change detection network based on a hybrid self-cross attention mechanism is proposed, which fuses the global and cross-scale information of the target area through the hybrid self-attention mechanism and the cross-attention mechanism;
[0034] 2) A hierarchical scale-aware feature module is proposed, which strengthens the key features of the target change information and suppresses redundant information by perceiving the differences between multi-scale features;
[0035] 3) An efficient target change detection network structure is constructed to adapt to multi-resolution remote sensing images, and its superiority is verified on two public datasets.
[0036] The specific process of the algorithm is as follows:
[0037] Step 1, input the first-phase remote sensing image T1 and the second-phase remote sensing image T2, and use the first convolutional structure F1, the second convolutional structure F2, the third convolutional structure F3, and the fourth convolutional structure F4 to extract features from the first and second-phase remote sensing images T1 and the second-phase remote sensing image T2 respectively;
[0038] Step 2: The features of the first-phase remote sensing image T1 and the features of the second-phase remote sensing image T2 extracted in Step 1 enter the first Hybrid Self-Cross Attention Fusion Module (HSC-AFM) simultaneously. The HSC-AFM includes a self-attention module, a cross-attention module, and a fusion module. First, the self-attention module is used to extract the global context information in the single-phase image and capture the correlation between different regions in the image. Secondly, the output features of the self-attention module and the output feature map of the fourth convolutional structure F4 are jointly input into the cross-attention module, and the cross-attention module is used to focus on the time-varying features and learn the spatio-temporal interaction information between images at different times. Finally, the output features of the cross-attention module and the output feature map of the third convolutional structure F3 are jointly input into the fusion module, and the fusion module integrates the two attention mechanisms of the self-attention module and the cross-attention module to achieve efficient fusion of spatio-temporal information.
[0039] Step 3: The result extracted in Step 2 and the output feature map of the second convolutional structure F2 are jointly input into the second Hybrid Self-Cross Attention Fusion Module (HSC-AFM) to strengthen the features of the change information.
[0040] Step 4: The result extracted in Step 3 and the output feature map of the first convolutional structure F1 are jointly input into the third Hybrid Self-Cross Attention Fusion Module (HSC-AFM). The third HSC-AFM perceives the differences between multi-scale features, strengthens the key features, and suppresses the redundant information simultaneously.
[0041] Step 5: The output features in Step 4 are input to the classifier, and after passing through the classifier, a binary prediction map is output to complete the remote sensing target change detection task.
[0042] As Figure 2 shown, HSC-AFM represents the Hybrid Self-Cross Attention Fusion Module. The first, second, and third Hybrid Self-Cross Attention Fusion Modules have the same composition. HSC-AFM is the core module in the present invention. HSC-AFM is composed of a hybrid self-cross attention module and a fusion module. In the hybrid self-cross attention module, the upper half with F4 as the input is the self-attention module, and the lower half with F4 as the input is the cross-attention module. The self-attention module includes a convolutional operation, self-attention calculation, and linear transformation.
[0043] Convolutional operation: Perform a convolutional operation on the input feature map F4 to obtain feature maps Q1, K1, and V1. The sizes of feature maps Q1 and V1 are both (h×w)×c, and the size of feature map K1 is c×(h×w).
[0044] Self-attention calculation: Calculate the dot product of the feature maps Q1 and K1 obtained above to obtain the attention weight matrix X1, with a size of c×(h×w); perform a Softmax operation on the attention weight matrix X1 to obtain the normalized attention weights; multiply the normalized attention weights by the feature map V1 to obtain the feature map B, with a size of (h×w)×c;
[0045] Linear transformation: Perform a linear transformation on the feature map B to obtain the feature map Q2, with a size of (h×w)×c. The linear transformation integrates features and provides decision-making information for the change detection task.
[0046] Between the two linear transformations, there is a skip connection that allows the network to directly transfer information from the convolutional layer to the subsequent layers. This structure is beneficial to the propagation of gradients and reduces the problem of gradient vanishing. After the skip connection, there is an addition operation that adds the output of the convolutional layer to the output of the linear layer, which helps the network learn richer feature representations. On the right side of the figure, there is a separate convolutional block that represents more complex convolutional operations for specific feature extraction or processing tasks. The design of this module enables it to efficiently extract and process features in the change detection task, capture spatial features through the convolutional layer, integrate features through the linear layer, and enhance the expressiveness of features through skip connections and addition operations. This structure helps improve the accuracy and robustness of change detection.
[0047] Cross-attention mechanism calculation: Perform a linear operation on another part of the output feature map of the fourth convolutional structure F4 to obtain the feature maps K2 and V2, calculate the dot product of the feature maps Q2 and K2 to obtain the attention weight matrix X2, with a size of c×(h×w); perform a softmax operation on the attention weight matrix X2 to obtain the normalized attention weights; multiply the normalized attention weights by the feature map V2 to obtain the feature map M, with a size of h×w×C.
[0048] The fusion module includes upsampling and a convolutional block.
[0049] Upsampling: Perform an upsampling operation on M to obtain a feature map with a size of 2h×2w×C.
[0050] Convolutional block: Perform a convolutional operation on the upsampled feature map and the output feature map of the third convolutional structure F3 to obtain the final output feature map O, with a size of 2h×2w×C.
[0051] After obtaining the prediction map in the above process, the difference between the model prediction result and the ground truth is measured by using the Dice Loss function, thereby guiding the training and optimization of the model. In the model effect verification stage, the Dice Loss function is used as the loss function, and its calculation is as follows.
[0052]
[0053] where y i and represent the true label and the predicted value of pixel i, respectively. N represents the total number of pixels, calculated as the product of the number of pixels per image and the batch size.
[0054] In summary, to solve these problems, the present invention proposes a remote sensing target change detection method based on a hybrid self-cross attention mechanism. The main innovations are as follows:
[0055] 1) A change detection network based on a hybrid self-cross attention mechanism is proposed, which fuses global and cross-scale information of the target region through a hybrid self-attention mechanism and a cross-attention mechanism;
[0056] 2) A hierarchical scale-aware feature module is proposed, which strengthens the key features of target change information and suppresses redundant information by perceiving the differences between multi-scale features;
[0057] 3) An efficient target change detection network structure is constructed to adapt to multi-resolution remote sensing images, and its superiority is verified on two public datasets.
[0058] Our proposed model was evaluated on two publicly available change detection (CD) datasets: WHU-CD and LEVIR-CD. For the WHU-CD dataset, it was divided into non-overlapping patches of size 256×256, generating 4536, 504, and 2760 patch pairs for training, validation, and testing, respectively. Similarly, the LEVIR-CD dataset was divided into non-overlapping patches of the same size, generating training, validation, and test sets containing 7120, 1024, and 2048 samples, respectively. Our model was implemented using PyTorch and trained on a single NVIDIA RTX 3090 GPU. We used the AdamW optimizer (weight decay of 0.0025 and initial learning rate of 5e-4) to optimize the loss function. The batch size was set to 8, and the training process spanned 50 epochs. For easier and more intuitive comparison, we used metrics such as F1-score (F1), precision (Pre.), recall (Rec.), overall accuracy (OA), and intersection over union (IoU) to evaluate the performance of our model relative to SOTA methods. These metrics were obtained by comparing the ground truth with the predicted map.
[0059] Table I Experimental results on the LEVIR-CD and WHU-CD datasets, with the first place in bold
[0060]
[0061] As shown in Table 1, the present invention exhibits excellent performance. On the LEVIR-CD dataset, the F1 score of the present invention reaches 91.96, the precision is 93.27, the recall rate is 90.68, the overall accuracy (OA) is as high as 99.46, and the intersection over union (IoU) is 85.11, ranking first among the comparative models. These results indicate that the present invention has obvious advantages in the accuracy, recall rate, and overall classification performance of change detection. In particular, the present invention performs excellently in terms of precision and IoU, being able to identify the changed areas with high precision and having a high degree of coincidence between the recognition results and the actual changed areas. On the WHU-CD dataset, the present invention also performs outstandingly, with an F1 score of 92.06, a precision of 94.02, a recall rate of 90.18, an OA of 99.45, and an IoU of 85.29. These results further confirm the leading position of the present invention in the change detection task. The high accuracy and IoU values are particularly noteworthy because they not only reflect the precision of the model in identifying the changed areas but also indicate a high degree of consistency between its predictions and the actual changed areas. Considering the performance on the two datasets, the advantages of the present invention in the field of change detection are obvious. Its excellent performance in key indicators not only proves the effectiveness and reliability of the present invention in the change detection task but also supports its potential in practical applications. These advantages make the present invention a strong candidate for change detection research and practical applications.
[0062] As Figure 3 shown, the present invention performs particularly well in change detection. On the LEVIR-CD dataset, the present invention can accurately identify the changed areas while maintaining a low number of false positive and false negative examples. On the WHU-CD dataset, the present invention also shows a high accuracy, with the detection results highly consistent with the ground truth (GT), and the number of false positive and false negative cases being the least among all models. This indicates that the present invention has obvious advantages in reducing false positives and false negatives. Compared with other models (such as FC-EF, FC-Siam-conc, FC-Siam-diff, STANet, SNUNet, MSPSNet, HANet, BIT, ChangeFormer, RSP-BIT, etc.), the present invention shows obvious advantages in both the accuracy and robustness of change detection. This further confirms the excellent performance of the present invention in quantitative comparison; on the LEVIR-CD and WHU-CD datasets, the present invention has achieved the highest scores in key indicators such as F1 score, precision, recall rate, overall accuracy, and intersection over union. These visualization results are consistent with the previous quantitative analysis, further highlighting the advantages of the present invention in change detection.
[0063] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "upper" etc. may be used herein to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that the spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, a device described as "above" or "over" other devices or structures will then be positioned "below" or "under" the other devices or structures. Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device may also be positioned in other different ways (rotated 90 degrees or in other orientations), and the corresponding explanations for the spatial relative descriptions used herein will be made accordingly.
[0064] In addition, it should be noted that the use of terms such as "first", "second" etc. to define components is only for the convenience of differentiating the corresponding components. Without additional statements, the above terms have no special meanings, and thus should not be construed as limiting the protection scope of the present invention.
[0065] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A remote sensing target change detection method based on a hybrid self-cross attention mechanism, characterized in that, The remote sensing target change detection method based on the hybrid self-cross attention mechanism includes: Step 1: Input the first-phase remote sensing image T1 and the second-phase remote sensing image T2, and use the first convolutional structure F1, the second convolutional structure F2, the third convolutional structure F3, and the fourth convolutional structure F4 to extract features from the first and second-phase remote sensing images T1 and the second-phase remote sensing image T2 respectively; Step 2: The features of the first-phase remote sensing image T1 and the features of the second-phase remote sensing image T2 extracted in Step 1 enter the first hybrid self-cross attention fusion module HSC-AFM at the same time. The first hybrid self-cross attention fusion module HSC-AFM includes a self-attention module, a cross-attention module, and a fusion module. First, use the self-attention module to extract the global context information in the single-phase image and capture the correlation between different regions in the image; Secondly, input the output features of the self-attention module and the output feature map of the fourth convolutional structure F4 into the cross-attention module together, and use the cross-attention module to focus on the time change features and learn the spatio-temporal interaction information between images at different times; Finally, input the output features of the cross-attention module and the output feature map of the third convolutional structure F3 into the fusion module together. The fusion module integrates the two attention mechanisms of the self-attention module and the cross-attention module to achieve efficient fusion of spatio-temporal information; Step 3: Input the result extracted in Step 2 and the output feature map of the second convolutional structure F2 into the second hybrid self-cross attention fusion module HSC-AFM together to strengthen the features of the change information; Step 4: Input the result extracted in Step 3 and the output feature map of the first convolutional structure F1 into the third hybrid self-cross attention fusion module HSC-AFM together. The third hybrid self-cross attention fusion module HSC-AFM perceives the differences between multi-scale features, strengthens the key features, and suppresses redundant information at the same time; Step 5: Output the features in Step 4 to the classifier, and after passing through the classifier, output a binary prediction map to complete the remote sensing target change detection task.
2. The remote sensing target change detection method based on the hybrid self-cross attention mechanism according to claim 1, wherein The self-attention module and the cross-attention module constitute the hybrid self-cross attention module. A part of the output feature map of the fourth convolutional structure F4 is input to the self-attention module, and another part of the output feature map of the fourth convolutional structure F4 is input to the cross-attention module.
3. The remote sensing target change detection method based on the hybrid self-cross attention mechanism according to claim 2, wherein The self-attention module is used to perform convolution operations, self-attention calculations, and linear transformations; The convolution operation specifically includes: performing a convolution operation on a part of the output feature map of the fourth convolutional structure F4 to obtain feature maps Q1, K1, V1. The sizes of feature maps Q1 and V1 are both (h×w)×c, and the size of feature map K1 is c×(h×w); The self-attention calculation specifically includes: calculating the dot product of the obtained feature maps Q1 and K1 to obtain an attention weight matrix X1 with a size of c×(h×w); performing a Softmax operation on the attention weight matrix X1 to obtain a normalized attention weight; multiplying the normalized attention weight by the feature map V1 to obtain a feature map B with a size of (h×w)×c; The linear transformation specifically includes: performing a linear transformation on the feature map B to obtain a feature map Q2 with a size of (h×w)×c, and the linear transformation integrates features and provides decision-making information for the change detection task.
4. The remote sensing target change detection method based on the hybrid self-cross attention mechanism according to claim 3, characterized in that The cross-attention mechanism calculation of the cross-attention module specifically includes: performing a linear operation on another part of the output feature map of the fourth convolutional structure F4 to obtain feature maps K2 and V2, calculating the dot product of the feature map Q2 and K2 to obtain an attention weight matrix X2 with a size of c×(h×w); performing a softmax operation on the attention weight matrix X2 to obtain a normalized attention weight; multiplying the normalized attention weight by the feature map V2 to obtain a feature map M with a size of h×w×C.
5. The remote sensing target change detection method based on the hybrid self-cross attention mechanism according to claim 4, wherein, The fusion module includes an upsampling unit and a convolutional block. The upsampling unit is used to perform an upsampling operation on the feature map M to obtain a feature map with a size of 2h×2w×C, and the convolutional block is used to perform a convolutional operation on the upsampled feature map and the output feature map of the third convolutional structure F3 to obtain a final output feature map O with a size of 2h×2w×C.
6. The remote sensing target change detection method based on the hybrid self-cross attention mechanism according to any one of claims 1 to 5, characterized in that The remote sensing target change detection method based on the hybrid self-cross attention mechanism further includes: verifying the model effect using a loss function.
7. The remote sensing target change detection method based on the hybrid self-cross attention mechanism according to claim 6, wherein The loss function uses the Dice loss function, and the Dice loss function is where y i is the predicted value of the true label, is the predicted value of pixel i, and N represents the total number of pixels.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the remote sensing target change detection method based on the hybrid self-cross attention mechanism as described in claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the remote sensing target change detection method based on the hybrid self-cross attention mechanism as described in any one of claims 1 to 7.