A method, device, equipment and medium for change detection of remote sensing images

By using feature extraction and space-time difference enhancement modules in remote sensing image change detection to generate multi-scale feature maps and perform edge refinement and feature fusion, the problem of insufficient accuracy in the existing methods is solved, and the accuracy of remote sensing image change detection is achieved.

CN119274071BActive Publication Date: 2025-07-22NANCHANG SURVEYING & MAPPING RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411333044.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-07-22
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

The existing remote sensing image change detection methods have shortcomings in terms of accuracy, especially ignoring the substantial differences between bi-time phase images and the correlation between different scale features, resulting in missed and missed detection problems.

Method used

By acquiring the front and back phase remote sensing images, the feature extraction module and the space-time difference enhancement module in the encoder generate multiple space-time difference enhancement feature maps of different scales, capture the remote context information and local cross-feature information of the changing area, and perform edge refinement and fuse with features through the decoder to generate predicted binary images.

Benefits of technology

Learning global change information and local fine-grained change information between the two-time phase feature maps at the same scale enhances the accuracy of remote sensing image change detection and improves the accuracy and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274071B_ABST
    Figure CN119274071B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and medium for change detection of remote sensing images, relating to the technical field of image processing. The method for change detection of remote sensing images includes: acquiring a pre-temporal remote sensing image and a post-temporal remote sensing image, inputting the pre-temporal remote sensing image and the post-temporal remote sensing image into an encoder to obtain multiple spatio-temporal difference enhancement feature maps of different scales, wherein the encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules, inputting the multiple spatio-temporal difference enhancement feature maps of different scales into a decoder, and performing edge refinement and feature fusion on the multiple spatio-temporal difference enhancement feature maps of different scales through the decoder to obtain a predicted binary image including change information between the pre-temporal remote sensing image and the post-temporal remote sensing image. The embodiments of the present application can improve the accuracy of change detection of remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, device, equipment and medium for change detection of remote sensing images. Background Art

[0002] Remote sensing image change detection is a technical means to analyze remotely sensed images acquired at different times to identify and quantify changes in land cover and land use in the same area, and it is currently widely used in many fields such as natural disaster and environmental monitoring, urban planning, agricultural management, and resource investigation.

[0003] The research objects of remote sensing image change detection methods in related technologies include detection methods based on pixels, objects, scenes, etc. There are still some problems with these detection methods, resulting in low accuracy for remote sensing image change detection. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to provide a method, device, equipment and medium for change detection of remote sensing images, aiming to improve the accuracy of remote sensing image change detection.

[0005] To achieve the above object, on the one hand, an embodiment of this application provides a method for change detection of remote sensing images, including the following steps: obtaining a pre-temporal remote sensing image and a post-temporal remote sensing image, where the pre-temporal remote sensing image is used to represent the surface feature image of the target area at the first time point, and the post-temporal remote sensing image is used to represent the surface feature image of the target area at the second time point; inputting the pre-temporal remote sensing image and the post-temporal remote sensing image into an encoder to obtain multiple spatio-temporal difference enhanced feature maps of different scales, where the encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules, the feature extraction module is used to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple different-scale bi-temporal feature maps, and the multiple spatio-temporal difference enhancement modules are used to capture the remote context information and local cross-feature information of the change areas in the multiple different-scale bi-temporal feature maps to generate the multiple different-scale spatio-temporal difference enhanced feature maps; inputting the multiple different-scale spatio-temporal difference enhanced feature maps into a decoder, and through the decoder, performing edge refinement and feature fusion on the multiple different-scale spatio-temporal difference enhanced feature maps to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

[0006] In some embodiments, the feature extraction module includes two parallel feature extraction branches. Inputting the pre-temporal remote sensing image and the post-temporal remote sensing image into the encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales, including: Inputting the pre-temporal remote sensing image and the post-temporal remote sensing image into two parallel feature extraction branches respectively, and extracting multiple different-scale dual-temporal feature maps through the feature extraction network in the feature extraction branches; Inputting the multiple different-scale dual-temporal feature maps into the corresponding spatio-temporal difference enhancement module respectively to obtain the multiple different-scale spatio-temporal difference enhanced feature maps.

[0007] In some embodiments, the spatio-temporal difference enhancement module includes a subtraction branch and a channel alternating splicing branch. Inputting the dual-temporal feature map into the spatio-temporal difference enhancement module to obtain the spatio-temporal difference enhanced feature map, including: Inputting the dual-temporal feature map into the subtraction branch, and calculating the absolute difference feature between the dual-temporal feature maps in the subtraction branch; Inputting the absolute difference feature into the dual attention mechanism module, and capturing the remote context information of the changed region in the dual-temporal feature map through the dual attention mechanism module to obtain a spatial attention enhanced feature map and a channel attention enhanced feature map; Splicing the spatial attention enhanced feature map and the channel attention enhanced feature map in channels to obtain a spatial-channel attention enhanced feature map; Inputting the dual-temporal feature map into the channel alternating splicing branch, inputting the feature map after channel alternating splicing into the channel adaptive enhancement module to enhance the model's ability to extract cross-feature interaction information, and obtaining a channel enhanced feature map; Adding the spatial-channel attention enhanced feature map and the channel enhanced feature map to obtain the spatio-temporal difference enhanced feature map.

[0008] In some embodiments, the dual attention mechanism module includes a spatial attention mechanism module and a channel attention mechanism module. Inputting the absolute difference feature into the dual attention mechanism module, and capturing the remote context information of the changed region in the dual-temporal feature map through the dual attention mechanism module to obtain a spatial attention enhanced feature map and a channel attention enhanced feature map, including: Inputting the absolute difference feature into the spatial attention mechanism module, and establishing the context relationship of local features through the spatial attention mechanism module to obtain the spatial attention enhanced feature map; Inputting the absolute difference feature into the channel attention mechanism module, and establishing the remote dependence relationship between the channels of the feature map through the channel attention mechanism module to obtain the channel attention enhanced feature map.

[0009] In some embodiments, the spatio-temporal difference enhancement module includes a channel alternating splicing branch. Inputting the dual-temporal feature map into the spatio-temporal difference enhancement module to obtain multiple spatio-temporal difference enhancement feature maps of different scales includes: inputting the dual-temporal feature map into the channel alternating splicing branch, and performing alternating splicing on the dual-temporal feature map in the channel alternating splicing branch to obtain a spliced feature; inputting the spliced feature into a channel adaptive enhancement mechanism module, and enhancing the local cross-feature interaction information of the changing region in the dual-temporal feature map through the channel adaptive enhancement mechanism module to obtain a channel enhanced feature map; adding the channel enhanced feature map and a spatial channel attention enhanced feature map to obtain the spatio-temporal difference enhancement feature map.

[0010] In some embodiments, inputting the multiple spatio-temporal difference enhancement feature maps of different scales into a decoder, and performing edge refinement and feature fusion on the multiple spatio-temporal difference enhancement feature maps of different scales through the decoder includes: respectively inputting the multiple spatio-temporal difference enhancement feature maps of different scales into multiple corresponding edge refinement residual modules, and enhancing the edges of the changing regions in the spatio-temporal difference enhancement feature maps of different scales through the edge refinement residual modules to obtain multiple edge refinement feature maps; inputting the multiple edge refinement feature maps into an adaptive bidirectional feature fusion module, and performing layer-by-layer fusion from top to bottom and from bottom to top on the edge refinement feature maps through the adaptive bidirectional feature fusion module to obtain a fused feature map; performing upsampling and channel adjustment on the fused feature map to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

[0011] In some embodiments, the edge refinement residual module includes a first branch and a second branch. Inputting the multiple spatio-temporal difference enhancement feature maps of different scales into multiple corresponding edge refinement residual modules, and enhancing the edges of the changing regions in the spatio-temporal difference enhancement feature maps of different scales through the edge refinement residual modules to obtain multiple edge refinement feature maps includes: inputting the multiple spatio-temporal difference enhancement feature maps of different scales into the corresponding first branch, and unifying the number of channels of the spatio-temporal difference enhancement feature maps of different scales through the first branch to obtain multiple standard channel feature maps; inputting the multiple standard channel feature maps into the corresponding second branch, and refining the spatio-temporal difference enhancement feature maps of different scales through the second branch to obtain multiple refined feature maps; respectively summing the multiple standard channel feature maps and the refined feature maps to obtain multiple edge refinement feature maps.

[0012] To achieve the above object, on the other hand, an embodiment of the present application provides a change detection device for remote sensing images. The change detection device for remote sensing images includes: an acquisition module configured to acquire a pre-temporal remote sensing image and a post-temporal remote sensing image, where the pre-temporal remote sensing image is used to represent the surface feature image of a target area at a first time point, and the post-temporal remote sensing image is used to represent the surface feature image of the target area at a second time point; a feature extraction and generation module configured to input the pre-temporal remote sensing image and the post-temporal remote sensing image into an encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales. The encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The feature extraction module is configured to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple dual-temporal feature maps at different scales, and the multiple spatio-temporal difference enhancement modules are configured to capture remote context information and local cross-feature information of the changed areas in the multiple dual-temporal feature maps at different scales to generate the multiple spatio-temporal difference enhanced feature maps at different scales; a feature refinement and fusion module configured to input the multiple spatio-temporal difference enhanced feature maps at different scales into a decoder, and perform edge refinement and feature fusion on the multiple spatio-temporal difference enhanced feature maps at different scales through the decoder to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

[0013] To achieve the above object, on yet another aspect, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned change detection method for remote sensing images is implemented.

[0014] To achieve the above object, on yet another aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the above computer program is executed by one or more processors, the steps of the above-mentioned change detection method for remote sensing images can be implemented.

[0015] The embodiments of the present application at least include the following beneficial effects:

[0016] The present application provides a method, apparatus, device, and medium for change detection of remote sensing images. In the embodiments of the present application, first, a pre-temporal remote sensing image and a post-temporal remote sensing image are obtained. The pre-temporal remote sensing image is used to represent the surface feature image of the target area at the first time point, and the post-temporal remote sensing image is used to represent the surface feature image of the target area at the second time point. Then, the pre-temporal remote sensing image and the post-temporal remote sensing image are input into an encoder to obtain multiple spatio-temporal difference enhancement feature maps at different scales. Among them, the encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The feature extraction module is used to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple different-scale bi-temporal feature maps. The multiple spatio-temporal difference enhancement modules are used to capture the remote context information and local cross-feature information of the changed areas in the multiple different-scale bi-temporal feature maps to generate multiple different-scale spatio-temporal difference enhancement feature maps. Finally, the multiple different-scale spatio-temporal difference enhancement feature maps are input into a decoder, and the decoder performs edge refinement and feature fusion on the multiple different-scale spatio-temporal difference enhancement feature maps to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image. In the embodiments of the present application, the feature extraction module extracts features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple different-scale bi-temporal feature maps. Then, the multiple spatio-temporal difference enhancement modules capture the remote context information and local cross-feature information of the changed areas in the multiple different-scale bi-temporal feature maps to generate multiple different-scale spatio-temporal difference enhancement feature maps, which can learn the global change information and local fine-grained change information between the bi-temporal feature maps at the same scale, enhance the spatio-temporal difference of the bi-temporal image features extracted by the encoder, and through the decoder, the rich position information in the shallow feature maps and the rich semantic information in the deep feature maps are effectively fused, thereby improving the accuracy of remote sensing image change detection.

[0017] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and / or additional aspects and advantages of the present application will become apparent and understandable from the description of the embodiments in conjunction with the following drawings, where:

[0019] Figure 1 is a flowchart of a method for change detection of remote sensing images provided by some embodiments of the present application;

[0020] Figure 2 is a schematic diagram of the architecture of a change detection model for remote sensing images provided by some embodiments of the present application;

[0021] Figure 3 It is a schematic diagram of the network structure of the encoder provided by some embodiments of the present application;

[0022] Figure 4 It is a schematic diagram of the network structure of the feature edge refinement residual module in the decoder part provided by some embodiments of the present application;

[0023] Figure 5 It is a schematic diagram of the network structure of the adaptive bidirectional feature fusion module in the decoder part provided by some embodiments of the present application;

[0024] Figure 6 It is a schematic diagram of the comparison of the visualization results of the remote sensing image change detection method provided by some embodiments of the present application and the existing remote sensing image change detection method on the WHU-CD dataset;

[0025] Figure 7 It is a schematic diagram of the comparison of the visualization results of the remote sensing image change detection method provided by some embodiments of the present application and the existing remote sensing image change detection method on the LEVIR-CD dataset;

[0026] Figure 8 It is a schematic diagram of the comparison of the visualization results of the remote sensing image change detection method provided by some embodiments of the present application and the existing remote sensing image change detection method on the SYSU-CD dataset;

[0027] Figure 9(a) is a schematic diagram of the comparison result of the computational complexity of the remote sensing image change detection method provided by some embodiments of the present application and the existing remote sensing image change detection method on the WHU-CD dataset;

[0028] Figure 9(b) is a schematic diagram of the comparison result of the computational complexity of the remote sensing image change detection method provided by some embodiments of the present application and the existing remote sensing image change detection method on the LEVIR-CD dataset;

[0029] Figure 9(c) is a schematic diagram of the comparison result of the computational complexity of the remote sensing image change detection method provided by some embodiments of the present application and the existing remote sensing image change detection method on the SYSU-CD dataset;

[0030] Figure 10 It is a schematic block diagram of the modules of the remote sensing image change detection device provided by some embodiments of the present application;

[0031] Figure 11 It is a schematic diagram of the hardware structure of the electronic device provided by some embodiments of the present application. Detailed implementation manners

[0032] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the mention of "embodiment" in this document means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase does not necessarily refer to the same embodiment at every occurrence in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0033] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0034] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present application, "at least one" includes one, two or more than two, "a plurality of" includes two or more than two, "each" refers to each of the corresponding plurality, and "any one" refers to any one of the plurality.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0036] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0037] To make the inventive concept of this application easy to understand, before the detailed description of the embodiments of this application, the English abbreviations (terms) / related concepts involved in the embodiments of this application are first described, and the English abbreviations (terms) / related concepts involved in the embodiments of this application are applicable to the following explanations.

[0038] Dual-temporal feature map: It involves comparing images acquired at two different time points to identify and analyze the changes that occur during this time interval. Phase one represents the image acquired at the first time point; phase two represents the image acquired at the second time point.

[0039] LEVIR-CD dataset: It is a large-scale remote sensing building change detection dataset. This dataset is designed as a new benchmark for evaluating change detection (CD) algorithms, especially deep learning-based algorithms.

[0040] WHU-CD dataset: It is a dataset for remote sensing image change detection, focusing on building change detection, and is usually used to evaluate and compare different remote sensing image change detection algorithms.

[0041] SYSU-CD dataset: It is a dataset for remote sensing image change detection, used to detect and analyze changes on the earth's surface, such as urban expansion, deforestation, or the impact of natural disasters.

[0042] Remote sensing image change detection is a technical means to identify and quantify the changes in surface cover and land use in the same area by analyzing remote sensing images acquired at different times. Currently, it has been widely used in many fields such as natural disaster and environmental monitoring, urban planning, agricultural management, and resource investigation. With the rapid development of remote sensing earth observation technology, image processing technology, and artificial intelligence, change detection driven by the joint of multi-source, multi-temporal remote sensing image data - model - knowledge has become an important direction in the remote sensing discipline. There are problems such as inconsistent ground object representations (such as pseudo-changes like shadow changes and seasonal changes), complex scene changes, and "same object with different spectra, different objects with the same spectra" in multi-source, multi-temporal, multi-spectral, and high-resolution remote sensing images, resulting in unsatisfactory detection effects such as missed detection and false detection in remote sensing image change detection, which brings many challenges to the change detection task.

[0043] The research objects on which traditional remote sensing image change detection methods are based include methods based on pixels, objects, scenes, etc.

[0044] Pixel-based remote sensing change detection methods mainly obtain remote sensing images at two time points, perform image registration on them to ensure spatial alignment, then use thresholding, statistical, time series analysis, or machine learning methods to perform change detection on each pixel, and finally perform post-processing optimization and analyze the results. However, this method is easily affected by noise, the detection of small-area changes in the image is not accurate enough, and it is unable to quantify the degree of relationship between pixels.

[0045] Object-based remote sensing change detection methods organize the pixels in the image into objects, which can better reflect the scenes in the real world. Its detection accuracy is higher than that of pixel-based methods. However, due to factors such as registration errors between multi-temporal images, complex landform shapes, and unclear boundaries in object-based detection methods, the change regions extracted by object-based change detection methods are inaccurate.

[0046] Scene-based remote sensing change detection methods use land cover class information and scene context to segment and classify remote sensing images, and finally, based on change analysis techniques for land cover classes, such as change detection models or classifiers, identify and analyze the changes in different land cover classes over time. However, the definition and extraction of scenes for this method are relatively complex, various factors need to be considered comprehensively, it is not sensitive enough to detect local subtle changes, and situations such as missed detection or false detection are likely to occur.

[0047] With the continuous development of artificial intelligence technology, deep learning models have been widely used in high-resolution remote sensing image interpretation tasks such as scene classification, object detection, and semantic segmentation, and their interpretation accuracy has exceeded that of traditional methods. In remote sensing image change detection tasks, the mainstream networks used include convolutional neural networks (CNNs), Transformers, etc. Although significant progress has been made in the field of remote sensing image change detection using deep learning, existing change detection methods still have certain limitations, including: (1) mainly focusing on indicators such as the optimal IoU or F1 score of the detection results, while ignoring the applicability between different lightweight and non-lightweight convolutional network models, Transformer models, and network architectures under the condition of the same model calculation efficiency; (2) overly focusing on feature extraction and not fully utilizing spatio-temporal difference information, ignoring the substantial differences between bi-temporal images; (3) multi-scale feature fusion relying too much on modules with specific structures and not fully quantifying the importance of different-scale features, ignoring the mutual correlation between different-scale features.

[0048] In view of this, the present application proposes a method, apparatus, device and medium for change detection of remote sensing images. The solution obtains a pre-temporal remote sensing image and a post-temporal remote sensing image. The pre-temporal remote sensing image is used to represent the surface feature image of the target area at the first time point, and the post-temporal remote sensing image is used to represent the surface feature image of the target area at the second time point. Then, the pre-temporal remote sensing image and the post-temporal remote sensing image are input into an encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales. Among them, the encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The feature extraction module is used to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple bi-temporal feature maps at different scales. The multiple spatio-temporal difference enhancement modules are used to capture the remote context information and local cross-feature information of the changed areas in the multiple bi-temporal feature maps at different scales to generate multiple spatio-temporal difference enhanced feature maps at different scales. Finally, the multiple spatio-temporal difference enhanced feature maps at different scales are input into a decoder. The decoder refines the edges and fuses the features of the multiple spatio-temporal difference enhanced feature maps at different scales to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image. In the embodiment of the present application, the feature extraction module extracts features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple bi-temporal feature maps at different scales. Then, the multiple spatio-temporal difference enhancement modules capture the remote context information and local cross-feature information of the changed areas in the multiple bi-temporal feature maps at different scales to generate multiple spatio-temporal difference enhanced feature maps at different scales, which can learn the global change information and local fine-grained change information between the bi-temporal feature maps at the same scale, enhance the spatio-temporal difference of the bi-temporal image features extracted by the encoder, and effectively fuse the rich position information in the shallow feature maps and the rich semantic information in the deep feature maps through the decoder to refine the edges and fuse the features, thereby improving the accuracy of remote sensing image change detection.

[0049] The method provided by the embodiment of the present application can be applied to the electronic device provided by the embodiment of the present application. Among them, the electronic device can be a terminal or a server.

[0050] The terminal can be a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto.

[0051] The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.

[0052] The implementation steps of a change detection method for remote sensing images provided by an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0053] Please refer to Figure 1 , Figure 1 which is a flowchart of the change detection method for remote sensing images provided by some embodiments of the present application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0054] The method of the embodiment of the present application includes the following steps:

[0055] Step 101: Obtain a pre-temporal remote sensing image and a post-temporal remote sensing image. The pre-temporal remote sensing image is used to represent the surface feature image of the target area at the first time point, and the post-temporal remote sensing image is used to represent the surface feature image of the target area at the second time point;

[0056] Step 102: Input the pre-temporal remote sensing image and the post-temporal remote sensing image into an encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales. Among them, the encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The feature extraction module is used to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple different-scale bi-temporal feature maps, and the multiple spatio-temporal difference enhancement modules are used to capture the remote context information and local cross-feature information of the changed areas in the multiple different-scale bi-temporal feature maps to generate multiple different-scale spatio-temporal difference enhanced feature maps;

[0057] Step 103: Input the multiple spatio-temporal difference enhanced feature maps at different scales into a decoder, and the decoder performs edge refinement and feature fusion on the multiple spatio-temporal difference enhanced feature maps at different scales to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

[0058] Steps 101 to 103 shown in the embodiments of the present application extract features from the pre-phase remote sensing image and the post-phase remote sensing image through a feature extraction module to generate multi-scale dual-phase feature maps. Then, through multiple spatio-temporal difference enhancement modules, the remote context information and local cross-feature information of the changed regions in the multi-scale dual-phase feature maps are captured to generate multi-scale spatio-temporal difference enhancement feature maps, which can learn the global change information and local fine-grained change information between the dual-phase feature maps at the same scale, enhance the spatio-temporal difference of the dual-phase image features extracted by the encoder, and through the decoder, edge refinement and feature fusion are performed on the multi-scale spatio-temporal difference enhancement feature maps, so that the rich position information in the shallow feature maps and the rich semantic information in the deep feature maps are effectively fused, thereby improving the accuracy of remote sensing image change detection.

[0059] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the architecture of a change detection model for remote sensing images provided by some embodiments of the present application. Figure 2 In, the change detection model for remote sensing images mainly includes an encoder module and a decoder module. The part enclosed by the dotted line box on the left is the encoder part, and the part enclosed by the dotted line box on the right is the decoder part.

[0060] The encoder part includes a feature extraction module and a spatio-temporal difference enhancement module (STDEM). T1 and T2 are the pre-phase remote sensing image and the post-phase remote sensing image obtained respectively. By separately extracting features from T1 and T2 through parallel feature extraction branches (feature extraction module), 4 groups of feature maps at different scales can be obtained, for example: the first group of feature maps, the second group of feature maps, the third group of feature maps, and the fourth group of feature maps. Each group of feature maps includes two feature maps, corresponding to the dual-phase feature maps. The 4 groups of feature maps at different scales are respectively input into 4 STDEM modules to obtain spatio-temporal difference enhancement feature maps at 4 different scales.

[0061] The decoder part includes an edge refinement residual module (ERRM) and an adaptive bidirectional feature fusion module (ABiFFM). The spatio-temporal difference enhancement feature maps at 4 different scales are respectively input into 4 corresponding ERRM to obtain 4 edge-refined feature maps at different scales. Then, the 4 edge-refined feature maps at different scales are input into ABiFFM to obtain 4 fusion feature maps at different scales. Finally, bilinear interpolation is used for 4-fold upsampling of the feature maps and 1×1 convolution to adjust the number of channels of the feature maps to 2, generating the final binary prediction map. Among them, Figure 2 the green planar rectangle in represents upsampling, and the red planar rectangle represents 1×1 convolution. Figure 2 In, there is also a label image participating in the supervised training.

[0062] The specific implementation methods of the above steps are introduced below.

[0063] In step 101, a pre-temporal remote sensing image and a post-temporal remote sensing image are obtained. The pre-temporal remote sensing image is used to represent the surface feature image of the target area at the first time point, and the post-temporal remote sensing image is used to represent the surface feature image of the target area at the second time point.

[0064] The surface feature image of the target area at the first time point can be the surface feature image of the target area before the change occurs, and the surface feature image of the target area at the second time point can be the surface feature image of the target area after the change occurs.

[0065] Among them, the remote sensing image can be the remote sensing image data provided by a space agency or organization, or the high-resolution remote sensing image provided by a commercial satellite company, or the remote sensing image can be obtained through an online platform, a data sharing community or forum. The present application does not limit the acquisition method of the remote sensing image.

[0066] The embodiment provided by the present application provides image support for subsequent processing of remote sensing images and obtaining change information between remote sensing images by obtaining a pre-temporal remote sensing image and a post-temporal remote sensing image.

[0067] In step 102, the pre-temporal remote sensing image and the post-temporal remote sensing image are input into an encoder to obtain multiple spatio-temporal difference enhanced feature maps of different scales. Among them, the encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The feature extraction module is used to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple different-scale two-temporal feature maps, and the multiple spatio-temporal difference enhancement modules are used to capture the remote context information and local cross-feature information of the change regions in the multiple different-scale two-temporal feature maps to generate multiple different-scale spatio-temporal difference enhanced feature maps.

[0068] In some implementation manners, the feature extraction module includes two parallel feature extraction branches. The pre-temporal remote sensing image and the post-temporal remote sensing image can be respectively input into the two parallel feature extraction branches, and multiple different-scale two-temporal feature maps are extracted through the feature extraction networks in the feature extraction branches; the multiple different-scale two-temporal feature maps are respectively input into the corresponding spatio-temporal difference enhancement modules to obtain multiple different-scale spatio-temporal difference enhanced feature maps.

[0069] Since in diverse change detection scenarios, feature extraction networks of different types and different parameter amounts will have different impacts on the performance and effect of the model in a specific scenario, it is necessary to analyze the adaptability of multiple different types and different parameter amounts of feature extraction networks to the network architecture, and then select a suitable backbone feature extraction network.

[0070] Exemplarily, the selected feature backbone extraction network can be an end-to-end feature extraction network built through the MobileNet V3-Large architecture. Among them, the MobileNet V3-Large architecture is a variant of the lightweight neural network architecture series MobileNet.

[0071] Exemplarily, according to the input remote sensing image group, on two parallel feature extraction branches, a feature extraction backbone network adapted to the network architecture can be used to extract four different-scale dual-temporal feature maps, including the first feature map group, the second feature map group, the third feature map group, and the fourth feature map group. Subsequently, the dual-temporal feature maps at different scales are respectively input into the spatio-temporal difference enhancement module to obtain multiple different-scale spatio-temporal difference enhancement feature maps, including the fifth feature map, the sixth feature map, the seventh feature map, and the eighth feature map.

[0072] The main purpose of selecting a feature extraction backbone network adapted to the network architecture to obtain the fifth, sixth, seventh, and eighth feature maps is that the generalization capabilities of different models are different, and feature extraction backbone networks of different types and numbers of parameters have different impacts on the performance and effects of the model in specific scenarios. Therefore, a lightweight network of the MobileNet V3-Large series can be selected to reduce the number of parameters, improve the feature extraction ability and computer efficiency. At the same time, the spatio-temporal difference enhancement module is used to process these dual-temporal feature maps respectively, so as to enhance the substantial difference of the dual-temporal features.

[0073] The main purpose of obtaining spatio-temporal difference enhancement feature maps at multiple different scales is that the channel adaptive enhancement mechanism and the dual attention mechanism can capture the remote context information and local cross-feature information of the changing regions in the dual-temporal feature maps, and enhance the spatio-temporal difference of the dual-temporal image features.

[0074] In some embodiments, the spatio-temporal difference enhancement module includes a subtraction branch and a channel alternating splicing branch. The dual-temporal feature map can be input into the subtraction branch, and the absolute difference feature between the dual-temporal feature maps is calculated in the subtraction branch; the absolute difference feature is input into the dual attention mechanism module, and the remote context information of the changing regions in the dual-temporal feature maps is captured through the dual attention mechanism module to obtain a spatial attention enhancement feature map and a channel attention enhancement feature map; the spatial attention enhancement feature map and the channel attention enhancement feature map are subjected to channel splicing to obtain a spatial-channel attention enhancement feature map; the dual-temporal feature map is input into the channel alternating splicing branch, and the feature map after channel alternating splicing is input into the channel adaptive enhancement module to enhance the model's ability to extract cross-feature interaction information to obtain a channel enhancement feature map; finally, the spatial-channel attention enhancement feature map and the channel enhancement feature map are added together to obtain a spatio-temporal difference enhancement feature map.

[0075] Generally, there are problems such as inconsistent ground object representations, the same object with different spectra, and different objects with the same spectra in dual-temporal remote sensing images. Directly subtracting, adding, or channel splicing the dual-temporal images will introduce a large amount of noise into change detection, and the detection results lack interpretability, leading to problems such as missed detection and false detection. Therefore, a spatio-temporal difference enhancement module with a dual-branch structure is introduced, which uses a channel adaptive enhancement mechanism and a dual attention mechanism to capture the remote context information and local cross-feature information of the changed regions in the dual-temporal feature maps, enhancing the spatio-temporal difference of the dual-temporal image features.

[0076] In the subtraction branch, a dual attention mechanism of a spatial attention mechanism and a channel attention mechanism is adopted. The spatial attention mechanism can model rich context relationships on local features, and the channel attention mechanism can explicitly model the long-range dependence relationships between the channels of the feature maps.

[0077] In some embodiments, the dual attention mechanism module includes a spatial attention mechanism module and a channel attention mechanism module. The absolute difference feature can be input into the spatial attention mechanism module to establish the context relationship of local features through the spatial attention mechanism module, obtaining a spatially attention-enhanced feature map; the absolute difference feature is input into the channel attention mechanism module to establish the long-range dependence relationship between the channels of the feature map through the channel attention mechanism module, obtaining a channel attention-enhanced feature map.

[0078] Optionally, for the subtraction branch, first calculate the absolute difference feature between the input first image and the second image group (dual-temporal feature map), and then input it into the spatial attention mechanism module and the channel attention mechanism module to capture the remote context information of the changed region, obtaining a spatially attention-enhanced feature map and a channel attention-enhanced feature map. Then, the two are spliced ​​on the channel, and after 3x3 convolution operation, normalization, and non-linear activation of the ReLu function, the third image (spatially and channel attention-enhanced feature map) is obtained.

[0079] Exemplarily, reference can be made to Figure 3 the spatial attention mechanism module in. The spatial attention mechanism can model rich context relationships on local features. For a given feature map A, A ∈ N H×W×C , where H represents the height, W represents the width, and C represents the number of channels. Use 1×1 convolution to generate feature maps B, C, and D of the same size respectively, and reset the sizes of B and C, D to N C×(H×W) , N (H×W)×C , and then perform matrix multiplication calculation on B and C, and calculate the spatial feature map K using the Softmax function, K ∈ N (H×W)×(H×W) :

[0080]

[0081] Kji Denote the elements in the spatial feature map K, where i and j are variables representing the elements.

[0082] Then perform matrix multiplication on K and D, and reset the calculation result to N H×W×C , multiply the reset result by the learnable scale parameter α, then perform a summation calculation with the original feature map A, and the final output is E, where E ∈ N H×W×C :

[0083]

[0084] The output E is the spatially attention-enhanced feature map.

[0085] Exemplarily, refer to Figure 3 the channel attention mechanism module in H×W×C , first reset the size of A to N C×(H×W) and N (H×W)×C , perform matrix multiplication on the two, and calculate the channel feature map T using the Softmax function, where T ∈ N C×C :

[0086]

[0087] t ji Denote the elements in the channel feature map T, where i and j are variables representing the elements.

[0088] Then perform matrix multiplication on T and the reset matrix of A, and reset the calculation result to N H×W×C , multiply the reset result by the learnable scale parameter β, then perform a summation calculation with the feature map A, and the final output is R, where R ∈ N H×W×C :

[0089]

[0090] The output R is the channel attention-enhanced feature map.

[0091] Concatenate the spatially attention-enhanced feature map and the channel attention-enhanced feature map on the channel dimension, then perform 3x3 convolution operation, normalization, and non-linear activation with the ReLu function to obtain the third image (spatially and channel attention-enhanced feature map).

[0092] In some embodiments of the present application, the spatio-temporal difference enhancement module includes a channel alternating splicing branch. The dual-temporal feature map can be input into the channel alternating splicing branch, where the dual-temporal feature map is alternately spliced to obtain a spliced feature. The spliced feature is input into the channel adaptive enhancement mechanism module to enhance the local cross-feature interaction information in the changed area of the dual-temporal feature map, obtaining a channel-enhanced feature map. The channel-enhanced feature map is added to the spatial channel attention enhanced feature map to obtain a spatio-temporal difference enhanced feature map.

[0093] In the channel alternating splicing branch, the channel adaptive enhancement mechanism can enhance the model's ability to extract cross-feature interaction information. In this way, spatio-temporal difference enhanced feature maps at multiple different scales can be obtained. Compared with other methods of directly subtracting and splicing, these feature maps have less noise and stronger spatio-temporal differences in the dual-temporal image features.

[0094] Optionally, for the channel alternating splicing branch, the first image and the second image group (dual-temporal feature map) are alternately spliced in the channel dimension, and then the channel adaptive enhancement mechanism is introduced to highlight the local cross-feature interaction information in the changed area. Immediately afterwards, it is passed through a 1x1 convolution operation, normalization, and non-linear activation of the ReLu function to obtain the fourth image (channel-enhanced feature map).

[0095] In the embodiments provided by the present application, the channel adaptive enhancement mechanism can enhance the model's ability to extract cross-feature interaction information. For a given feature map A, A ∈ N H×W×C , first, the dual-temporal feature map is alternately spliced in the channel dimension, and then the feature map is subjected to global average pooling to obtain a weight matrix w containing 2C channels, w ∈ N 1 ×1×2C :

[0096]

[0097] i, j, and t are variables representing the elements in the weight matrix.

[0098] Perform a 1D convolution operation on w with a convolution kernel size of k, and then use the Sigmoid activation function to limit the weight size within the interval (0, 1):

[0099] w′ = Sigmoid(Conv1D k (w));

[0100] Among them, the convolution kernel size k is adaptively determined by the number of input channels 2C:

[0101]

[0102] Among them, int represents rounding down, odd represents calculating the nearest odd number, and γ and b are respectively valued at 2 and 1.

[0103] Reset w′∈N 1×1×2C to w″∈N H×W×2C , multiply w″ by A to obtain the feature map after cross-channel feature interaction. After passing this feature map through a 1x1 convolution operation, normalization, and non-linear activation of the ReLu function, the fourth image (channel-enhanced feature map) is obtained.

[0104] Add the channel-enhanced feature map and the spatio-channel attention-enhanced feature map to obtain the spatio-temporal difference-enhanced feature map, that is, sum the third image and the fourth image processed by the double-branch structure to obtain the spatio-temporal difference-enhanced feature map. Among them, the third image is the enhanced feature map of the subtraction branch, and the fourth image is the enhanced feature map of the channel alternating splicing branch. The mathematical formula can be expressed as:

[0105]

[0106] i represents the variable of the elements in the feature map.

[0107] Sum the feature maps processed by the double-branch structure to obtain the spatio-temporal difference-enhanced feature map (S i ) that contains long-distance spatio-temporal change information and local fine-grained change information.

[0108] In the embodiment provided by this application, by introducing the spatio-temporal difference enhancement module of the double-branch structure and using the channel adaptive enhancement mechanism and the dual attention mechanism to capture the remote context information and local cross-feature information of the changing regions in the double-temporal feature maps, the spatio-temporal difference of the double-temporal image features can be enhanced.

[0109] In step 103, input multiple spatio-temporal difference-enhanced feature maps of different scales into the decoder. Through the decoder, edge refinement and feature fusion are performed on the multiple spatio-temporal difference-enhanced feature maps of different scales to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

[0110] An edge refinement residual module and an adaptive bidirectional feature fusion module are introduced into the decoder. By using the residual learning of double-branch edge refinement, the feature maps at different scales can be refined and the representation ability and learning ability of the network can be enhanced to obtain the feature maps processed by the edge refinement residual block. Subsequently, feature fusion at different scales is performed. Through the fusion of the top-down processing level and the bottom-up processing level, the semantic information in the deep network and the position information in the shallow network are fully fused. At the same time, weight parameter η is added during the fusion process to quantify the importance of the feature maps at different levels and different scales, and a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image is obtained.

[0111] In some embodiments, spatio-temporal difference enhancement feature maps of multiple different scales are respectively input into multiple corresponding edge refinement residual modules. The edges of the changing regions in the spatio-temporal difference enhancement feature maps at different scales are enhanced through the edge refinement residual modules to obtain multiple edge-refined feature maps. The multiple edge-refined feature maps are input into an adaptive bidirectional feature fusion module, and the edge-refined feature maps are fused layer by layer from top to bottom and from bottom to top through the adaptive bidirectional feature fusion module to obtain a fused feature map. The fused feature map is upsampled and channel-adjusted to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

[0112] The edge refinement residual module can perform double-branch convolution and basic residual operations on the input spatio-temporal difference enhancement feature maps of multiple different scales, such as the fifth feature map, the sixth feature map, the seventh feature map, and the eighth feature map, respectively, to obtain edge-refined feature maps: the fifth feature map, the sixth feature map, the seventh feature map, and the eighth feature map.

[0113] In the embodiments provided by the present application, by introducing the edge refinement residual module, the model can learn the differences between feature maps of different scales, thereby optimizing the edges of the changing regions in the spatio-temporal difference enhancement feature maps at different scales.

[0114] The adaptive bidirectional weighted feature fusion module (adaptive bidirectional feature fusion module) is used to perform feature fusion on the edge-refined fifth, sixth, seventh, and eighth feature maps input, with a top-down processing level and a bottom-up processing level, so that the semantic information in the deep network and the position information in the shallow network are fully fused to obtain fused feature maps of different scales.

[0115] Exemplarily, the adaptive bidirectional weighted feature fusion module can adopt the PANet model. The PANet model uses a path aggregation module, which ignores the feature weights at different levels and different scales. The features at different levels and different scales have inconsistent effects on the fused output result. Therefore, adopting a top-down processing level combined with a bottom-up processing level can fully fuse the semantic information in the deep network and the position information in the shallow network. By adding a weight parameter η during the fusion process, the importance of feature maps at different levels and different scales can be quantified.

[0116] In the decoder module, first, the fifth, sixth, seventh, and eighth feature maps at different scales output by the encoder are input into the edge refinement residual module to obtain the edge-refined fifth, sixth, seventh, and eighth feature maps. Then, they are input into the adaptive bidirectional weighted feature fusion module to perform feature fusion in a way of dynamically adjusting the feature weights at different scales to obtain an output fused feature map. Finally, bilinear interpolation is used for features Figure 4The upsampling by a factor of two and the 1×1 convolution adjust the number of channels of the feature map to 2, generating the final binary prediction map.

[0117] In the embodiments provided by this application, a decoder is constructed by combining the edge refinement residual module and the adaptive bidirectional feature fusion module. The residual connection and feature refinement operations in the edge refinement residual module can suppress the interference of noise, thereby further improving the representation ability and generalization ability of the network; the adaptive bidirectional feature fusion module performs feature fusion by dynamically adjusting the weights of features at different scales, quantifying the importance of features at different scales, and promoting the interaction between the rich semantic information in the deep network and the rich location information in the shallow network, so as to optimize the spatio-temporal differences at different scales to the greatest extent, enhance the edges of the changing regions in the feature map, and improve the representation ability and learning ability of the network. At the same time, taking into account the feature weights at different levels and different scales, the details of the contribution loss of features at different levels and different scales to the fused feature output are considered, and an accurate prediction binary image with a low false detection rate is generated.

[0118] In some embodiments of this application, the edge refinement residual module includes a first branch and a second branch. Multiple spatio-temporal difference enhanced feature maps at different scales can be input into the corresponding first branch. The first branch unifies the number of channels of the spatio-temporal difference enhanced feature maps at different scales to obtain multiple standard channel feature maps; the multiple standard channel feature maps are input into the corresponding second branch, and the second branch refines the spatio-temporal difference enhanced feature maps at different scales to obtain multiple refined feature maps; the multiple standard channel feature maps and the refined feature maps are summed respectively to obtain multiple edge refinement feature maps.

[0119] The feature edge refinement residual module adopts a dual-branch structure, which can further optimize the edges of the changing regions in the spatio-temporal difference enhanced feature maps at different scales, and improve the representation ability and learning ability of the network.

[0120] Optionally, reference can be made to Figure 4 , Figure 4 FIG. is a schematic diagram of the network structure of the feature edge refinement residual module of the decoder part provided by some embodiments of this application. Multiple spatio-temporal difference enhanced feature maps at different scales, such as the fifth feature map, the sixth feature map, the seventh feature map, and the eighth feature map, are respectively input into the dual-branch structure of the edge refinement residual module. The first branch uses a 1×1 convolutional layer to unify the number of channels of the feature maps F i at different scales to C2, and this process can be expressed as:

[0121]

[0122] where Conv1 represents the convolution operation with a convolution kernel size of 1, and i represents the variable of the elements in the feature map.

[0123] After that, the feature map that has passed through the first branch is input into the second branch. The second branch is a basic residual block operation. By performing a 3×3 convolution operation, batch normalization, ReLU activation function, and 3×3 convolution operation on the input feature map, the feature map at different scales can be refined, and the representation ability and learning ability of the network can be enhanced. The mathematical formula is as follows:

[0124]

[0125] Among them, Conv3 represents a convolution operation with a convolution kernel size of 3, and BN and ReLU represent batch normalization and activation operations respectively. Finally, F1 i and F2 i are summed to obtain the feature map after being processed by the edge refinement residual block:

[0126]

[0127] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the network structure of the decoder part adaptive bidirectional feature fusion module provided by some embodiments of the present application. Figure 5 It includes an input feature map, a top-down branch, and a bottom-up branch. In the adaptive bidirectional weighted feature fusion module, the process of obtaining the fusion feature map at different scales can be: performing top-down hierarchical fusion on the fifth, sixth, seventh, and eighth feature maps with edge refinement at four different scales to obtain the fusion feature map at any scale in the second level.

[0128] As Figure 5 shown, in the top-down branch, starting from the first level, the edge-refined feature maps between different scales are operated on and fused layer by layer from top to bottom to obtain the feature fusion map F 2,i in the second level. The specific formula is:

[0129]

[0130] Among them, Conv1 represents a convolution operation with a convolution kernel size of 1, η is a learnable weight parameter, the value of ∈ is 0.00001 to avoid the sum of weights being 0, and Down is a downsampling operation.

[0131] Optionally, in the bottom-up branch, the feature fusion maps at different scales in the second level are operated on by a formula and fused layer by layer from bottom to top to obtain the feature fusion map in the third level. The specific formula is:

[0132]

[0133] Among them, Conv1 represents a convolution operation with a convolution kernel size of 1, η’ is a learnable weight parameter, the value of ∈ is 0.00001 to avoid the sum of weights being 0, and Up is an upsampling operation.

[0134] Exemplarily, the decoder includes four edge refinement residual modules, an adaptive bidirectional weighted feature fusion module, a 1x1 convolution layer, and an upsampling module. The fifth feature map, the sixth feature map, the seventh feature map, and the eighth feature map are input into the edge refinement residual modules to obtain the edge-refined fifth feature map, sixth feature map, seventh feature map, and eighth feature map. The edge-refined fifth feature map, sixth feature map, seventh feature map, and eighth feature map are input into the adaptive bidirectional weighted feature fusion module to perform feature fusion at different levels at different scales, so as to output a fusion feature map with a size and number of channels of H / 4×W / 4×C2. Subsequently, the fusion feature map is upsampled 4 times using bilinear interpolation, and finally the number of channels of the feature map is adjusted to 2 using a 1×1 convolution to obtain the final predicted binary image.

[0135] In the embodiments provided by the present application, in the decoder, an edge refinement residual module is introduced, enabling the remote sensing image change detection model to optimize the edges of the changed regions in the spatio-temporal difference enhanced feature maps at different scales and improving the model's ability to capture various image features. An adaptive bidirectional weighted feature fusion module is introduced to quantify the importance of features at different scales by using learnable weight parameters to achieve effective fusion of multi-scale features.

[0136] Please refer to Figure 6 , Figure 6 is a schematic diagram comparing the visualization results of the remote sensing image change detection method provided by some embodiments of the present application and the existing remote sensing image change detection method on the WHU-CD dataset. The WHU-CD dataset consists of a pair of remote sensing images with a spatial resolution of 0.7m, and the image size is 32507×15354. The original images can be cropped into non-overlapping images with a size of 512×512 pixels, and 1461, 183, and 183 pairs of images are randomly divided for training, validation, and testing. Figure 6 The images presented in are the remote sensing images, labels, various comparison methods, and the predicted binary images of the method of the present invention at different time points on the WHU-CD dataset. Among them, the white areas of the binary images are the changed regions, the black areas are the unchanged regions, the green areas are the missed detection regions, and the red areas are the misdetection regions.

[0137] Please refer to Figure 7 , Figure 7It is a schematic diagram for comparing the visualization results of the remote sensing image change detection method provided by some embodiments of this application and the existing remote sensing image change detection method on the LEVIR-CD dataset. The LEVIR-CD dataset is a large-scale dataset for building change detection tasks. The original training set, validation set, and test set contain 445, 64, and 128 pairs of images respectively, with a spatial resolution of 0.5m and an image size of 1024×1024. Each image can be cropped into non-overlapping images with a size of 512×512 pixels. The final training set, validation set, and test set sizes are 1780, 256, and 512 pairs of images respectively. Figure 7 The images presented in Figure 7 are remote sensing images, labels, various comparison methods, and the predicted binary images of the method of the present invention at different time points on the LEVIR-CD dataset. Among them, the white area in the binary image is the changed area, the black area is the unchanged area, the green area is the missed detection area, and the red area is the misdetection area.

[0138] Please refer to Figure 8 , Figure 8 It is a schematic diagram for comparing the visualization results of the remote sensing image change detection method provided by some embodiments of this application and the existing remote sensing image change detection method on the SYSU-CD dataset. The SYSU-CD dataset covers various types of land cover changes, including buildings, vegetation, roads, and water areas, etc. The size of each pair of images is 256×256. The dataset contains 20,000 pairs of 0.5m aerial images taken during the period from 2007 to 2014. The training set, validation set, and test set sizes are 12,000, 4,000, and 4,000 pairs of images respectively. Figure 8 The images presented in Figure 8 are remote sensing images, labels, various comparison methods, and the predicted binary images of the method of the present invention at different time points on the SYSU-CD dataset. Among them, the white area in the binary image is the changed area, the black area is the unchanged area, the green area is the missed detection area, and the red area is the misdetection area.

[0139] Figures 6 to 8 In Figures 6 to 8 , the comparison methods adopted are divided into the following types. Among them, Tiny-CD and RFANet are lightweight networks with a differential feature semantic enhancement module. FC Siam diff and SNUNet are fully convolutional neural networks with a skip connection structure. ChangeFormer is a pure Transformer architecture network with a multi-head attention mechanism, and BIT is an integrated deep convolutional neural network and Transformer architecture network. TFI_GR is a network with a spatio-temporal differential feature interaction module, and DMINet introduces a cross-attention mechanism and a self-attention mechanism network as well as the network SEAFNet of the present invention.

[0140] From Figure 6, Figure 7 and Figure 8 It can be seen that in terms of large-scale ground object changes, compared with the other six methods, the method SEAFNet of this application has the best extraction effect on changes such as large buildings and large-area surface vegetation, and the edge structure of the extracted change area is relatively complete, without obvious false detections or missed detections. For example, in Figure 6 the newly built buildings, Figure 7 the newly built parks, and Figure 8 the extraction of newly added vegetation, except for the method SEAFNet of this invention, other methods all showed varying degrees of missed detections. And in Figure 6 the extraction of newly added buildings, except for the comparison method DMINet and the method SEAFNet of this invention, other methods all showed varying degrees of false detections.

[0141] In terms of medium and small-scale ground object changes, from Figure 6 , Figure 7 and Figure 8 it can be seen that the method SEAFNet of this invention has the lowest degree of missed detections and false detections in changes such as small buildings and narrow buildings. And in Figure 7 the newly added slender buildings and Figure 8 the newly added small building change scenarios, the degree of missed detections of other methods is more serious than that of the method SEAFNet of this invention.

[0142] In terms of pseudo-changes, from Figure 7 and Figure 8 it can be seen that compared with other methods, the method SEAFNet of this invention has the lowest degree of false detections in pseudo-change scenarios. For example, in Figure 7 the pseudo-change scenarios caused by seasonal factors, other methods all showed varying degrees of false detections, and the method ChangeFormer with long-distance image dependence relationship performed the worst, while the method SEAFNet of this invention performed the best.

[0143] Please refer to Figures 9(a) to 9(c) , Figures 9(a) to 9(c)It respectively shows the schematic diagram of the comparison results of the computational complexity of the change detection method SEAFNet for remote sensing images provided by the embodiments of the present application and the existing change detection methods for remote sensing images on the WHU-CD dataset, LEVIR-CD dataset, and SYSU-CD dataset. Among them, the abscissa represents the FLOPs of the model, and the ordinate represents the Iou value of the model. It can be seen from this that the method of the present invention achieves the optimal Iou performance with the third-best FLOPs, and has a certain advantage in the number of parameters. Although the FLOPs and the number of parameters of the Tiny-CD and REFNet models are lower than those of SEAFNet, the quantitative and qualitative analysis results of SEAFNet are better. The TFI-GR and DMINet models embedded with the spatio-temporal difference interaction module perform relatively well on all datasets, but there are still certain deficiencies in processing multi-scale feature maps, and their FLOPs are higher than the method of this article. Comprehensive above analysis fully shows that in the task of remote sensing image change detection, enhancing the difference between bi-temporal images and effectively fusing feature maps at different scales is particularly crucial, and the change detection method for remote sensing images provided by the embodiments of the present application achieves the best balance in terms of considering the computational efficiency and accuracy of the model.

[0144] The above is the introduction to the embodiments of the change detection method for remote sensing images of the present application.

[0145] Next, the implementation manner of the change detection device for remote sensing images provided by the embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0146] For the change detection method for remote sensing images provided in the above embodiments, the embodiments of the present application also provide a change detection device for remote sensing images for implementing the above method, as Figure 10 shown, Figure 10 is the module schematic block diagram of the change detection device for remote sensing images of the embodiments of the present application. The change detection device for remote sensing images includes:

[0147] An acquisition module, configured to acquire a pre-temporal remote sensing image and a post-temporal remote sensing image. The pre-temporal remote sensing image is used to represent the surface feature image of the target area at the first time point, and the post-temporal remote sensing image is used to represent the surface feature image of the target area at the second time point;

[0148] A feature extraction and generation module is configured to input a pre-temporal remote sensing image and a post-temporal remote sensing image into an encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales. The encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The feature extraction module is configured to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple bi-temporal feature maps at different scales. The multiple spatio-temporal difference enhancement modules are configured to capture the remote context information and local cross-feature information of the changing regions in the multiple bi-temporal feature maps at different scales to generate multiple spatio-temporal difference enhanced feature maps.

[0149] A feature refinement and fusion module is configured to input the multiple spatio-temporal difference enhanced feature maps at different scales into a decoder. The decoder performs edge refinement and feature fusion on the multiple spatio-temporal difference enhanced feature maps at different scales to obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

[0150] It can be understood that the content in the above method embodiments is applicable to the present device embodiment. The functions specifically implemented in the present device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0151] As Figure 11 shown, the embodiment of the present application further provides an electronic device. The electronic device includes a memory, one or more processors ( Figure 11 only one is shown in the figure) and a computer program stored in the memory and executable on the processor. Among them: the memory is used to store software programs and units. The processor executes various functional applications and data processing by running the software programs and units stored in the memory to obtain the resources corresponding to the above preset events. Optionally, when the processor runs the above computer program stored in the memory, the above method for change detection of remote sensing images is implemented.

[0152] As a non-transitory computer-readable medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network.

[0153] It can be understood that the content in the above method embodiments is applicable to the present electronic device embodiment. The functions specifically implemented in the present electronic device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0154] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method for change detection of remote sensing images is implemented.

[0155] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0156] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program. When the above computer program is executed by one or more processors, the steps of the method for change detection of remote sensing images as described above can be implemented.

[0157] It can be understood that the content in the above method embodiments is applicable to this computer program product. The functions specifically implemented by the embodiments of this computer program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0158] The method, device, electronic device, medium, and computer program product for change detection of remote sensing images provided by the embodiments of the present application extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image through a feature extraction module to generate multiple dual-temporal feature maps of different scales. Then, through multiple spatio-temporal difference enhancement modules, the remote context information and local cross-feature information of the change regions in the multiple dual-temporal feature maps of different scales are captured to generate multiple spatio-temporal difference enhancement feature maps of different scales, which can learn the global change information and local fine-grained change information between the dual-temporal feature maps at the same scale, enhance the spatio-temporal difference of the dual-temporal image features extracted by the encoder. Through the decoder, edge refinement and feature fusion are performed on the multiple spatio-temporal difference enhancement feature maps of different scales, so that the rich position information in the shallow feature map and the rich semantic information in the deep feature map are effectively fused, while taking into account the feature weights at different levels and different scales, and the details of the contributions of the features at different levels and different scales to the fused feature output, thereby improving the accuracy of remote sensing image change detection.

[0159] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0160] Although specific embodiments are described herein, those of ordinary skill in the art will recognize that many other modifications or alternative embodiments are also within the scope of the present disclosure. For example, any one of the functions and / or processing capabilities described in connection with a particular device or component can be performed by any other device or component. Additionally, although various illustrative implementations and architectures have been described in accordance with embodiments of the present disclosure, those of ordinary skill in the art will recognize that many other modifications to the illustrative implementations and architectures described herein are also within the scope of the present disclosure.

[0161] Certain aspects of the present disclosure have been described above with reference to block diagrams and flowcharts of systems, methods, systems, and / or computer program products according to exemplary embodiments. It should be understood that one or more blocks in the block diagrams and flowcharts, and combinations of blocks in the block diagrams and flowcharts, can be implemented respectively by executing computer-executable program instructions. Similarly, according to some embodiments, some blocks in the block diagrams and flowcharts may not need to be executed in the order shown, or may not need to be executed at all. Additionally, additional components and / or operations beyond those shown in the blocks of the block diagrams and flowcharts may exist in certain embodiments.

[0162] Accordingly, the blocks in the block diagrams and flowcharts support combinations of means for performing the specified functions, combinations of elements or steps for performing the specified functions, and means for program instructions for performing the specified functions. It should also be understood that each block in the block diagrams and flowcharts, and combinations of blocks in the block diagrams and flowcharts, can be implemented by a special-purpose hardware computer system that performs a particular function, element, or step, or by a combination of special-purpose hardware and computer instructions.

[0163] The program modules, applications, etc. described herein may include one or more software components, including, for example, software objects, methods, data structures, etc. Each such software component may include computer-executable instructions that, in response to execution, cause at least a portion of the functions described herein (e.g., one or more operations of the illustrative methods described herein) to be performed.

[0164] Software components can be coded in any of a variety of programming languages. An exemplary programming language can be a low-level programming language, such as an assembly language associated with a particular hardware architecture and / or operating system platform. Software components including assembly language instructions may need to be converted by an assembler into executable machine code before being executed by the hardware architecture and / or platform. Another exemplary programming language can be a higher-level programming language, which can be portable across multiple architectures. Software components including a higher-level programming language may need to be converted by an interpreter or compiler into an intermediate representation before execution. Other examples of programming languages include, but are not limited to, macro languages, shell or command languages, job control languages, scripting languages, database query or search languages, or report writing languages. In one or more exemplary embodiments, a software component including instructions in one of the above examples of programming languages can be directly executed by an operating system or other software component without first being converted into another form.

[0165] Software components can be stored as files or other data storage constructs. Software components with similar types or related functions can be stored together in, for example, a particular directory, folder, or library. Software components can be static (e.g., pre-set or fixed) or dynamic (e.g., created or modified at execution time).

[0166] The embodiments of the present application have been described in detail above with reference to the accompanying drawings. However, the present application is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art to which the present application pertains, various changes can be made without departing from the gist of the present application.

Claims

1. A method for change detection of remote sensing images, characterized in that, Including the following steps: Obtain a pre-temporal remote sensing image and a post-temporal remote sensing image. The pre-temporal remote sensing image is used to characterize the surface feature image of the target area at the first time point, and the post-temporal remote sensing image is used to characterize the surface feature image of the target area at the second time point; Input the pre-temporal remote sensing image and the post-temporal remote sensing image into an encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales. Among them, the encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The spatio-temporal difference enhancement module includes a channel alternating splicing branch. The feature extraction module is used to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple double-temporal feature maps at different scales. The multiple spatio-temporal difference enhancement modules are used to capture the remote context information and local cross-feature information of the changing areas in the multiple double-temporal feature maps at different scales to generate the multiple spatio-temporal difference enhanced feature maps at different scales; Perform alternating splicing on the double-temporal feature maps in the channel alternating splicing branch, and then input them into the channel adaptive enhancement mechanism module to obtain channel enhanced feature maps; Add the channel enhanced feature maps to the spatial-channel attention enhanced feature maps to obtain spatio-temporal difference enhanced feature maps; Input the multiple spatio-temporal difference enhanced feature maps at different scales into the corresponding first branch, and unify the number of channels of the spatio-temporal difference enhanced feature maps at different scales through the first branch to obtain multiple standard channel feature maps; Input the multiple standard channel feature maps into the corresponding second branch, and refine the spatio-temporal difference enhanced feature maps at different scales through the second branch to obtain multiple refined feature maps. Among them, the second branch includes a convolution operation, batch normalization, activation function, and convolution operation arranged in sequence; Sum the multiple standard channel feature maps and the refined feature maps respectively to obtain multiple edge refined feature maps; Input the multiple edge refined feature maps into the adaptive bidirectional feature fusion module, and perform top-down and bottom-up layer-by-layer fusion on the edge refined feature maps through the adaptive bidirectional feature fusion module to obtain a fused feature map, and add learnable weight parameters during the fusion process; Obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

2. The change detection method for remote sensing images according to claim 1, wherein The feature extraction module includes two parallel feature extraction branches. The step of inputting the pre-temporal remote sensing image and the post-temporal remote sensing image into the encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales includes: Input the pre-temporal remote sensing image and the post-temporal remote sensing image into two parallel feature extraction branches respectively, and extract multiple double-temporal feature maps at different scales through the feature extraction network in the feature extraction branch; Input the multiple double-temporal feature maps at different scales into the corresponding spatio-temporal difference enhancement modules respectively to obtain multiple spatio-temporal difference enhanced feature maps at different scales.

3. The change detection method for remote sensing images according to claim 2, wherein The spatio-temporal difference enhancement module includes a subtraction branch. The step of inputting the double-temporal feature maps into the spatio-temporal difference enhancement module to obtain the spatio-temporal difference enhanced feature maps includes: Input the dual - temporal feature map into the subtraction branch, and calculate the absolute difference feature between the dual - temporal feature maps in the subtraction branch; Input the absolute difference feature into the dual - attention mechanism module, and capture the long - range context information of the changing regions in the dual - temporal feature maps through the dual - attention mechanism module to obtain a spatially attention - enhanced feature map and a channel - attention - enhanced feature map; Perform channel concatenation on the spatially attention - enhanced feature map and the channel - attention - enhanced feature map to obtain a spatially and channel - attention - enhanced feature map; Add the spatially and channel - attention - enhanced feature map to the channel - enhanced feature map to obtain the spatio - temporal difference - enhanced feature map.

4. The change detection method for remote sensing images according to claim 3, characterized in that, The dual - attention mechanism module includes a spatial attention mechanism module and a channel attention mechanism module. The step of inputting the absolute difference feature into the dual - attention mechanism module and capturing the long - range context information of the changing regions in the dual - temporal feature maps through the dual - attention mechanism module to obtain a spatially attention - enhanced feature map and a channel - attention - enhanced feature map includes: Input the absolute difference feature into the spatial attention mechanism module, and establish the context relationship of local features through the spatial attention mechanism module to obtain the spatially attention - enhanced feature map; Input the absolute difference feature into the channel attention mechanism module, and establish the long - range dependence relationship between the channels of the feature maps through the channel attention mechanism module to obtain the channel - attention - enhanced feature map.

5. The change detection method for remote sensing images according to claim 2, characterized in that, The spatio - temporal difference - enhanced module includes a channel - alternating concatenation branch. The step of inputting the dual - temporal feature map into the spatio - temporal difference - enhanced module to obtain multiple spatio - temporal difference - enhanced feature maps of different scales includes: Input the dual - temporal feature map into the channel - alternating concatenation branch, and perform alternating concatenation on the dual - temporal feature maps in the channel - alternating concatenation branch to obtain a concatenated feature; Input the concatenated feature into the channel - adaptive enhancement mechanism module, and enhance the local cross - feature interaction information of the changing regions in the dual - temporal feature maps through the channel - adaptive enhancement mechanism module to obtain a channel - enhanced feature map; Add the channel - enhanced feature map to the spatially and channel - attention - enhanced feature map to obtain the spatio - temporal difference - enhanced feature map.

6. The change detection method for remote sensing images according to claim 1, wherein Obtain a predicted binary image including the change information between the pre - temporal remote - sensing image and the post - temporal remote - sensing image, including: Upsample and adjust the channels of the fused feature map to obtain a predicted binary image including the change information between the pre - temporal remote - sensing image and the post - temporal remote - sensing image.

7. A change detection device for remote sensing images, characterized in that, The device includes: An acquisition module, configured to acquire a pre - temporal remote - sensing image and a post - temporal remote - sensing image. The pre - temporal remote - sensing image is used to represent the surface feature image of the target area at the first time point, and the post - temporal remote - sensing image is used to represent the surface feature image of the target area at the second time point; A feature extraction and generation module is configured to input the pre-temporal remote sensing image and the post-temporal remote sensing image into an encoder to obtain multiple spatio-temporal difference enhanced feature maps at different scales. The encoder includes a feature extraction module and multiple spatio-temporal difference enhancement modules. The spatio-temporal difference enhancement module includes a channel alternating splicing branch. The feature extraction module is configured to extract features from the pre-temporal remote sensing image and the post-temporal remote sensing image to generate multiple dual-temporal feature maps at different scales. The multiple spatio-temporal difference enhancement modules are configured to capture remote context information and local cross-feature information of the changing regions in the multiple dual-temporal feature maps at different scales to generate the multiple spatio-temporal difference enhanced feature maps. In the channel alternating splicing branch, the dual-temporal feature maps are alternately spliced and input into a channel adaptive enhancement mechanism module to obtain channel enhanced feature maps. The channel enhanced feature maps are added to the spatial-channel attention enhanced feature maps to obtain the spatio-temporal difference enhanced feature maps; Input the multiple spatio-temporal difference enhanced feature maps at different scales into the corresponding first branch, and unify the number of channels of the spatio-temporal difference enhanced feature maps at different scales through the first branch to obtain multiple standard channel feature maps; Input the multiple standard channel feature maps into the corresponding second branch, and refine the spatio-temporal difference enhanced feature maps at different scales through the second branch to obtain multiple refined feature maps. The second branch includes a convolutional operation, batch normalization, an activation function, and a convolutional operation arranged in sequence; Sum the multiple standard channel feature maps and the multiple refined feature maps respectively to obtain multiple edge-refined feature maps; Input the multiple edge-refined feature maps into an adaptive bidirectional feature fusion module, and perform top-down and bottom-up layer-by-layer fusion on the edge-refined feature maps through the adaptive bidirectional feature fusion module to obtain a fused feature map, and add learnable weight parameters during the fusion process; Obtain a predicted binary image including the change information between the pre-temporal remote sensing image and the post-temporal remote sensing image.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the change detection method of the remote sensing image according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the change detection method of the remote sensing image according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Illegal building change detection method and device based on unmanned aerial vehicle, equipment and medium

    CN118212517A

  • Small sample hyperspectral remote sensing image change detection method based on graph convolution

    CN118447395A