A natural reserve remote sensing image change detection network, detection method and detection device
By combining the feature extraction network with the backbone network and utilizing the SAM pre-trained weights and multi-scale feature interaction module, the problems of insufficient feature extraction and false changes in change detection in remote sensing images of nature reserves are solved, thereby improving detection accuracy and training efficiency.
Patent Information
- Application Number
- CN202510133629.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-02-06
AI Technical Summary
Existing remote sensing image change detection technologies based on convolutional neural networks and Transformers have insufficient feature extraction capabilities and poor robustness in nature reserves, and are prone to missing complex environmental changes. In addition, the amount of training data for large SAM models is insufficient, and it is difficult to effectively alleviate the phenomenon of pseudo-changes.
The feature extraction network is combined with the backbone network, the weights are initialized through SAM pre-training, the multi-scale feature interaction module and the change guidance unit are combined, and the mask extraction and position attention modules are used to enhance the feature extraction and pseudo-change suppression capabilities.
The accuracy of change detection in remote sensing images of nature reserves is improved, the missed detection rate and the impact of false changes are reduced, the training efficiency is improved, and overfitting is avoided.
Smart Images

Figure CN120125991B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image change detection, and in particular relates to a remote sensing image change detection network, a detection method and a detection device for nature reserves. Background Art
[0002] Nature reserves are crucial for ecosystem conservation and biodiversity maintenance. In recent years, the ecological environment in these areas has faced increasing threats due to human activities and natural disasters. Therefore, it is crucial to better detect changes within nature reserves caused by human activities.
[0003] Currently, AI-based change detection in remote sensing images has become a common technique for identifying changes within nature reserves caused by human activities. This technique often uses one or more deep learning models, such as convolutional neural networks, Transformer neural networks, and the SAM model, to construct image change detection networks, helping to support the precise management of nature reserves.
[0004] However, existing image change detection networks built using the aforementioned convolutional neural networks and Transformers have limited feature extraction capabilities and are prone to missed detections when detecting changes in nature reserves in complex environments. This results in insufficient robustness when extracting the complex and diverse environmental features of nature reserves. Furthermore, large SAM models typically contain tens to billions of parameters, and training often requires training sets many times larger than those of ordinary models to support their feature extraction performance and avoid overfitting. Consequently, using large SAM models to build network models often faces the problem of insufficient training data.
[0005] Furthermore, nature reserves often experience significant artifacts due to various factors beyond actual ground features. These include image differences caused by varying lighting conditions and seasonal variations. These AI-based remote sensing image change detection techniques typically directly extract difference features after extracting image features. This approach fails to accurately capture actual changes in the inspected area and mitigates the severe artifacts within nature reserves. Consequently, the accuracy of these models in detecting changes in remote sensing imagery of nature reserves is limited. Summary of the Invention
[0006] The present invention provides a network, a detection method and a detection device for detecting changes in remote sensing images of nature reserves, so as to solve at least one of the above technical problems.
[0007] In a first aspect, the present invention provides a nature reserve remote sensing image change detection network, comprising:
[0008] Backbone network, used to extract the input area to be detected for each dual-phase remote sensing image X1, X2 in the dual-phase remote sensing image pair for remote sensing image change detection t Four feature maps of different preset scales X1 and X2 are the remote sensing images before and after the time phase, Represents X t The i-th feature map, i = 1, 2, 3, 4, The scale of is reduced successively; X1 and X2 are an image pair after spatial registration;
[0009] The feature extraction network uses a SAM feature extractor to extract the input image pair X1, X2 for each dual-phase remote sensing image X t The feature map containing global information The model parameters of the SAM feature extractor are the SAM pre-trained weights;
[0010] The first feature enhancement unit is connected to the backbone network and the feature extraction network, and is used to extract the feature based on the feature extraction network. Extraction of backbone network Enhance and obtain the first enhanced feature map and the second enhanced feature map
[0011] The multi-scale feature interaction module is connected to the first feature enhancement unit and is used to enhance the first feature enhancement unit Perform feature interaction to obtain Enhanced feature map
[0012] The second feature enhancement unit is connected to the backbone network, the first feature enhancement unit, and the multi-scale feature interaction module, and is used to obtain the feature information based on the multi-scale feature interaction module. right Perform feature enhancement to obtain enhanced feature map
[0013] The first change feature extraction module is connected to the multi-scale feature interaction module and the second feature enhancement unit to extract and The changing characteristics of D i , i=1,2,3,4;
[0014] A first mask extraction module, connected to the first change feature extraction module, is used to obtain a mask Guidemap of the change feature D4;
[0015] A change guidance unit, connected to the first change feature extraction module and the first mask extraction module, is used to obtain a change feature G3 of the change features D2, D3, and D4 after the change guidance of the mask Guidemap and self-attention;
[0016] The fourth feature fusion module is connected to the change guidance unit and the first change feature extraction module, and is used to fuse the features of G3 and D1 to obtain the feature G′3;
[0017] The second change feature extraction unit is connected to the feature extraction network and is used to extract the Get the change feature D corresponding to X1 and x2 sam ;
[0018] Position attention module, connected to the second change feature extraction unit, for sam Perform feature extraction, and the obtained feature map is recorded as the position attention weight map;
[0019] The third feature enhancement module is connected to the position attention module and the fourth feature fusion module, and is used to multiply the position attention weight map with the feature G′3 to obtain the feature D defined ;
[0020] The second mask extraction module is connected to the third feature enhancement module to extract feature D defined The change mask M is obtained, that is, the remote sensing image change of the area to be detected is obtained; where t=1,2.
[0021] Furthermore, the first feature enhancement unit includes:
[0022] The first 3×3 convolutional layer is connected to the feature extraction network and is used to adjust the features extracted by the feature extraction network. The scale and number of channels are obtained by The same feature map
[0023] The second 3×3 convolutional layer, connected to the first 3×3 convolutional layer, is used to adjust The scale and number of channels are obtained by The same feature map
[0024] The first feature enhancement module is connected to the first 3×3 convolutional layer and the backbone network to and Add to get the first enhanced feature map
[0025] The second feature enhancement module is connected to the second 3×3 convolutional layer and the backbone network to and Add to get the second enhanced feature map
[0026] The second change feature extraction unit includes the second 3×3 convolutional layer and the first 3×3 convolutional layer; the second change feature extraction unit also includes:
[0027] The third 3×3 convolutional layer, connected to the second 3×3 convolutional layer, is used to adjust The scale and number of channels of G′3 are obtained to obtain the feature map with the same scale and number of channels as G′3
[0028] The second change feature extraction module is connected to the third 3×3 convolution layer to extract feature maps With feature map The characteristics of the change feature D are obtained sam ;
[0029] The change guidance unit includes:
[0030] The first change guidance module is connected to the first mask extraction module and the first change feature extraction module, and is used to calculate the change feature G1 of D4 after the mask Guidemap and self-attention change guidance;
[0031] The first channel splicing layer is connected to the first change guidance module and the first change feature extraction module, and is used to splice the feature G1 and the feature D3 to obtain the first splicing feature;
[0032] The second change guidance module is connected to the first mask extraction module and the first channel splicing layer, and is used to calculate the change feature G2 of the first splicing feature after the mask Guidemap and self-attention change guidance;
[0033] The second channel splicing layer is connected to the second change guidance module and the first change feature extraction module, and is used to splice the feature G2 and the feature D2 to obtain the second splicing feature;
[0034] A third change guidance module is connected to the first mask extraction module and the second channel splicing layer, and is used to calculate the change characteristics of the second splicing feature after the mask Guidemap and the self-attention change guidance, that is, to obtain the G3;
[0035] The second feature enhancement unit includes:
[0036] The first upsampling module is connected to the multi-scale feature interaction module to Upsampling to get the scale and The same feature map
[0037] The first feature fusion module is connected to the first upsampling module and the first feature enhancement module, and is used to and Perform feature fusion to obtain the enhanced feature map
[0038] The second upsampling module is connected to the first feature fusion module and is used to Upsampling to get the scale and The same feature map
[0039] The second feature fusion module is connected to the second upsampling module and the backbone network to and Perform feature fusion to obtain the enhanced feature map
[0040] The third upsampling module is connected to the second feature fusion module to Upsampling is performed to obtain the scale and The same feature map
[0041] The third feature fusion module is connected to the third upsampling module and the backbone network to and Perform feature fusion to obtain the enhanced feature map
[0042] Furthermore, the first mask extraction module and the second mask extraction module each include:
[0043] The second convolution combination unit is used to reduce the number of channels of the input change feature to 1 / 2 of the original number, thereby obtaining a change feature with reduced channel number;
[0044] The fourth 3×3 convolutional layer is connected to the second convolution combination unit and is used to perform mask extraction on the change feature with reduced channel number to obtain a mask with a channel number of 1, which is a mask of the change feature of the input.
[0045] Furthermore, the methods for calculating the change features of the target features after the mask Guidemap and the self-attention change guidance by the first change guidance module, the second change guidance module, and the third change guidance module all include:
[0046] Calculate the product of the target feature and the mask Guidemap to obtain the features of the enhanced change area;
[0047] Based on the features of the enhanced change area and the mask Guidemap, a self-attention of the enhanced change weight is obtained;
[0048] The features of the enhanced change region are input into the self-attention of the enhanced change weight to obtain the change features after change guidance.
[0049] Furthermore, the position attention module includes a first 3×3 convolution unit and a first 1×1 convolution unit;
[0050] The first 3×3 convolution unit is connected to the second change feature extraction module, and is used to extract the D sam Perform feature extraction and obtain the number of channels D sam The feature map with one quarter of the number of channels is recorded as the first feature map;
[0051] The first 1×1 convolution unit is connected to the first 3×3 convolution unit and is used to adjust the number of channels of the first feature map to generate a feature map with the same number of channels as G′3, which is the position attention weight map.
[0052] Furthermore, the multi-scale feature interaction module includes a channel stacking unit, four dilated convolutions with different dilation rates, a third channel splicing layer, a second 3×3 convolution unit and a feature interaction unit;
[0053] The channel stacking unit is connected to the first feature enhancement unit and is used to Perform channel stacking;
[0054] Four dilated convolutions with different expansion rates are connected to the channel stacking units to extract features of four different scales from the channel stacking results;
[0055] The third channel splicing layer is connected to four dilated convolutions with different expansion rates. It is used to perform channel splicing on the convolution results of the four dilated convolutions to obtain the third splicing feature.
[0056] The second 3×3 convolution unit is connected to the third channel splicing layer, and is used to perform 3×3 convolution processing on the third splicing feature and record the processing result as the bi-temporal feature interaction weight w;
[0057] The feature interaction unit is connected to the second 3×3 convolution unit to combine the bi-temporal feature interaction weight w with the feature Multiply to obtain enhanced features after feature interaction
[0058] Furthermore, the method for acquiring the remote sensing image change detection network includes:
[0059] Building a network model of the remote sensing image change detection network;
[0060] Construct a training set;
[0061] Initialize the built network model to obtain the initialized network model of the remote sensing image change detection network; in the initialized network model, the model parameters of each backbone network are initialized using the pre-trained weights of VGG16, and the model parameters of each SAM feature extractor are initialized using the pre-trained weights of SAM;
[0062] During training, the model parameters of each SAM feature extractor in the initialized network model are frozen, and then the data in the training set is input into the initialized network model for iterative training. During each iterative training process, the predicted output is obtained through forward propagation, and then the loss function is used to calculate the loss between the predicted output and its corresponding true label. Then, the network is back-propagated according to the calculated loss, and then the gradient descent algorithm is used to update the parameters in the network model during back-propagation until the loss function converges or the training reaches a preset number of iterations, thereby obtaining the remote sensing image change detection network.
[0063] Furthermore, the training set construction method includes:
[0064] Select two temporal remote sensing images of the nature reserve, download the data, and obtain the original dual-temporal remote sensing image pair;
[0065] Perform spatial registration on the original dual-temporal remote sensing image pairs, and perform data cropping on the registered original dual-temporal remote sensing image pairs to obtain several dual-temporal remote sensing image pairs;
[0066] For each cropped dual-temporal remote sensing image pair: mark the changed areas in the latter-phase image relative to the former-phase image to obtain the change mask labels;
[0067] Each cropped dual-temporal remote sensing image pair and its corresponding change mask label is taken as a sample, and the samples are collected to obtain the training set.
[0068] In a second aspect, the present invention provides a method for detecting changes in remote sensing images of a nature reserve, the method comprising:
[0069] Collect a pair of dual-temporal remote sensing images of the area to be detected for remote sensing image change detection, which is recorded as a detection image pair;
[0070] Perform spatial registration on the detection image pair to obtain the registered detection image pair, which is recorded as the target image pair;
[0071] The target image pair is input into the remote sensing image change detection network described in the above aspects to perform operations, and the image changes of the to-be-detected area of the to-be-detected nature reserve are obtained.
[0072] In a third aspect, the present invention provides a device for detecting changes in remote sensing images of nature reserves, the device comprising a data acquisition device and an image change detection device;
[0073] The data acquisition device is used to acquire a pair of dual-temporal remote sensing images of the area to be detected for remote sensing image change detection, thereby obtaining a target dual-temporal remote sensing image pair;
[0074] An image change detection device integrates an image preprocessing unit and the remote sensing image change detection network described in the above aspects; the image change detection device is connected to a data acquisition device, and is used to call the image preprocessing unit to perform spatial registration on the target dual-phase remote sensing image pair, and then input the target dual-phase remote sensing image pair that has completed spatial registration into the remote sensing image change detection network for calculation to obtain the image changes of the area to be detected.
[0075] It can be seen from the above technical solutions that the present invention has the following advantages:
[0076] The present invention adopts a feature extraction network as an auxiliary means of feature extraction, and enhances the feature map extracted by the backbone network based on the feature map extracted by the feature extraction network, which helps to enhance the feature extraction capability and reduce the missed detection rate in change detection.
[0077] The feature extraction network of the present invention adopts the SAM large model, and its model parameters are the SAM pre-trained weights. It can be seen that the present invention does not need to train and update the model parameters of SAM, which helps to solve the problem of insufficient data when using the SAM large model to build a model, and then helps to improve training efficiency and avoid overfitting.
[0078] The present invention is equipped with a multi-scale feature interaction module, which can perform feature interaction and enable information interaction between features, which helps to a certain extent to the spatiotemporal correlation of features, and then helps to increase the attention of the entire detection network to the real change area and reduce the impact of pseudo-changes.
[0079] The present invention is provided with a change guidance unit, which helps to use the mask obtained from the deep change features to guide the multi-scale change information for feature fusion, further enhance the detection network's attention to the change area, and reduce the impact of pseudo-changes on the detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0081] Figure 1 This is a schematic block diagram of an embodiment of the remote sensing image change detection network of the present invention.
[0082] Figure 2This is a schematic block diagram of another embodiment of the remote sensing image change detection network of the present invention.
[0083] Figure 3 This is a schematic block diagram of an embodiment of the multi-scale feature interaction module of the present invention.
[0084] Figure 4 This is a schematic block diagram of an embodiment of the position attention module described in the present invention.
[0085] Figure 5 This is a schematic block diagram of another embodiment of the remote sensing image change detection network of the present invention.
[0086] Figure 6 The present invention is a schematic flow chart of an embodiment of the method.
[0087] Figure 7 The figure is a schematic block diagram of an embodiment of the apparatus of the present invention.
[0088] Figure 8 1 is a front-phase remote sensing image of a target dual-phase remote sensing image pair that has completed spatial registration.
[0089] Figure 9 for Figure 8 The front-phase remote sensing image shown corresponds to the back-phase remote sensing image in the target dual-temporal remote sensing image pair that has completed spatial registration.
[0090] Figure 10 For the general Figure 8 、 Figure 9 The target dual-temporal remote sensing image pair input for completing spatial registration is shown Figure 5 The figure shows a schematic diagram of image changes corresponding to the area to be detected calculated by the remote sensing image change detection network. DETAILED DESCRIPTION
[0091] In the detailed description that follows, various embodiments of the present disclosure will be described more fully.
[0092] It should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather that the present disclosure should be construed to cover all modifications, equivalents and / or alternatives falling within the spirit and scope of the various embodiments of the present disclosure.
[0093] The following is a brief introduction to some of the terms and technologies involved in the embodiments of the present invention:
[0094] Dual-temporal remote sensing image pair / image pair: It is two remote sensing images / images of the same geographical area at two different time points, consisting of a front-phase remote sensing image / image and a back-phase remote sensing image / image. The front-phase remote sensing image / image is taken earlier than the back-phase remote sensing image / image.
[0095] Spatial registration: refers to the alignment of remote sensing images / images of different phases in the same geographical area so that the images / images of different phases in the same geographical area are aligned.
[0096] SAM: Segment Anything Model, also known as Segment Anything Model, is an efficient and general image segmentation model that is used to automatically segment the input image and obtain a feature map of a pre-set scale.
[0097] VGG16: It is a deep convolutional neural network architecture, also known as Visual Geometry Group 16, used for image feature extraction.
[0098] like Figure 1 As shown, this embodiment provides a change detection network for remote sensing images of nature reserves, which includes: a backbone network, a feature extraction network, a first feature enhancement unit 300, a multi-scale feature interaction module (MFI), a second feature enhancement unit 400, a first change feature extraction module, a first mask extraction module, a change guidance unit 200, a fourth feature fusion module, a second change feature extraction unit 100, a position attention module, a third feature enhancement module, and a second mask extraction module. Specifically:
[0099] The backbone network is used to extract the input area to be detected for each dual-phase remote sensing image X1, X2 for remote sensing image change detection. t Four feature maps of different preset scales X1 is the remote sensing image of the previous phase, and X2 is the remote sensing image of the next phase. The dual-temporal remote sensing image X extracted by the backbone network t The i-th feature map of , i = 1, 2, 3, 4, t = 1, 2. X1 and X2 are two remote sensing images after spatial registration.
[0100] In this embodiment, the four different preset scales used to extract the feature map of image X1 are the same as the four different preset scales used to extract the feature map of image X2. and The same scale, and The same scale, and The same scale, and The same scale. In this embodiment, The scale decreases successively, t=1,2. For deep features, and is the middle-level feature, In specific implementation, those skilled in the art can set the above four different preset scales according to actual conditions.
[0101] The feature extraction network adopts SAM feature extractor, the model parameters of SAM feature extractor are SAM pre-trained weights, which are used to extract the input dual-phase remote sensing image X1, X2 for each dual-phase remote sensing image X1. t Feature map Each feature map All contain their corresponding remote sensing images X t Global information.
[0102] In the present invention, the model parameters of the feature extraction network adopt SAM pre-training weights. It can be seen that the model parameters of the feature extraction network in the present invention do not need to be trained and updated, which in turn helps to solve the problem of insufficient training data when using the SAM large model to build a model to a certain extent.
[0103] The first feature enhancement unit 300 is connected to the backbone network and the feature extraction network, and is used to extract the feature based on the feature extraction network. Extraction of backbone network Enhance and obtain the first enhanced feature map and the second enhanced feature map t = 1, 2. This design helps to enhance the feature extraction capability of the entire detection network and reduce the missed detection rate in change detection.
[0104] It can be understood that the first feature enhancement unit 300 is based on right Enhanced to get based on right Enhanced to get
[0105] The multi-scale feature interaction module is connected to the first feature enhancement unit 300 and is used to enhance the features obtained by the first feature enhancement unit 300. Perform feature interaction to obtain Enhanced feature map
[0106] Specifically, the multi-scale feature interaction module Perform feature interaction to obtain Enhanced feature map as well as Enhanced feature map Enhanced feature map Right now
[0107] The use of the multi-scale feature interaction module helps to learn the feature interaction patterns between different scales and phases, so that the output of the entire feature interaction module reflects the relationship between these phases and scales, which helps to increase the attention of the entire detection network to the truly changed areas in the area to be detected, thereby helping to alleviate the impact of pseudo-changes in the environment of the area to be detected.
[0108] The second feature enhancement unit 400 is connected to the backbone network, the first feature enhancement unit 300, and the multi-scale feature interaction module, and is used to obtain the feature information based on the multi-scale feature interaction module. right Perform feature enhancement to obtain t=1,2. This design helps to further enhance the feature extraction capability of the entire detection network.
[0109] The first change feature extraction module is connected to the multi-scale feature interaction module and the second feature enhancement unit 400 to extract and The changing characteristics of D i , i=1,2,3,4.
[0110] Specifically, the first change feature extraction module is used to extract and The change feature D1 is extracted and Extract the change feature D2 and The change feature D3, and the extraction and The changes in feature D4 are as follows:
[0111]
[0112] Where i = 1, 2, 3, 4, Concat() represents the concatenation of channel dimensions, CBR 3×3 () represents 3×3 convolution processing, specifically, CBR 3×3 ()=Relu(BatchNorm(Conv 3×3 ())),Conv 3×3 () represents 3×3 convolution, BatchNorm() represents batch normalization, and Relu() represents the Relu activation function.
[0113] The first mask extraction module is connected to the first change feature extraction module and is used to obtain a mask guidemap of the change feature D4. Optionally, the first mask extraction module obtains a mask guidemap of the change feature D4 as follows:
[0114] Guidemap=Sigmoid(Conv 3×3 (CBR 3×3 (D4))),
[0115] Among them, Conv 3×3 is 3×3 convolution, Sigmoid is the activation function, CBR 3×3 It is a 3×3 convolution process, which is a cascaded 3×3 convolution, batch normalization and Relu activation process. That is, D4 is first subjected to 3×3 convolution, followed by batch normalization, and finally Relu activation.
[0116] The change guidance unit 200 is connected to the first change feature extraction module and the first mask extraction module, and is used to obtain the change feature G3 of the change features D2, D3, and D4 after the mask Guidemap and the self-attention change guidance.
[0117] The fourth feature fusion module is connected to the change guidance unit 200 and the first change feature extraction module, and is used to fuse the features of G3 and D1 to obtain feature G′3. G′3 can enhance the change information.
[0118] The second change feature extraction unit 100 is connected to the feature extraction network and is used to extract the feature based on the feature extraction network. Get X1, X2 (ie {X t |t=1,2}) corresponding to the change feature D sam The feature extraction network passes X1 and X2 to unit 100.
[0119] The position attention module is connected to the second change feature extraction unit 100, and is used to sam Feature extraction is performed and the obtained feature map is recorded as the position attention weight map.
[0120] The third feature enhancement module is connected to the position attention module and the fourth feature fusion module, and is used to multiply the position attention weight map with G′3 to obtain the feature D defined . D defined To enhance the features of changing position information.
[0121] Specifically, D Refine =Ws×G′3, where Ws represents the position attention weight map.
[0122] The second mask extraction module is connected to the third feature enhancement module to extract feature D defined The mask of is obtained to obtain the change mask M, that is, the image change of the area to be detected is obtained.
[0123] The image change of the area to be detected mentioned in this specification is a mask, which is a binary change result image.
[0124] The foreground of the change mask M is the area where image changes occur in the area to be detected by the present invention.
[0125] When in use, first obtain the above-mentioned dual-phase remote sensing image pair X1 and X2, and then input the obtained X1 and X2 into the above-mentioned nature reserve remote sensing image change detection network for calculation, so as to obtain the image changes occurring in the corresponding area to be detected.
[0126] It can be understood that based on the change mask M, the area where the image change occurs in the area to be detected can be located.
[0127] Optionally, the first mask extraction module and the second mask extraction module both include:
[0128] The second convolution combination unit is used to reduce the number of channels of the input change feature to 1 / 2 of the original number (that is, the number of channels of the input change feature is halved), thereby obtaining a change feature with a reduced number of channels;
[0129] The fourth 3×3 convolutional layer is connected to the second convolution combination unit and is used to perform mask extraction on the change feature with reduced channel number to obtain a mask with a channel number of 1, which is a mask of the change feature of the input.
[0130] Optionally, the number of channels of the input change feature is reduced to 1 / 2 of the original number, including: first performing a 3×3 convolution on the input change feature, then performing normalization processing, and then performing activation processing using an activation function (such as a ReLU activation function).
[0131] It can be understood that when the second mask extraction module is used, its second convolution combination unit performs the input feature D defined First, perform 3×3 convolution, then perform normalization (such as batch normalization), and finally perform activation function activation to convert the input feature D defined The number of channels is halved to obtain a change feature with a reduced number of channels. Then, the fourth 3×3 convolutional layer performs mask extraction on the change feature with a reduced number of channels to obtain a change mask M with a channel number of 1.
[0132] It can be understood that when the first mask extraction module is used, its second convolution combination unit first performs 3×3 convolution on the input feature D4, then performs normalization processing, and finally performs activation function activation processing, thereby halving the number of channels of D4. After that, its fourth 3×3 convolution layer performs mask extraction on the D4 with halved channel number to obtain the above-mentioned mask map Guidemap.
[0133] The mask map Guidemap is used to guide the model of the above remote sensing image change detection network to focus on the changed area.
[0134] In an exemplary embodiment of the present invention, Figure 2 As shown, the first feature enhancement unit 300 includes: a first 3×3 convolutional layer, a second 3×3 convolutional layer, a first feature enhancement module and a second feature enhancement module.
[0135] The first 3×3 convolutional layer is connected to the feature extraction network to adjust The scale and number of channels are obtained by The same feature map
[0136] That is: the first 3×3 convolutional layer is used to adjust The scale and number of channels are obtained by The same feature map Used to adjust The scale and number of channels are obtained by The same feature map
[0137] The second 3×3 convolutional layer is connected to the first 3×3 convolutional layer to adjust The scale and number of channels are obtained by The same feature map
[0138] Understandably, when using, the second 3×3 convolutional layer adjusts The scale and number of channels are obtained by The same feature map Adjustment The scale and number of channels are obtained by The same feature map
[0139] The first feature enhancement module is connected to the first 3×3 convolutional layer and the backbone network to and Add to get the first enhanced feature map It is understandable that when used, the first feature enhancement module will and Add to get the first enhanced feature map Will and Add to get another first enhanced feature map
[0140] The second feature enhancement module is connected to the second 3×3 convolutional layer and the backbone network to and Add to get the second enhanced feature map It is understandable that when used, the second feature enhancement module will and Add to get the second enhanced feature map Will and Add to get another second enhanced feature map
[0141] For example, please refer to Figure 2 The multi-scale feature interaction module is connected to the second feature enhancement module to enhance the feature map Perform feature interaction to obtain Enhanced feature map
[0142] For example, Figure 3 As shown in the figure, the multi-scale feature interaction module consists of a channel stacking unit, four dilated convolutions with different expansion rates, a third channel splicing layer, a second 3×3 convolution unit and a feature interaction unit;
[0143] The channel stacking unit is connected to the second feature enhancement module and is the starting point of the multi-scale feature interaction module. Perform channel stacking;
[0144] Four dilated convolutions with different expansion rates are connected to the channel stacking units to extract features of four different scales from the channel stacking results;
[0145] The third channel splicing layer is connected to the four dilated convolutions with different expansion rates and is used to perform channel splicing on the convolution results of the four dilated convolutions to obtain the third splicing feature;
[0146] The second 3×3 convolution unit is connected to the third channel splicing layer, and is used to perform 3×3 convolution processing on the third splicing feature and record the processing result as the bi-temporal feature interaction weight w;
[0147] The feature interaction unit is connected to the second 3×3 convolution unit to combine the bi-temporal feature interaction weight w with the feature Multiply to obtain enhanced features after feature interaction
[0148] Specifically, the channel stacking unit pair To perform channel stacking:
[0149] Among them, f is the result of channel stacking, and Stack() represents the channel dimension. Channels are stacked together.
[0150] The four dilated convolutions with different expansion rates are: f′ h =DilatedConv 3×3 (f,d h ),
[0151] Among them, f′ h represents the convolution result of the hth dilated convolution in four dilated convolutions with different expansion rates, h = 1, 2, 3, 4; d h Represents the dilation rate of the h-th dilated convolution, d1 = 1, d2 = 3, d3 = 5, d4 = 7; DilatedConv 3×3 Represents a 3×3 dilated convolution. Dilated convolution is used to insert gaps between the elements of the convolution kernel to expand the receptive field. The dilation rate is the size of the gap between the convolution kernel elements.
[0152] The h-th dilated convolution is recorded as dilated convolution h, and there are four dilated convolutions with different dilation rates: dilated convolution 1, dilated convolution 2, dilated convolution 3, and dilated convolution 4. f′1, f′2, f′3, and f′4 are the convolution results of dilated convolution 1, dilated convolution 2, dilated convolution 3, and dilated convolution 4, respectively.
[0153] The convolution result f′ of each hole convolution h in the third channel splicing layer h Perform channel splicing to obtain the third splicing feature:
[0154] Concat(f′1,f′2,f′3,f′4), where Concat() represents the concatenation of the channel dimension.
[0155] The second 3×3 convolution unit performs 3×3 convolution on the third concatenated feature to obtain the bi-temporal feature interaction weight w:
[0156] w=Relu(BatchNorm(Conv 3×3 (Concat(f′1,f′2,f′3,f′4))),
[0157] Among them, Relu(BatchNorm(Conv 3×3 ())) represents 3×3 convolution processing, Conv 3×3represents 3×3 convolution, BatchNorm() represents batch normalization, and Relu represents the Relu activation function.
[0158] The feature interaction unit combines the above w with the feature Multiply to obtain enhanced features after feature interaction As shown below:
[0159] Where t=1,2.
[0160] That is: the feature interaction unit combines the bi-phase feature interaction weight w with the feature Multiply to obtain enhanced features after feature interaction The bi-temporal feature interaction weight w is combined with the feature Multiply to obtain enhanced features after feature interaction
[0161] For example, Figure 2 As shown, the second feature enhancement unit 400 includes a first upsampling module, a first feature fusion module, a second upsampling module, a second feature fusion module, a third upsampling module and a third feature fusion module.
[0162] The first upsampling module is connected to the multi-scale feature interaction module to Upsampling to get the scale and The same feature map Specifically, the first upsampling module is used to Upsampling to get the scale and The same feature map right Upsampling to get the scale and The same feature map
[0163] The first feature fusion module is connected to the first upsampling module and the first feature enhancement module, and is used to and Perform feature fusion to obtain Enhanced feature map That is, when used, the first feature fusion module can and Perform feature fusion to obtain Enhanced feature map Will and Perform feature fusion to obtain Enhanced feature map
[0164] The second upsampling module is connected to the first feature fusion module to Upsampling to get the scale and The same feature map That is, when used, the second upsampling module Upsampling to get the scale and The same feature map right Upsampling to get the scale and The same feature map
[0165] The second feature fusion module is connected to the second upsampling module and the backbone network to and Perform feature fusion to obtain Enhanced feature map It can be understood that when used, the second feature fusion module can and Perform feature fusion to obtain Enhanced feature map Will and Perform feature fusion to obtain Enhanced feature map
[0166] The third upsampling module is connected to the second feature fusion module to Upsampling is performed to obtain the scale and The same feature map It can be understood that when using, the third upsampling module Upsampling is performed to obtain the scale and The same feature map right Upsampling is performed to obtain the scale and The same feature map
[0167] The third feature fusion module is connected to the third upsampling module and the backbone network to and Perform feature fusion to obtain Enhanced feature map Specifically, the third feature fusion module is used to: and Perform feature fusion to obtain Enhanced feature map Will and Perform feature fusion to obtain Enhanced feature map
[0168] Alternatively, as Figure 2As shown, the change guiding unit 200 includes: a first change guiding module, a first channel splicing layer, a second change guiding module, a second channel splicing layer and a third change guiding module.
[0169] like Figure 2 As shown, the first change feature extraction module passes the extracted D1 to the fourth feature fusion module, passes the extracted D2 to the second channel splicing layer, passes the extracted D3 to the first channel splicing layer, and passes the extracted D4 to the first mask extraction module and the first change guidance module. Specifically:
[0170] The first change guidance module is connected to the first mask extraction module and the first change feature extraction module, and is used to calculate the change feature G1 of D4 after the change guidance of the mask Guidemap and self-attention.
[0171] The first channel splicing layer is connected to the first change guiding module and the first change feature extraction module, and is used to splice the feature G1 and the feature D3 to obtain the first splicing feature.
[0172] The second change guidance module is connected to the first mask extraction module and the first channel splicing layer, and is used to calculate the change feature G2 of the first splicing feature after the change guidance of the mask Guidemap and self-attention.
[0173] The second channel splicing layer is connected to the second change guiding module and the first change feature extraction module, and is used to splice the feature G2 and the feature D2 to obtain a second splicing feature.
[0174] The third change guidance module is connected to the first mask extraction module and the second channel splicing layer, and is used to calculate the change feature G3 of the second splicing feature after the change guidance of the mask Guidemap and self-attention.
[0175] The method for calculating the change feature of the target feature after the mask guidemap and the self-attention change guidance by the first change guidance module, the second change guidance module, and the third change guidance module includes:
[0176] Calculate the product of the target feature and the mask Guidemap to obtain the features of the enhanced change area;
[0177] Based on the features of the enhanced change area and the mask Guidemap, the self-attention of the enhanced change weight is obtained. The specific calculation expression is as follows:
[0178] Q=x×W q ×Guidemap,
[0179] K=x×W k ×Guidemap,
[0180] V=x×Wv ×Guidemap,
[0181]
[0182] Where W q 、W k 、W v is a learnable weight matrix, Q, K, V are the query vector Query, key vector Key and value vector Value of self-attention, softmax is the activation function, is the scaling factor, K T represents the transpose of K, x represents the features of the enhanced change area, and Attention(Q,K,V) represents the self-attention of the enhanced change weight;
[0183] The features of the enhanced change area are input into the self-attention of the enhanced change weight to obtain the change features after change guidance. The specific calculation expression is as follows:
[0184] y=x+Conv 3×3 (Attention(Q, K, V)),
[0185] Among them, Conv 3×3 represents a 3×3 convolution, and y represents the change characteristics after the change guidance.
[0186] It can be understood that the target feature is the feature of the input for change guidance. D4, the first splicing feature and the second splicing feature are each a target feature.
[0187] Specifically, the method for the first change guidance module to calculate the change feature G1 of D4 after the change guidance of the mask Guidemap and the self-attention includes:
[0188] Step 110: Calculate the product of D4 and the mask Guidemap to obtain the feature D4′ of the enhanced change area, that is, D4′=D4×Guidemap;
[0189] Step 120: Based on the D4′ and the mask Guidemap, obtain the self-attention with enhanced change weight, as shown below:
[0190] Q=D4′×W q ×Guidemap,
[0191] K=D4′×W k ×Guidemap,
[0192] V=D4′×W v ×Guidemap,
[0193]
[0194] Where W q 、W k 、W v is a learnable weight matrix, Q, K, V are the query vector Query, key vector Key and value vector Value of self-attention, softmax is the activation function, is the scaling factor, KT represents the transpose of K, and Attention(Q, K, V) is the self-attention that enhances the change weight;
[0195] Step 130: Input the enhanced change region feature D4′ into the self-attention of the enhanced change weight obtained in step 120 to obtain the change feature G1 after change guidance, as shown below:
[0196] G1=D4′+Conv 3×3 (Attention(Q, K, V)),
[0197] Among them, Conv 3×3 Represents a 3×3 convolution.
[0198] The method for the second change guidance module to calculate the change feature G2 of the first splicing feature after the mask Guidemap and the self-attention change guidance includes:
[0199] Step 1: Calculate the product of the first splicing feature and the mask Guidemap to obtain the feature D3′ of the enhanced change area, that is, D3′=F1×Guidemap, where F1 is the first splicing feature;
[0200] Step 2: Based on the feature D3′ of the enhanced change area and the mask Guidemap, the self-attention of the enhanced change weight is obtained, as shown below:
[0201] Q=D3′×W q ×Guidemap,
[0202] K=D3′×W k ×Guidemap,
[0203] V=D3′×W v ×Guidemap,
[0204]
[0205] Where W q 、W k 、W v is a learnable weight matrix, Q, K, V are the query vector Query, key vector Key and value vector Value of self-attention, softmax is the activation function, is the scaling factor, K T represents the transpose of K, and Attention(Q, K, V) is the self-attention that enhances the change weight;
[0206] Step 3: Input the feature D3′ of the enhanced change area into the self-attention of the enhanced change weight obtained in step 2 to obtain the change feature G2 after change guidance, as shown below:
[0207] G2=D3′+Conv 3×3 (Attention(Q, K, V)),
[0208] Among them, Conv 3×3 Represents a 3×3 convolution.
[0209] The method for the third change guidance module to calculate the change feature G3 of the second splicing feature after the mask Guidemap and the self-attention change guidance includes:
[0210] Step 1) Calculate the product of the second stitching feature and the mask Guidemap to obtain the feature D2′ of the enhanced change area, that is, D2′=F2×Guidemap, where F2 is the second stitching feature;
[0211] Step 2) Based on the feature D2′ of the enhanced change area and the mask Guidemap, the self-attention of the enhanced change weight is obtained, as shown below:
[0212] Q=D2′×W q ×Guidemap,
[0213] K=D2′×W k ×Guidemap,
[0214] V=D2′×W v ×Guidemap,
[0215]
[0216] Where W q 、W k 、W v is a learnable weight matrix, Q, K, V are the query vector Query, key vector Key and value vector Value of self-attention, softmax is the activation function, is the scaling factor, K T represents the transpose of K, and Attention(Q, K, V) is the self-attention that enhances the change weight;
[0217] Step 3) The feature D2′ of the enhanced change area is input into the self-attention of the enhanced change weight obtained in step 2) to obtain the change feature G3 after change guidance.
[0218] It can be understood that in the first change guidance module, G1 is y and D4′ is x. In the second change guidance module, G2 is y and D′3 is x. In the third change guidance module, G3 is y and D′3 is x.
[0219] Optionally, the method in which the first feature fusion module, the second feature fusion module, the third feature fusion module, and the fourth feature fusion module fuse two features to obtain a feature that enhances change information includes:
[0220] The two features to be fused are channel-spliced to obtain a spliced feature; the spliced feature is then convolved through a 3×3 convolution calculation unit to obtain a feature of enhanced change information corresponding to the two features to be fused; the 3×3 convolution calculation unit includes a cascaded 3×3 convolution, batch normalization, and ReLU activation function.
[0221] The first feature fusion module will and Perform feature fusion to obtain Enhanced feature map The methods include: and Perform channel splicing and then splice and The obtained splicing features are input into its 3×3 convolution calculation unit for feature fusion to obtain the feature
[0222] It can be understood that the first feature fusion module will and The feature fusion specifically includes: and Channel splicing is performed, and then the obtained splicing features are input into the 3×3 convolution calculation unit for convolution processing to obtain the feature right and Perform channel splicing, and then input the obtained splicing features into its 3×3 convolution calculation unit for convolution processing to obtain the feature
[0223] It is understandable that the second feature fusion module will and Perform feature fusion to obtain enhanced feature map Method: and Perform channel splicing and then input it into its 3×3 convolution calculation unit for convolution processing to obtain the feature t=1,2.
[0224] The third feature fusion module will and Perform feature fusion to obtain enhanced feature map Method: and Perform channel splicing, and then input the splicing result into its 3×3 convolution calculation unit for convolution processing to obtain the feature
[0225] The fourth feature fusion module obtains the feature G′3 of enhanced change information by performing channel concatenation on G3 and D1, and then inputting the convolution operation into a 3×3 convolution calculation unit to obtain the feature G′3, as follows:
[0226] G′3=CBR 3×3 (Concat(G3,D1)),
[0227] Concat means channel concatenation, CBR 3×3 Represents the convolution processing of the 3×3 convolution computing unit, that is, the cascaded 3×3 convolution, batch normalization, and ReLU activation function activation.
[0228] It can be understood that the first feature fusion module, the second feature fusion module, the third feature fusion module and the fourth feature fusion module each include a 3×3 convolution calculation unit.
[0229] Alternatively, as Figure 2 As shown, the second variation feature extraction unit 100 includes the second 3×3 convolution layer and the first 3×3 convolution layer, and also includes: a third 3×3 convolution layer and a second variation feature extraction module.
[0230] The third 3×3 convolutional layer is connected to the second 3×3 convolutional layer to adjust The scale and number of channels of G′3 are obtained to obtain the feature map with the same scale and number of channels as G′3
[0231] Specifically include: adjustment The scale and number of channels of G′3 are obtained to obtain the feature map with the same scale and number of channels as G′3 Adjustment The scale and number of channels of G′3 are obtained to obtain the feature map with the same scale and number of channels as G′3
[0232] The second change feature extraction module is connected to the third 3×3 convolutional layer to extract feature maps and feature maps The changing characteristics of D sam .
[0233] Optionally, the second change feature extraction module extracts and The changing characteristics of D sam :
[0234]
[0235] in, Express and Perform channel splicing, CBR 3×3 Express Perform 3×3 convolution processing, that is, First, perform 3×3 convolution, then normalize, and finally activate using the ReLU activation function.
[0236] In this embodiment, the position attention module is connected to the second change feature extraction module.
[0237] Alternatively, as Figure 4 As shown, the position attention module includes the first 3×3 convolution unit and the first 1×1 convolution unit.
[0238] The first 3×3 convolution unit is connected to the second change feature extraction module, and is used to extract the D sam Perform feature extraction and obtain the number of channels D sam The feature map with one quarter of the number of channels is recorded as the first feature map.
[0239] For example, D sam The number of channels is m, and the first 3×3 convolution unit is used to convert D sam The number of channels is reduced to m / 4.
[0240] The first 1×1 convolution unit is connected to the first 3×3 convolution unit and is used to adjust the number of channels of the first feature map to generate a feature map with the same number of channels as G′3, which is the position attention weight map.
[0241] The first 3×3 convolution unit includes a cascaded 3×3 convolution, batch normalization, and ReLU activation function to sam First, a 3×3 convolution is performed, followed by batch normalization, and then a ReLU activation process is performed to obtain the first feature map.
[0242] The 1×1 convolution combination includes a cascaded 1×1 convolution, batch normalization, and Relu activation function, which is used to first perform 1×1 convolution on the first feature map, then perform batch normalization, and finally perform Relu activation to obtain the position attention weight map.
[0243] As an illustrative embodiment of the present invention, the spatial registration of remote sensing images X1 and X2 includes: first performing atmospheric correction, geometric correction, and radiometric correction on the original images of X1 and X2, and then performing spatial registration.
[0244] As an exemplary embodiment of the present invention, the method for acquiring the remote sensing image change detection network includes:
[0245] Building a network model of the remote sensing image change detection network;
[0246] Construct a training set;
[0247] Initialize the built network model to obtain the initialized network model of the remote sensing image change detection network; in the initialized network model, the model parameters of each backbone network are initialized using the pre-trained weights of VGG16, and the model parameters of each SAM feature extractor are initialized using the pre-trained weights of SAM;
[0248] During training, the model parameters of each SAM feature extractor in the initialized network model are frozen, and then the data in the training set is input into the initialized network model for iterative training. During each iterative training process, the predicted output is obtained through forward propagation, and then the loss function is used to calculate the loss between the predicted output and its corresponding true label. Then, the network is back-propagated according to the calculated loss, and then the gradient descent algorithm is used to update the parameters in the network model during back-propagation until the loss function converges or the training reaches a preset number of iterations, thereby obtaining the remote sensing image change detection network.
[0249] It can be understood that building the network model of the remote sensing image change detection network includes: building the above-mentioned backbone network, building the above-mentioned feature extraction network, building the above-mentioned first feature enhancement unit 300, building the above-mentioned multi-scale feature interaction module, building the above-mentioned second feature enhancement unit 400, building the above-mentioned first change feature extraction module, building the above-mentioned first mask extraction module, building the above-mentioned change guidance unit 200, building the above-mentioned fourth feature fusion module, building the above-mentioned second change feature extraction unit 100, building the above-mentioned position attention module, building the above-mentioned third feature enhancement module, and building the above-mentioned second mask extraction module.
[0250] The pre-trained weights of VGG16 are provided by the VGG16 official website.
[0251] When training the network model, the present invention adopts a frozen parameter strategy to freeze the model parameters of each SAM feature extractor of the initialized remote sensing image change detection network, which helps to ensure that the parameters of the SAM feature extractor will not be updated during the iterative training process, thereby improving training efficiency and avoiding overfitting.
[0252] The SAM pre-training weights mentioned in this manual are all SAM pre-training weights provided by the SAM official website.
[0253] Exemplarily, the steps of constructing the training set include the following steps S1 to S4: Step S1: selecting two temporal remote sensing images of a nature reserve, downloading the data, and obtaining an original dual-temporal remote sensing image pair.
[0254] It can be understood that the nature reserve mentioned in step S1 can be any nature reserve.
[0255] In specific implementation, in order to increase the accuracy of remote sensing image change detection network detection, a training set can be constructed based on the remote sensing images of the nature reserve where the area to be detected is located to train a remote sensing image change detection network specifically for detecting the nature reserve where the area to be detected is located. Specifically, two phases of remote sensing images of the nature reserve where the area to be detected is located, collected by high-resolution satellites, can be selected and downloaded as the original dual-phase remote sensing image pair.
[0256] In specific implementation, the area to be detected can be specified by using latitude and longitude coordinates and place names.
[0257] Step S2: spatially registering the original bi-temporal remote sensing image pairs, and performing data cropping on the registered original bi-temporal remote sensing image pairs to obtain a plurality of bi-temporal remote sensing image pairs.
[0258] The original dual-temporal remote sensing image pair is spatially registered, including: performing atmospheric correction, geometric correction and radiation correction on the original dual-temporal remote sensing image pair to obtain the pre-processed original dual-temporal remote sensing image pair; and performing spatial registration on the pre-processed original dual-temporal remote sensing image pair.
[0259] Optionally, the data cropping step is 256 pixels, and the size of the remote sensing image obtained by data cropping is 256 pixels × 256 pixels.
[0260] Step S3: for each cropped dual-temporal remote sensing image pair, respectively mark the changed regions in the subsequent temporal image relative to the previous temporal image to obtain a change mask label corresponding to each cropped dual-temporal remote sensing image pair.
[0261] Step S4: taking each cropped dual-temporal remote sensing image pair and its corresponding change mask label as a sample, and collecting the samples to obtain the training set.
[0262] Preferably, data enhancement can be performed on the training set, and then the data-enhanced training set can be used to train the network model of the remote sensing image change detection network constructed above.
[0263] Perform data augmentation on the training set, including but not limited to: enhancing the images in the training set by flipping, rotating, translating, cropping, adjusting brightness and contrast, and performing random occlusion. Using image augmentation to enhance the training set data can provide richer training data for subsequent model training.
[0264] As a preference, Figure 5 As shown, the backbone network uses the first backbone network and the second backbone network. The first backbone network is used to extract the feature map of image X1 The second backbone network is used to extract the feature map of image X2 i = 1, 2, 3, 4. The first and second backbone networks both use the VGG16 network.
[0265] As a preference, Figure 5 As shown, the feature extraction network uses two SAM feature extractors, which are denoted as SAM1 and SAM2. The model parameters of the SAM feature extractors corresponding to SAM1 and SAM2 are all SAM pre-trained weights. SAM1 is used to extract the feature map of X1 SAM2 is used to extract the feature map of X2
[0266] Optionally, the first 3×3 convolutional layer is implemented using two convolutional layers, which are denoted as C1-1 and C1-2, as shown in Figure 5 As shown. C1-1 is used to adjust get C1-2 is used to adjust get The second 3×3 convolutional layer is implemented using two convolutional layers, which are denoted as C2-1 and C2-2. Figure 5 As shown. C2-1 is used to adjust get C2-2 is used to adjust get The third 3×3 convolutional layer is implemented using two convolutional layers, which are denoted as C3-1 and C3-2. C3-1 is used to adjust get C3-2 is used to adjust get like Figure 5 shown.
[0267] Optionally, C1-1, C1-2, C2-1, C2-2, C3-1 and C3-2 each include a 3×3 convolution (for feature scale and channel number adjustment), a batch normalization and an activation function in a cascade setting.
[0268] Optionally, the first feature enhancement module is implemented using two feature enhancement modules, which are denoted as TZ1-1 and TZ1-2. and Add together to get TZ1-2 is used to and Add together to get The second feature enhancement module is implemented using two feature enhancement modules, which are denoted as TZ2-1 and TZ2-2. and Add together to get TZ2-2 is used to and Add together to get like Figure 5 shown.
[0269] Optionally, the first upsampling module is implemented using two upsampling modules, which are respectively denoted as U1-1 and U1-2. U1-1 is used to Upsampling to get the scale and The same feature map U1-2 pairs Upsampling to get the scale and The same feature map The second upsampling module uses two upsampling modules, which are denoted as U2-1 and U2-2. U2-1 is used to Upsampling to get the scale and The same feature map U2-2 is used for Upsampling to get the scale and The same feature map The third up-sampling module uses two up-sampling modules, which are denoted as U3-1 and U3-2. U3-1 is used to Upsampling is performed to obtain the scale and The same feature map U3-2 is used for Upsampling is performed to obtain the scale and The same feature map like Figure 5 shown.
[0270] It is understandable that the first change feature extraction module can be implemented using one or more change feature extraction modules. Exemplarily, the first change feature extraction module is implemented using four change feature extraction modules. The four change feature extraction modules are denoted as aj, j = 1, 2, 3, 4, and aj represents the jth change feature extraction module among the four change feature extraction modules. Figure 5In this embodiment, a1 is used to extract and The change features D1 and a2 are used to extract and The change features D2 and a3 are used to extract and The change features D3 and a4 are used to extract and The change characteristics of D4.
[0271] Exemplarily, the second change feature extraction module is the same as each change feature extraction module aj, and both include a first channel splicing unit and a first convolution combination unit, wherein:
[0272] The first channel splicing unit is used to perform channel splicing on the two input features whose change features are to be obtained, and to input the spliced features into the first convolution combination unit;
[0273] The first convolution combination unit is used to first perform 3×3 convolution on the concatenated features of the input, then perform normalization processing, and then perform activation function activation to obtain the change features of the two features of the input whose change features are to be obtained.
[0274] In this embodiment, the expressions for obtaining D1, D2, D3, and D4 are:
[0275]
[0276] Among them, Concat represents channel splicing, CBR 3×3 Represents the convolution processing of the first convolutional combination unit (specifically: first perform 3×3 convolution on the input, then perform batch normalization, and then activate the activation function).
[0277] Optionally, the first feature fusion module is implemented using two feature fusion modules, which are denoted as R1-1 and R1-2. R1-1 is used to and Perform feature fusion to obtain R1-2 is used to and Perform feature fusion to obtain The second feature fusion module is implemented by two feature fusion modules, which are denoted as R2-1 and R2-2. and Perform feature fusion to obtain R2-2 is used to and Perform feature fusion to obtain The third feature fusion module is implemented by two feature fusion modules, which are denoted as R3-1 and R3-2. and Perform feature fusion to obtain Enhanced feature map R3-2 is used to and Perform feature fusion to obtain Enhanced feature map like Figure 5 shown.
[0278] It should be noted that in order to simplify the view structure, the present invention does not Figure 1 、 Figure 2 、 Figure 5 All t=1,2 are marked, but it is understandable that Figure 1 、 Figure 2 、 Figure 5 The values of each t shown in are 1 and 2, that is, t = 1, 2.
[0279] like Figure 6 As shown, the present invention provides a method for detecting changes in remote sensing images of nature reserves, the method comprising the following steps:
[0280] Collect a pair of dual-temporal remote sensing images of the area to be detected (the image pair is a pair of dual-temporal remote sensing images of the area to be detected used for remote sensing image change detection), which is recorded as a detection image pair;
[0281] Perform spatial registration on the detection image pair, and record the registered detection image pair as the target image pair;
[0282] The target image pair is input into the remote sensing image change detection network in any one of the above embodiments to perform calculations to obtain image changes in the area to be detected.
[0283] It is understandable that the area to be inspected may be the entire nature reserve to be inspected, or may be a designated area of the nature reserve to be inspected.
[0284] As can be understood, during image alignment, the pre-phase remote sensing image is the baseline remote sensing image, representing the baseline state of the corresponding area of the nature reserve for image change detection. The post-phase remote sensing image shows changes that have occurred in the area for image change detection since the baseline remote sensing image was captured. As can be understood, the resulting image changes in the target area represent the areas of change in the post-phase remote sensing image relative to the pre-phase remote sensing image.
[0285] The present invention also provides a device for detecting changes in remote sensing images of nature reserves, which includes a data acquisition device and an image change detection device. Figure 7 As shown. The data acquisition device is used to acquire a pair of dual-temporal remote sensing images of the area to be detected for remote sensing image change detection, thereby obtaining a target dual-temporal remote sensing image pair. The image change detection device integrates an image preprocessing unit and the remote sensing image change detection network of any of the above embodiments.
[0286] The image change detection device is connected to the data acquisition device, and is used to call the image preprocessing unit to perform spatial registration on the target dual-phase remote sensing image pair, and then input the target dual-phase remote sensing image pair that has completed spatial registration into the remote sensing image change detection network for calculation to obtain the image changes of the area to be detected.
[0287] During use, the data acquisition device collects a pair of dual-phase remote sensing images of the area to be detected for remote sensing image change detection to obtain a target dual-phase remote sensing image pair and transmits it to the image preprocessing unit of the image change detection device. The image preprocessing unit of the image change detection device performs spatial registration on the transmitted target dual-phase remote sensing image pair, and then inputs the spatially registered target dual-phase remote sensing image pair into the remote sensing image change detection network, which performs calculations to obtain the image changes of the area to be detected. For example, Figure 8 The front-phase remote sensing image of a target dual-phase remote sensing image pair that has completed spatial registration is shown as follows: Figure 9 for Figure 8 The example of the front-phase remote sensing image corresponds to the back-phase remote sensing image in the target dual-phase remote sensing image pair that has completed spatial registration. Figure 10 For the general Figure 8 、 Figure 9 The target dual-phase remote sensing image pair with completed spatial registration is input into the remote sensing image change detection network of the present invention to calculate the image change of the corresponding detection area. Figure 10 In , black represents the image background, white represents the image foreground, and the area corresponding to the image foreground is the image change detected in the area to be detected.
[0288] It can be understood that in the original acquired image corresponding to the later phase remote sensing image in the image pair input to the remote sensing image change detection network, the areas corresponding to the foreground of the image detected by the remote sensing image change detection network are marked. These areas are the locations where image changes occur in the corresponding areas to be detected based on the remote sensing image change detection network.
[0289] Preferably, the image change detection device provides a graphical user interface (GUI), through which a user can view the obtained image changes of the area to be detected.
[0290] Optionally, the graphical user interface (GUI) includes an API for users to upload remote sensing image change detection networks. The GUI includes an image change detection network display area for displaying all remote sensing image change detection networks uploaded to the device for user selection. Optionally, the image change detection device includes a network loading module for automatically loading a user-selected remote sensing image change detection network.
[0291] Optionally, the image change detection device also includes an integrated logging module for recording the work logs of each remote sensing image change detection network loaded within the device; the image change detection device also includes an integrated log storage module for storing the logs recorded by the logging module. A virtual log export button is integrated into the graphical user interface, allowing the user to export the logs stored in the log storage module. Optionally, the image change detection device is equipped with a data transmission unit for uploading the target dual-phase remote sensing image pair acquired by the data acquisition device and the image changes (i.e., change masks) corresponding to the area to be detected to a host computer.
[0292] Optionally, after receiving the target dual-phase remote sensing image pair and the image changes of its corresponding area to be detected uploaded by the image change detection device, the upper computer marks the area corresponding to the uploaded image foreground of the corresponding image change (i.e., the corresponding change mask) in the later-phase remote sensing image of the uploaded target dual-phase remote sensing image pair, thereby obtaining a remote sensing image with the location where the image change occurs in the corresponding area to be detected.
[0293] Optionally, the data transmission unit includes but is not limited to: a USB interface, a Wi-Fi interface, and an Ethernet interface. The data acquisition device is a remote sensing image acquisition device, such as a sensor capable of acquiring remote sensing images. The image change detection device is a computer.
[0294] It should be noted that the same or similar parts in this specification can be referenced to each other. Figure 1 、 Figure 2 The crossing lines in do not intersect. Figure 5 The cross lines with black dots at the center intersection intersect, and the cross lines without black dots do not intersect.
[0295] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A nature reserve remote sensing image change detection network, characterized in that: The detection network includes: Backbone network, used to extract the input area to be detected for remote sensing image change detection dual-phase remote sensing image pair 、 Each dual-temporal remote sensing image Four feature maps of different preset scales , 、 are remote sensing images of the before and after phases, represent No. feature maps, , 、 、 、 The scale decreases successively; 、 is an image pair after spatial registration; Feature extraction network, using SAM feature extractor, is used to extract the input image pair 、 Each dual-temporal remote sensing image Feature map containing global information ; The model parameters of the SAM feature extractor are the SAM pre-training weights; The first feature enhancement unit is connected to the backbone network and the feature extraction network, and is used to extract the feature based on the feature extraction network. Extraction of backbone network 、 Enhance and obtain the first enhanced feature map and the second enhanced feature map ; The multi-scale feature interaction module is connected to the first feature enhancement unit and is used to enhance the first feature enhancement unit 、 Perform feature interaction to obtain Enhanced feature map ; The second feature enhancement unit is connected to the backbone network, the first feature enhancement unit, and the multi-scale feature interaction module, and is used to obtain the feature information based on the multi-scale feature interaction module. ,right 、 、 Perform feature enhancement to obtain enhanced feature map 、 、 ; The first change feature extraction module is connected to the multi-scale feature interaction module and the second feature enhancement unit to extract and Changing characteristics , ; The first mask extraction module is connected to the first change feature extraction module and is used to obtain the change feature Mask Guidemap; The change guidance unit is connected to the first change feature extraction module and the first mask extraction module to obtain the change feature Change features after mask guidemap and self-attention changes ; The fourth feature fusion module is connected to the change guidance unit and the first change feature extraction module, and is used to and Perform feature fusion to obtain features ; The second change feature extraction unit is connected to the feature extraction network and is used to extract the , get 、 Corresponding change characteristics ; Position attention module, connected to the second change feature extraction unit, for Perform feature extraction, and the obtained feature map is recorded as the position attention weight map; The third feature enhancement module is connected to the position attention module and the fourth feature fusion module, and is used to combine the position attention weight map with the feature Multiply to get features ; The second mask extraction module is connected to the third feature enhancement module to extract features The change mask M of the remote sensing image of the area to be detected is obtained; in, ; The change guidance unit calculates the target feature in the mask And the methods of changing features after self-attention change guidance include: Calculate target features and masks The product of is used to obtain the features of the enhanced change area; Based on the features of the enhanced change region and the mask , get the self-attention with enhanced change weight; The features of the enhanced change region are input into the self-attention of the enhanced change weight to obtain the change features after change guidance.
2. The remote sensing image change detection network according to claim 1, characterized in that: The first feature enhancement unit includes: The first 3×3 convolutional layer is connected to the feature extraction network and is used to adjust the features extracted by the feature extraction network. The scale and number of channels are obtained by The same feature map ; The second 3×3 convolutional layer, connected to the first 3×3 convolutional layer, is used to adjust The scale and number of channels are obtained by The same feature map ; The first feature enhancement module is connected to the first 3×3 convolutional layer and the backbone network to and Add to get the first enhanced feature map ; The second feature enhancement module is connected to the second 3×3 convolutional layer and the backbone network to and Add to get the second enhanced feature map ; The second change feature extraction unit includes the second 3×3 convolutional layer and the first 3×3 convolutional layer; the second change feature extraction unit also includes: The third 3×3 convolutional layer, connected to the second 3×3 convolutional layer, is used to adjust The scale and number of channels are obtained by The same feature map ; The second change feature extraction module is connected to the third 3×3 convolution layer to extract feature maps With feature map The characteristics of the change are obtained ; The change guidance unit includes: The first change guiding module is connected to the first mask extraction module and the first change feature extraction module, and is used to calculate Change features after mask guidemap and self-attention change guidance ; The first channel splicing layer is connected to the first change guidance module and the first change feature extraction module for splicing features and features , get the first splicing feature; The second change guidance module is connected to the first mask extraction module and the first channel splicing layer, and is used to calculate the change characteristics of the first splicing feature after the mask guidemap and self-attention change guidance ; The second channel splicing layer is connected to the second change guidance module and the first change feature extraction module to splice features and features , get the second splicing feature; The third change guidance module is connected to the first mask extraction module and the second channel splicing layer, and is used to calculate the change characteristics of the second splicing feature after the mask Guidemap and the self-attention change guidance, that is, to obtain the ; The second feature enhancement unit includes: The first upsampling module is connected to the multi-scale feature interaction module to Upsampling to get the scale and The same feature map ; The first feature fusion module is connected to the first upsampling module and the first feature enhancement module, and is used to and Perform feature fusion to obtain the enhanced feature map ; The second upsampling module is connected to the first feature fusion module and is used to Upsampling to get the scale and The same feature map ; The second feature fusion module is connected to the second upsampling module and the backbone network to and Perform feature fusion to obtain the enhanced feature map ; The third upsampling module is connected to the second feature fusion module to Upsampling is performed to obtain the scale and The same feature map ; The third feature fusion module is connected to the third upsampling module and the backbone network to and Perform feature fusion to obtain the enhanced feature map ; in, .
3. The remote sensing image change detection network according to claim 2, characterized in that: The first mask extraction module and the second mask extraction module both include: The second convolution combination unit is used to reduce the number of channels of the input change feature to 1 / 2 of the original number, thereby obtaining a change feature with reduced channel number; The fourth 3×3 convolutional layer is connected to the second convolution combination unit and is used to perform mask extraction on the change feature with reduced channel number to obtain a mask with a channel number of 1, which is a mask of the change feature of the input.
4. The remote sensing image change detection network according to claim 1, characterized in that: The position attention module includes the first 3×3 convolution unit and the first 1×1 convolution unit; The first 3×3 convolution unit is connected to the second change feature extraction module, and is used to extract the Perform feature extraction and get the number of channels as The feature map with one quarter of the number of channels is recorded as the first feature map; The first 1×1 convolution unit is connected to the first 3×3 convolution unit to adjust the number of channels of the first feature map and generate the number of channels and The same feature map, which is the position attention weight map.
5. The remote sensing image change detection network according to claim 1, characterized in that: The multi-scale feature interaction module consists of a channel stacking unit, four dilated convolutions with different expansion rates, a third channel splicing layer, a second 3×3 convolution unit, and a feature interaction unit; The channel stacking unit is connected to the first feature enhancement unit and is used to 、 Perform channel stacking; Four dilated convolutions with different expansion rates are connected to the channel stacking units to extract features of four different scales from the channel stacking results; The third channel splicing layer is connected to four dilated convolutions with different expansion rates. It is used to perform channel splicing on the convolution results of the four dilated convolutions to obtain the third splicing feature. The second 3×3 convolution unit is connected to the third channel splicing layer, which is used to perform 3×3 convolution processing on the third splicing feature and record the processing result as the bi-temporal feature interaction weight. ; The feature interaction unit is connected to the second 3×3 convolution unit to convert the bi-phase feature interaction weights and features Multiply to obtain enhanced features after feature interaction , t=1,2.
6. The remote sensing image change detection network according to any one of claims 1 to 5, characterized in that: The method for acquiring the remote sensing image change detection network includes: Building a network model of the remote sensing image change detection network; Construct a training set; Initialize the built network model to obtain the initialized network model of the remote sensing image change detection network; in the initialized network model, the model parameters of each backbone network are initialized using the pre-trained weights of VGG16, and the model parameters of each SAM feature extractor are initialized using the pre-trained weights of SAM; During training, the model parameters of each SAM feature extractor in the initialized network model are frozen, and then the data in the training set is input into the initialized network model for iterative training. During each iterative training process, the predicted output is obtained through forward propagation, and then the loss function is used to calculate the loss between the predicted output and its corresponding true label. Then, the network is back-propagated according to the calculated loss, and then the gradient descent algorithm is used to update the parameters in the network model during back-propagation until the loss function converges or the training reaches a preset number of iterations, thereby obtaining the remote sensing image change detection network.
7. The remote sensing image change detection network according to claim 6, characterized in that: The training set construction methods include: Select two temporal remote sensing images of the nature reserve, download the data, and obtain the original dual-temporal remote sensing image pair; Perform spatial registration on the original dual-temporal remote sensing image pairs, and perform data cropping on the registered original dual-temporal remote sensing image pairs to obtain several dual-temporal remote sensing image pairs; For each cropped dual-temporal remote sensing image pair: mark the changed areas in the latter-phase image relative to the former-phase image to obtain the change mask labels; Each cropped dual-temporal remote sensing image pair and its corresponding change mask label is taken as a sample, and the samples are collected to obtain the training set.
8. A method for detecting changes in remote sensing images of nature reserves, characterized in that: Methods include: Collect a pair of dual-temporal remote sensing images of the area to be detected for remote sensing image change detection, which is recorded as a detection image pair; Perform spatial registration on the detection image pair to obtain the registered detection image pair, which is recorded as the target image pair; The target image pair is input into the remote sensing image change detection network according to any one of claims 1 to 7 to perform calculations to obtain image changes of the area to be detected in the nature reserve to be detected.
9. A device for detecting changes in remote sensing images of nature reserves, characterized in that: The device includes a data acquisition device and an image change detection device; The data acquisition device is used to acquire a pair of dual-temporal remote sensing images of the area to be detected for remote sensing image change detection, thereby obtaining a target dual-temporal remote sensing image pair; An image change detection device, which integrates an image preprocessing unit and a remote sensing image change detection network according to any one of claims 1 to 7; The image change detection device is connected to the data acquisition device, and is used to call the image preprocessing unit to perform spatial registration on the target dual-phase remote sensing image pair, and then input the target dual-phase remote sensing image pair that has completed spatial registration into the remote sensing image change detection network for calculation to obtain the image changes of the area to be detected.
Citation Information
Patent Citations
Remote sensing image detection method based on dual-time-phase interaction enhancement CNN-Transform
CN118135392A
Method for extracting building change area in double-time-phase remote sensing image based on twinborn mixed attention mechanism and multi-scale feature fusion
CN118212532A