High-resolution remote sensing image change detection method, electronic device and storage medium based on multi-scale convolutional decoding

By using multi-scale convolutional decoding method in high-resolution remote sensing image change detection, MambaVision encoder and multi-scale convolutional decoder are built, which solves the problems of limited spatial context capture and high computational cost in the prior art, and achieves more accurate change detection and more efficient computing performance.

CN119418216BActive Publication Date: 2025-05-13HARBIN AEROSPACE STAR DATA SYST TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411394160.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-05-13
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

The prior art has problems such as limited spatial context capture and high computational cost in high resolution remote sensing image change detection.

Method used

A high-resolution remote sensing image change detection method based on multi-scale convolutional decoding is proposed. By constructing MambaVision encoder, time feature interaction module TFIM and multi-scale convolutional decoder EMCAD, it can effectively capture global spatial context information, reduce the computational burden, and improve the training and inference efficiency of the model.

Benefits of technology

This method can capture changing areas more accurately, reduce false detection of pseudo-change changes, and improve the ability to detect slight changes in high-resolution images. It is suitable for change detection tasks of complex terrain or subtle targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418216B_ABST
    Figure CN119418216B_ABST
Patent Text Reader

Abstract

A high-resolution remote sensing image change detection method, electronic device and storage medium based on multi-scale convolution decoding belong to the technical field of remote sensing image change detection. To meet the needs of high-resolution remote sensing image change detection, the present invention includes constructing a remote sensing image data set; constructing a MambaVision encoder, the first and second stages are based on CNN modules, the third and fourth stages are based on MambaVision modules and Transformer modules, and four dual-phase multi-scale feature maps are obtained through four-stage processing; constructing a time feature interaction module TFIM; constructing a multi-scale convolution decoder EMCAD; constructing a loss function, using the constructed remote sensing image data set for training, testing and evaluation, and applying the obtained trained optimal model to high-resolution remote sensing change detection. The present invention improves the ability to detect small changes in high-resolution images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image change detection, and in particular relates to a high-resolution remote sensing image change detection method based on multi-scale convolution decoding, an electronic device and a storage medium. Background Art

[0002] With the continuous development of remote sensing technology, change detection in remote sensing images has become a hot research field in the remote sensing community. The goal of this task is to monitor surface changes in the same area using remote sensing images acquired at different times. Change detection plays an important role in various fields such as urban planning, land cover analysis, disaster assessment, ecosystem monitoring, and resource management. Optical high-resolution remote sensing images are widely used in the field of change detection because they can provide detailed features such as rich texture and geometric structure information. At present, CNN and Transformer have made remarkable progress in the field of remote sensing change detection. However, both architectures have their inherent shortcomings: CNN is limited by a limited receptive field, which prevents them from capturing a wider range of spatial contexts, while Transformer is computationally intensive, which makes them expensive to train and deploy on large datasets. In recent years, the Mamba architecture based on the state-space model has shown remarkable performance in a series of natural language processing tasks, which can effectively make up for the shortcomings of the above two architectures. At present, there are also many studies that introduce Mamba into the field of vision and achieve good results. Summary of the invention

[0003] The problem to be solved by the present invention is to meet the demand for high-resolution remote sensing image change detection, and propose a high-resolution remote sensing image change detection method, electronic equipment and storage medium based on multi-scale convolution decoding.

[0004] To achieve the above object, the present invention is implemented through the following technical solutions:

[0005] A high-resolution remote sensing image change detection method based on multi-scale convolution decoding comprises the following steps:

[0006] S1. Collect remote sensing images through sensors carried by satellites or drones, perform preprocessing, and then construct a remote sensing image dataset;

[0007] S2. Construct a MambaVision encoder, which consists of four different stages. The first and second stages are based on CNN modules, and the third and fourth stages are based on MambaVision modules and Transformer modules. The data in the remote sensing image dataset obtained in step S1 is input into the MambaVision encoder, and four dual-phase multi-scale feature maps are obtained through four-stage processing.

[0008] S3. Construct a time feature interaction module TFIM, input the dual-phase multi-scale feature map obtained in step S2 into the time feature interaction module TFIM, and obtain a feature fusion feature map;

[0009] S4. Construct a multi-scale convolution decoder EMCAD, including using four multi-scale convolution attention modules MSCAM to optimize the feature fusion feature map obtained in step S3. After each MSCAM, use the segmentation head SH to generate the segmentation map of this stage, upsample the optimized feature fusion feature map through the efficient upconvolution block EUCB, and add the upsampled feature map to the corresponding group attention gate LGAG output. Finally, add the four different segmentation maps to generate the final prediction map;

[0010] S5. Construct a loss function, use the remote sensing image dataset constructed in step S1 to train, test, and evaluate the MambaVision encoder obtained in step S2, the temporal feature interaction module TFIM obtained in step S3, and the multi-scale convolutional decoder EMCAD obtained in step S4, and apply the trained optimal model to high-resolution remote sensing change detection.

[0011] Furthermore, the preprocessing method of step S1 includes geometric correction, cloud and shadow removal, clutter cleaning, image enhancement, scale normalization and noise suppression;

[0012] The remote sensing image dataset is divided into a training set, a validation set, and a test set according to a data ratio of 6:2:2.

[0013] Furthermore, the specific implementation method of step S2 includes the following steps:

[0014] S2.1. Set the first and second stages based on CNN modules, including input images of size H×W×3, H is the height, W is the width, and the input is first converted to size The layers are overlapped and projected into a C-dimensional embedding space by two consecutive 3×3 convolutional layers with a stride of 2. The downsampler between stages consists of a batch normalized 3×3 convolutional layer with a stride of 2 to halve the image resolution.

[0015] The CNN module follows the following general residual block formula:

[0016]

[0017] in, is the partial output after the residual block, z is the input, GELU is the Gaussian error linear unit activation function, BN is the batch normalization function, Conv 3×3(·) denotes a 3×3 convolutional layer with batch normalization and ReLU activation function;

[0018] S2.3. Set the third and fourth stages to be based on the MambaVision module and the Transformer module;

[0019] S2.3.1. Set the given input to X1 and the output of the MambaVision module to X out Calculated according to the following formula:

[0020]

[0021] Among them, Linear(C in ,C out ) represents a linear layer, where C in and C out are the embedding dimensions of input and output respectively, Scan is the selective scanning operation, σ is the Sigmoid Linear activation function, and Concat is the concatenation operation;

[0022] S2.3.2. Set the expression of the multi-head self-attention mechanism Transformer module to:

[0023]

[0024] Where Q, K, and V represent query, key, and value, respectively. h represents the number of attention heads, and Softmax is the activation function;

[0025] S2.3.3. Setting input Where the sequence length is T, the calculation formula for the output of the nth layer in the third stage and the output of the nth layer in the fourth stage is:

[0026]

[0027] in, is the output after being processed by the MambaVision module, X n is the output of the multi-head self-attention mechanism, X n-1 is the input, Norm and Mixer represent the selection of layer normalization and token mixing modules respectively;

[0028] Set the given N layers, before The MambaVision module is used in the layer, and then The layer adopts a multi-head self-attention mechanism;

[0029] S2.4. The dual-phase multi-scale feature map F is obtained through four-stage processing. i, i=1,2,3,4.

[0030] Furthermore, the specific implementation method of step S3 includes the following steps:

[0031] S3.1. For the i-th bi-phase multi-scale feature map obtained in step S2, the difference feature map is calculated by element-by-element subtraction and then absolute value operation. The expression is:

[0032]

[0033] in, is the first phase characteristic diagram, is the characteristic diagram of the second phase, |·| represents the absolute value operation, Represents element-by-element subtraction operation, Conv 3×3 (·) denotes a 3×3 convolutional layer with batch normalization and ReLU activation function;

[0034] S3.2. Applied to the time feature, the change area in the time feature is highlighted by element-by-element multiplication operation, and the enhanced time feature map is added to the original feature map. The expression is:

[0035]

[0036] in, is the refined first phase feature, For the refined second phase characteristics, and ⊕ represent element-wise multiplication and addition operations, respectively;

[0037] S3.3. Through the splicing of refined dual-phase features, a 3×3 convolutional layer is used to mine the time difference information, and the number of channels of the spliced ​​feature is halved. Then, the time difference representation is added to the previous coarse time difference feature. Finally, a 1×1 convolutional layer is used to reduce the number of channels. The expression is:

[0038]

[0039] Among them, D i To obtain the temporal difference features, Cat(·) represents the feature concatenation operation;

[0040] S3.4. Apply the methods of steps S3.1 to S3.3 to the four dual-phase multi-scale feature maps to generate four multi-feature fused feature maps.

[0041] Furthermore, the specific implementation method of step S4 includes the following steps:

[0042] S4.1. Construct the group attention gate LGAG, specifically set it at q att (.) function, first by applying independent 3×3 grouped convolution GC g (.) and GC x (.) processes the 1×1 convolution gating signal g and the input feature map x, and the convolved features are then normalized by batch normalization BN(.) and merged by element-by-element addition; the generated feature map is activated by a ReLU(R(.)) layer, after which a 1×1 convolution C(.) is applied followed by a BN(.) layer to obtain a single-channel feature map; then, the obtained single-channel feature map is activated by the Sigmoidσ(.) activation function to generate the attention coefficient; the output of this transformation is used to scale the input feature x by element-by-element multiplication to generate the attention gate feature LGAG(g,x), expressed as:

[0043] q att (g,x)=R(BN(GC g (g)+BN(GC x (x)))))

[0044]

[0045] S4.2. Construct a multi-scale convolutional attention module MSCAM, which consists of a channel attention block CAB, a spatial attention block SAB, and a multi-scale convolution block MSCB, and the expression is:

[0046] MSCAM(x)=MSCB(SAB(CAB(x)))

[0047] Among them, x represents the input tensor;

[0048] S4.2.1. Construct a multi-scale convolution block MSCB. First, expand the number of channels by a point-by-point convolution layer PWC1(·) with an expansion factor of 2, and then perform batch normalization BN(·) and ReLU6 activation function R6(.); then, capture multi-scale and multi-resolution contextual information through multi-scale depth convolution MSDC(.); introduce channel shuffling operation to enhance the correlation between channels; then, use another point-by-point convolution PWC2(.) and combine BN(.) to restore the number of channels to the original number, and encode the dependency between channels, expressed as:

[0049] MSCB(x)=BN(PWC2(CS(MSDC(R6(BN(PWC1(x)))))))

[0050] MSDC(x)=∑ ks∈KS DWCB ks (x)

[0051] x=x+DWCB ks (x);

[0052] S4.2.2. Construct a channel attention block CAB. First, apply adaptive maximum pooling Pm(·) and adaptive average pooling Pa(·) to the spatial dimension to extract the most significant features of the entire feature map in each channel. Then, for each pooled feature map, reduce the number of channels to r = 1 / 16 of the original through point-by-point convolution C1(·), and then perform ReLU activation R. After that, use another point-by-point convolution C2(·) to restore the original number of channels. Next, add the two restored feature maps and apply the Sigmoid (σ) activation function to estimate the attention weight. Finally, use the Hadamard product Incorporating the weights into the input x, the expression is:

[0053]

[0054] S4.2.3. Construct a spatial attention block SAB. First, perform maximum pooling Chmax(·) and average pooling Chavg(·) along the channel dimension to focus on local features. Then, use a 7×7 large kernel convolution layer to enhance the local contextual relationship between features. Then, apply the Sigmoid activation function (σ) to calculate the attention weight. Finally, use the Hadamard product Apply the weights to the input x, expressed as:

[0055]

[0056] S4.3. Construct an efficient upconvolution block EUCB. First, use UpSampling Up(·) to perform upsampling by a factor of 2. Then, by applying a 3×3 depth convolution DWC(·), combined with batch normalization BN(·) and ReLU(.) activation function, the upsampled feature map is enhanced; finally, a 1×1 convolution C1×1(.) is used to reduce the number of channels to match the number of the next stage, expressed as:

[0057] EUCB(x)=C 1×1 (ReLU(BN(DWC(Up(x))))).

[0058] Furthermore, the specific implementation method of step S5 is to generate four prediction graphs p1, p2, p3 and p4 at each stage based on the four segmentation heads of the multi-scale convolutional decoder EMCAD, calculate the loss of all possible prediction combinations obtained from the four segmentation heads, minimize this cumulative combination loss during the training process, and obtain the loss function The expression is:

[0059]

[0060] in, and are the losses of each prediction graph, α=β=γ=ζ=δ=1.0 is the weight assigned to each loss, is an additional loss item;

[0061] The prediction map p1 at the final stage of the decoder is regarded as the final segmentation map, and then the final segmentation output is obtained by performing binary classification segmentation using the Sigmoid function.

[0062] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a high-resolution remote sensing image change detection method based on multi-scale convolution decoding when executing the computer program.

[0063] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a high-resolution remote sensing image change detection method based on multi-scale convolution decoding.

[0064] Beneficial effects of the present invention:

[0065] The present invention discloses a method for detecting changes in high-resolution remote sensing images based on multi-scale convolutional decoding, constructs a MambaVision encoder, and can effectively capture global spatial context information in high-resolution remote sensing images, while reducing the computational burden on large-scale data sets and improving the training and reasoning efficiency of the model. By utilizing the temporal feature interaction module TFIM, the complementarity and dependency between image features at different time points are enhanced, so that the model can more accurately capture the changed area and reduce false detection of pseudo-changes. At the same time, the constructed multi-scale convolutional decoder EMCAD effectively utilizes local detail information and improves the ability to detect small changes in high-resolution images, and is particularly suitable for change detection tasks for complex terrain or subtle targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 A flowchart of a high-resolution remote sensing image change detection method based on multi-scale convolution decoding according to the present invention;

[0067] Figure 2 It is a schematic diagram of the principle flow of the present invention;

[0068] Figure 3 This is the schematic diagram of the MambaVision encoder;

[0069] Figure 4 This is a schematic diagram of the MambaVision module architecture;

[0070] Figure 5 It is a schematic diagram of the temporal feature interaction module TFIM;

[0071] Figure 6 It is the schematic diagram of the multi-scale convolutional decoder EMCAD. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0073] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0074] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples, and the attached Figure 1 -Attached Figure 6 The detailed instructions are as follows:

[0075] Embodiment 1:

[0076] A high-resolution remote sensing image change detection method based on multi-scale convolution decoding comprises the following steps:

[0077] S1. Collect remote sensing images through sensors carried by satellites or drones, perform preprocessing, and then construct a remote sensing image dataset;

[0078] Furthermore, the preprocessing method of step S1 includes geometric correction, cloud and shadow removal, clutter cleaning, image enhancement, scale normalization and noise suppression;

[0079] The remote sensing image data set is divided into a training set, a validation set, and a test set according to a data ratio of 6:2:2;

[0080] S2. Construct a MambaVision encoder, which consists of four different stages. The first and second stages are based on CNN modules, and the third and fourth stages are based on MambaVision modules and Transformer modules. The data in the remote sensing image dataset obtained in step S1 is input into the MambaVision encoder, and four dual-phase multi-scale feature maps are obtained through four-stage processing.

[0081] Furthermore, the specific implementation method of step S2 includes the following steps:

[0082] S2.1. Set the first and second stages based on CNN modules, including input images of size H×W×3, H is the height, W is the width, and the input is first converted to size The layers are overlapped and projected into a C-dimensional embedding space by two consecutive 3×3 convolutional layers with a stride of 2. The downsampler between stages consists of a batch normalized 3×3 convolutional layer with a stride of 2 to halve the image resolution.

[0083] The expression of the CNN module is:

[0084]

[0085] in, is the partial output after the residual block, z is the input, GELU is the Gaussian error linear unit activation function, BN is the batch normalization function, Conv 3×3 (·) denotes a 3×3 convolutional layer with batch normalization and ReLU activation function;

[0086] S2.3. Set the third and fourth stages to be based on the MambaVision module and the Transformer module;

[0087] S2.3.1. Set the given input to X1 and the output of the MambaVision module to X out Calculated according to the following formula:

[0088]

[0089] Among them, Linear(C in ,C out ) represents a linear layer, where C in and C out are the embedding dimensions of input and output respectively,

[0090] Scan is a selective scanning operation, σ is a Sigmoid Linear activation function, and Concat is a concatenation operation;

[0091] S2.3.2. Set the expression of the multi-head self-attention mechanism Transformer module to:

[0092]

[0093] Where Q, K, and V represent query, key, and value, respectively. h represents the number of attention heads, and Softmax is the activation function;

[0094] S2.3.3. Setting input Where the sequence length is T, the calculation formula for the output of the nth layer in the third stage and the output of the nth layer in the fourth stage is:

[0095]

[0096] in, is the output after being processed by the MambaVision module, X n is the output of the multi-head self-attention mechanism, X n-1 is the input, Norm and Mixer represent the selection of layer normalization and token mixing modules respectively;

[0097] Set the given N layers, before The MambaVision module is used in the layer, and then The layer adopts a multi-head self-attention mechanism;

[0098] S2.4. The dual-phase multi-scale feature map F is obtained through four-stage processing. i , i=1,2,3,4;

[0099] S3. Construct a time feature interaction module TFIM, input the dual-phase multi-scale feature map obtained in step S2 into the time feature interaction module TFIM, and obtain a feature fusion feature map;

[0100] Furthermore, the specific implementation method of step S3 includes the following steps:

[0101] S3.1. For the i-th bi-phase multi-scale feature map obtained in step S2, the difference feature map is calculated by element-by-element subtraction and then absolute value operation. The expression is:

[0102]

[0103] in, is the first phase characteristic diagram, is the characteristic diagram of the second phase, |·| represents the absolute value operation, Represents element-by-element subtraction operation, Conv3×3 (·) denotes a 3×3 convolutional layer with batch normalization and ReLU activation function;

[0104] S3.2. Applied to the time feature, the change area in the time feature is highlighted by element-by-element multiplication operation, and the enhanced time feature map is added to the original feature map. The expression is:

[0105]

[0106] in, is the refined first phase feature, For the refined second phase characteristics, and ⊕ represent element-wise multiplication and addition operations, respectively;

[0107] S3.3. Through the splicing of refined dual-phase features, a 3×3 convolutional layer is used to mine the time difference information, and the number of channels of the spliced ​​feature is halved. Then, the time difference representation is added to the previous coarse time difference feature. Finally, a 1×1 convolutional layer is used to reduce the number of channels. The expression is:

[0108]

[0109] Among them, D i To obtain the temporal difference features, Cat(·) represents the feature concatenation operation;

[0110] S3.4. Apply the methods of steps S3.1 to S3.3 to the four dual-phase multi-scale feature maps to generate four multi-feature fused feature maps.

[0111] S4. Construct a multi-scale convolution decoder EMCAD, including using four multi-scale convolution attention modules MSCAM to optimize the feature fusion feature map obtained in step S3. After each MSCAM, use the segmentation head SH to generate the segmentation map of this stage, upsample the optimized feature fusion feature map through the efficient upconvolution block EUCB, and add the upsampled feature map to the corresponding group attention gate LGAG output. Finally, add the four different segmentation maps to generate the final prediction map;

[0112] Furthermore, the specific implementation method of step S4 includes the following steps:

[0113] S4.1. Construct the group attention gate LGAG, specifically set it at q att (.) function, first by applying independent 3×3 grouped convolution GC g (.) and GC x(.) processes the 1×1 convolution gating signal g and the input feature map x, and the convolved features are then normalized by batch normalization BN(.) and merged by element-by-element addition; the generated feature map is activated by a ReLU(R(.)) layer, after which a 1×1 convolution C(.) is applied followed by a BN(.) layer to obtain a single-channel feature map; then, the obtained single-channel feature map is activated by the Sigmoidσ(.) activation function to generate the attention coefficient; the output of this transformation is used to scale the input feature x by element-by-element multiplication to generate the attention gate feature LGAG(g,x), expressed as:

[0114] q att (g,x)=R(BN(GC g (g)+BN(GC x (x)))))

[0115]

[0116] S4.2. Construct a multi-scale convolutional attention module MSCAM, which consists of a channel attention block CAB, a spatial attention block SAB, and a multi-scale convolution block MSCB, and the expression is:

[0117] MSCAM(x)=MSCB(SAB(CAB(x)))

[0118] Among them, x represents the input tensor;

[0119] S4.2.1. Construct a multi-scale convolution block MSCB. First, expand the number of channels by a point-by-point convolution layer PWC1(·) with an expansion factor of 2, and then perform batch normalization BN(·) and ReLU6 activation function R6(.); then, capture multi-scale and multi-resolution contextual information through multi-scale depth convolution MSDC(.); introduce channel shuffling operation to enhance the correlation between channels; then, use another point-by-point convolution PWC2(.) and combine BN(.) to restore the number of channels to the original number, and encode the dependency between channels, expressed as:

[0120] MSCB(x)=BN(PWC2(CS(MSDC(R6(BN(PWC1(x)))))))

[0121] MSDC(x)=∑ ks∈KS DWCB ks (x)

[0122] x=x+DWCB ks (x);

[0123] S4.2.2. Construct a channel attention block CAB. First, apply adaptive maximum pooling Pm(·) and adaptive average pooling Pa(·) to the spatial dimension to extract the most significant features of the entire feature map in each channel. Then, for each pooled feature map, reduce the number of channels to r = 1 / 16 of the original through point-by-point convolution C1(·), and then perform ReLU activation R. After that, use another point-by-point convolution C2(·) to restore the original number of channels. Next, add the two restored feature maps and apply the Sigmoid (σ) activation function to estimate the attention weight. Finally, use the Hadamard product Incorporating the weights into the input x, the expression is:

[0124]

[0125] S4.2.3. Construct a spatial attention block SAB. First, perform maximum pooling Chmax(·) and average pooling Chavg(·) along the channel dimension to focus on local features. Then, use a 7×7 large kernel convolution layer to enhance the local contextual relationship between features. Then, apply the Sigmoid activation function (σ) to calculate the attention weight. Finally, use the Hadamard product Apply the weights to the input x, expressed as:

[0126]

[0127] S4.3. Construct an efficient upconvolution block EUCB. First, use UpSampling Up(·) to perform upsampling by a factor of 2. Then, by applying a 3×3 depth convolution DWC(·), combined with batch normalization BN(·) and ReLU(.) activation function, the upsampled feature map is enhanced; finally, a 1×1 convolution C1×1(.) is used to reduce the number of channels to match the number of the next stage, expressed as:

[0128] EUCB(x)=C 1×1 (ReLU(BN(DWC(Up(x))))).

[0129] S5. Construct a loss function, use the remote sensing image dataset constructed in step S1 to train, test, and evaluate the MambaVision encoder obtained in step S2, the temporal feature interaction module TFIM obtained in step S3, and the multi-scale convolutional decoder EMCAD obtained in step S4, and apply the trained optimal model to high-resolution remote sensing change detection.

[0130] Furthermore, the specific implementation method of step S5 is to generate four prediction graphs p1, p2, p3 and p4 at each stage based on the four segmentation heads of the multi-scale convolutional decoder EMCAD, calculate the loss of all possible prediction combinations obtained from the four segmentation heads, minimize this cumulative combination loss during the training process, and obtain the loss function The expression is:

[0131]

[0132] in, and are the losses of each prediction graph, α=β=γ=ζ=δ=1.0 is the weight assigned to each loss, is an additional loss item;

[0133] The prediction map p1 at the final stage of the decoder is regarded as the final segmentation map, and then the final segmentation output is obtained by performing binary classification segmentation using the Sigmoid function.

[0134] Embodiment 2:

[0135] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a high-resolution remote sensing image change detection method based on multi-scale convolution decoding described in Example 1 are implemented.

[0136] The computer device of the present invention may be a device including a processor and a memory, such as a single chip microcomputer including a central processing unit. Furthermore, the processor is used to implement the steps of the above-mentioned high-resolution remote sensing image change detection method based on multi-scale convolution decoding when executing the computer program stored in the memory.

[0137] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf programmable gate arrays (FPGAs), or a processor.

[0138] Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0139] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0140] Embodiment 3:

[0141] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the high-resolution remote sensing image change detection method based on multi-scale convolutional decoding described in Example 1 is implemented.

[0142] The computer-readable storage medium of the present invention can be any form of storage medium that can be read by a processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. The computer-readable storage medium stores a computer program. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned high-resolution remote sensing image change detection method based on multi-scale convolutional decoding can be implemented.

[0143] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0144] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0145] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and parts thereof may be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A high-resolution remote sensing image change detection method based on multi-scale convolutional decoding, characterized in that: The steps include: S1. Collect remote sensing images through sensors carried by satellites or drones, perform preprocessing, and then construct a remote sensing image dataset; S2. Construct a MambaVision encoder, which consists of four different stages. The first and second stages are based on CNN modules, and the third and fourth stages are based on MambaVision modules and Transformer modules. The data in the remote sensing image dataset obtained in step S1 is input into the MambaVision encoder, and four dual-phase multi-scale feature maps are obtained through four-stage processing. S3. Construct a time feature interaction module TFIM, input the dual-phase multi-scale feature map obtained in step S2 into the time feature interaction module TFIM, and obtain a feature fusion feature map; S4. Construct a multi-scale convolution decoder EMCAD, including using four multi-scale convolution attention modules MSCAM to optimize the feature fusion feature map obtained in step S3. After each MSCAM, use the segmentation head SH to generate the segmentation map of this stage, upsample the optimized feature fusion feature map through the efficient upconvolution block EUCB, and add the upsampled feature map to the corresponding group attention gate LGAG output. Finally, add the four different segmentation maps to generate the final prediction map; S5. Construct a loss function, use the remote sensing image dataset constructed in step S1 to train, test, and evaluate the MambaVision encoder obtained in step S2, the temporal feature interaction module TFIM obtained in step S3, and the multi-scale convolutional decoder EMCAD obtained in step S4, and apply the trained optimal model to high-resolution remote sensing change detection.

2. The high-resolution remote sensing image change detection method based on multi-scale convolutional decoding according to claim 1, characterized in that: The preprocessing method of step S1 includes geometric correction, cloud and shadow removal, clutter cleaning, image enhancement, scale normalization and noise suppression; The remote sensing image dataset is divided into a training set, a validation set, and a test set according to a data ratio of 6:2:

2.

3. The high-resolution remote sensing image change detection method based on multi-scale convolutional decoding according to claim 2 is characterized in that: The specific implementation method of step S2 includes the following steps: S2.

1. Set the first and second stages based on CNN modules, including input images of size H×W×3, H is the height, W is the width, and the input is first converted to size The layers are overlapped and projected into a C-dimensional embedding space by two consecutive 3×3 convolutional layers with a stride of 2. The downsampler between stages consists of a batch normalized 3×3 convolutional layer with a stride of 2 to halve the image resolution. The expression of the CNN module is: in, is the partial output after the residual block, z is the input, GELU is the Gaussian error linear unit activation function, BN is the batch normalization function, Conv 3×3 (·) denotes a 3×3 convolutional layer with batch normalization and ReLU activation function; S2.

3. Set the third and fourth stages to be based on the MambaVision module and the Transformer module; S2.3.

1. Set the given input to X1 and the output of the MambaVision module to X out Calculated according to the following formula: Among them, Linear(C in , C out ) represents a linear layer, where C in and C out are the embedding dimensions of input and output respectively, Scan is the selective scanning operation, σ is the Sigmoid Linear activation function, and Concat is the concatenation operation; S2.3.

2. Set the expression of the multi-head self-attention mechanism Transformer module to: Where Q, K, and V represent query, key, and value, respectively. h represents the number of attention heads, and Softmax is the activation function; S2.3.

3. Setting input Where the sequence length is T, the calculation formula for the output of the nth layer in the third stage and the output of the nth layer in the fourth stage is: in, is the output after processing by the MambaVision module, Xn is the output after the multi-head self-attention mechanism, and X n-1 is the input, Norm and Mixer represent the selection of layer normalization and token mixing modules respectively; Set the given N layers, before The MambaVision module is used in the layer, and then The layer adopts a multi-head self-attention mechanism; S2.

4. The dual-phase multi-scale feature map F is obtained through four-stage processing. i , i=1,2,3,4.

4. The high-resolution remote sensing image change detection method based on multi-scale convolutional decoding according to claim 3 is characterized in that: The specific implementation method of step S3 includes the following steps: S3.

1. For the i-th bi-phase multi-scale feature map obtained in step S2, the difference feature map is calculated by element-by-element subtraction and then absolute value operation. The expression is: in, is the first phase characteristic diagram, is the characteristic diagram of the second phase, |·| represents the absolute value operation, Represents element-by-element subtraction operation, Conv 3×3 (·) denotes a 3×3 convolutional layer with batch normalization and ReLU activation function; S3.

2. Applied to the time feature, the change area in the time feature is highlighted by element-by-element multiplication operation, and the enhanced time feature map is added to the original feature map. The expression is: in, is the refined first phase feature, For the refined second phase characteristics, and Represent element-wise multiplication and addition operations respectively; S3.

3. Through the splicing of refined dual-phase features, a 3×3 convolutional layer is used to mine the time difference information, and the number of channels of the spliced ​​feature is halved. Then, the time difference representation is added to the previous coarse time difference feature. Finally, a 1×1 convolutional layer is used to reduce the number of channels. The expression is: Among them, D i To obtain the temporal difference features, Cat(·) represents the feature concatenation operation; S3.

4. Apply the methods of steps S3.1 to S3.3 to the four dual-phase multi-scale feature maps to generate four multi-feature fused feature maps.

5. The high-resolution remote sensing image change detection method based on multi-scale convolutional decoding according to claim 4 is characterized in that: The specific implementation method of step S4 includes the following steps: S4.

1. Construct the group attention gate LGAG, which is set in the qatt(.) function. First, apply independent 3×3 group convolution GC g (.) and GC x (.) processes the 1×1 convolution gating signal g and the input feature map x, and the convolved features are then normalized by batch normalization BN(.) and merged by element-by-element addition; the generated feature map is activated by a ReLU (R(.)) layer, after which a 1×1 convolution C(.) is applied followed by a BN(.) layer to obtain a single-channel feature map; then, the obtained single-channel feature map is activated by the Sigmoidσ(.) activation function to generate the attention coefficient; the output of this transformation is used to scale the input feature x by element-by-element multiplication to generate the attention gate feature LGAG(g, x), expressed as: q att (g,x)=R(BN(GC g (g)+BN(GC x (x))))) S4.

2. Construct a multi-scale convolutional attention module MSCAM, which consists of a channel attention block CAB, a spatial attention block SAB, and a multi-scale convolution block MSCB, and the expression is: MSCAM(x)=MSCB(SAB(CAB(x))) Among them, x represents the input tensor; S4.2.

1. Construct a multi-scale convolution block MSCB. First, expand the number of channels by a point-by-point convolution layer PWC1(·) with an expansion factor of 2, and then perform batch normalization BN(·) and ReLU6 activation function R6(.); then, capture multi-scale and multi-resolution contextual information through multi-scale depth convolution MSDC(.); introduce channel shuffling operation to enhance the correlation between channels; then, use another point-by-point convolution PWC2(.) and combine BN(.) to restore the number of channels to the original number, and encode the dependency between channels, expressed as: MSCB(x)=BN(PWC2(CS(MSDC(R6(BN(PWC1(x))))))) MSDC(x)=∑ ks∈KS DWCB ks (x) x=x+DWCB ks (x); S4.2.

2. Construct a channel attention block CAB. First, apply adaptive maximum pooling Pm(·) and adaptive average pooling Pa(·) to the spatial dimension to extract the most significant features of the entire feature map in each channel. Then, for each pooled feature map, reduce the number of channels to r = 1 / 16 of the original through point-by-point convolution C1(·), and then perform ReLU activation R. After that, use another point-by-point convolution C2(·) to restore the original number of channels. Next, add the two restored feature maps and apply the Sigmoid (σ) activation function to estimate the attention weight. Finally, use the Hadamard product Incorporating the weights into the input x, the expression is: S4.2.

3. Construct a spatial attention block SAB. First, perform maximum pooling Chmax(·) and average pooling Chavg(·) along the channel dimension to focus on local features. Then, use a 7×7 large kernel convolution layer to enhance the local contextual relationship between features. Then, apply the Sigmoid activation function (σ) to calculate the attention weight. Finally, use the Hadamard product Apply the weights to the input x, expressed as: S4.

3. Construct an efficient upconvolution block EUCB. First, use UpSampling Up(·) to perform upsampling by a factor of 2. Then, by applying a 3×3 depth convolution DWC(·), combined with batch normalization BN(·) and ReLU(.) activation function, the upsampled feature map is enhanced; finally, a 1×1 convolution C1×1(.) is used to reduce the number of channels to match the number of the next stage, expressed as: EUCB(x)=C 1×1 (ReLU(BN(DWC(Up(x)))))。 6. The high-resolution remote sensing image change detection method based on multi-scale convolutional decoding according to claim 5, characterized in that: The specific implementation method of step S5 is to generate four prediction graphs p1, p2, p3 and p4 at each stage based on the four segmentation heads of the multi-scale convolutional decoder EMCAD, calculate the loss of all possible prediction combinations obtained from the four segmentation heads, minimize this cumulative combination loss during the training process, and obtain the loss function The expression is: in, and are the losses of each prediction graph, α=β=γ=ζ=δ=1.0 is the weight assigned to each loss, is an additional loss item; The prediction map p1 at the final stage of the decoder is regarded as the final segmentation map, and then the final segmentation output is obtained by performing binary classification segmentation using the Sigmoid function.

7. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a high-resolution remote sensing image change detection method based on multi-scale convolution decoding as described in any one of claims 1 to 6 when executing the computer program.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the high-resolution remote sensing image change detection method based on multi-scale convolution decoding described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Micro-expression detection method, system and device and storage medium

    CN116071810A

  • Double-temporal remote sensing image change detection network based on attention and multiple scales

    CN116343052A