Multi-Feature Fusion Remote Sensing Image Change Detection Method and System
By combining the multi-feature fusion method of Transformer and high-resolution convolutional network, the problem of loss of detailed information caused by excessive global information in remote sensing image change detection is solved, and more accurate change area segmentation and edge detection are achieved.
Patent Information
- Application Number
- CN202310340120.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-03-29
AI Technical Summary
In the existing remote sensing image change detection, when using the Transformer model, there is a problem of excessive global information attention, resulting in loss of detailed information, especially the inaccurate segmentation of the edge area of the changing object.
Using a multi-feature fusion remote sensing image change detection method, combined with Transformer and high-resolution convolution network, multiple preliminary features are extracted and feature fusion is performed to enhance the segmentation accuracy of edge information by using Transformer layer and convolutional layer in parallel.
While maintaining global information, the segmentation accuracy of the changing area is improved, the detection effect of edge information is enhanced, and the shortcomings of the Transformer model in remote sensing image change detection are overcome.
Smart Images

Figure CN116310692B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing and remote sensing image change detection technologies, and particularly to a multi-feature fusion remote sensing image change detection method and system. Background Art
[0002] The statements in this section only mention the background technologies related to the present invention and do not necessarily constitute prior art.
[0003] Remote sensing image change detection is an application in the field of remote sensing images, aiming to identify the change information between two-temporal remote sensing images. By comparing the remote sensing images obtained at different times in the same area and assigning binary labels of changed and unchanged to each pixel, the change information of the area is obtained. This technology is widely used in urban development planning, land management, disaster assessment and other fields to analyze and solve corresponding specific problems. With the development of remote sensing satellite technology, the acquisition of remote sensing data has become more and more convenient, and the related fields of remote sensing images have also been better developed accordingly.
[0004] Transformer in the field of natural language processing has been migrated to Vision Transformer (ViT) and applied in the field of computer image processing, achieving quite good performance. By adopting the self-attention method, the model pays more attention to global information and can learn the pixel relationships at a relatively long distance in the whole image. In remote sensing images, there are often situations where objects at a relatively long distance have similar features, such as houses, roads, trees, etc. at different positions. Therefore, the idea of Transformer has also been introduced into the direction of processing related fields of remote sensing images and achieved good results. Since Transformer adopts the self-attention method and obtains more global information, although each pixel can pay attention to the information relationships with all other pixels, it blurs a large amount of relevant information that adjacent pixels should have, resulting in the loss of detailed information and inaccurate segmentation of the edge regions of changed objects. Summary of the Invention
[0005] To solve the deficiencies of the prior art, the present invention provides a multi-feature fusion remote sensing image change detection method and system; by using Transformer and high-resolution convolution in parallel, the segmentation accuracy of the changed area can be improved while obtaining global information, so as to effectively obtain a changed image with accurate detection and better edge information.
[0006] In the first aspect, the present invention provides a multi-feature fusion remote sensing image change detection method;
[0007] The multi-feature fusion remote sensing image change detection method includes:
[0008] Obtain two-temporal remote sensing images to be detected;
[0009] Preprocess the dual-temporal remote sensing images to be detected;
[0010] Input the preprocessed dual-temporal remote sensing images into the trained remote sensing image change detection model and output the detection results;
[0011] Among them, the trained remote sensing image change detection model is used to perform preliminary feature extraction on the input dual-temporal remote sensing images to be detected, extract several different preliminary features, further extract feature maps for each preliminary feature, then perform multi-feature fusion on different feature maps to obtain fusion features, and finally perform prediction on the fusion features to output a change map.
[0012] In a second aspect, the present invention provides a multi-feature fusion remote sensing image change detection system;
[0013] The multi-feature fusion remote sensing image change detection system includes:
[0014] An acquisition module configured to acquire dual-temporal remote sensing images to be detected;
[0015] A preprocessing module configured to preprocess the dual-temporal remote sensing images to be detected;
[0016] An output module configured to input the preprocessed dual-temporal remote sensing images into the trained remote sensing image change detection model and output the detection results;
[0017] Among them, the trained remote sensing image change detection model is used to perform preliminary feature extraction on the input dual-temporal remote sensing images to be detected, extract several different preliminary features, further extract feature maps for each preliminary feature, then perform multi-feature fusion on different feature maps to obtain fusion features, and finally perform prediction on the fusion features to output a change map.
[0018] In a third aspect, the present invention further provides an electronic device, including:
[0019] A memory for non-temporarily storing computer-readable instructions; and
[0020] A processor for running the computer-readable instructions,
[0021] Among them, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.
[0022] In a fourth aspect, the present invention further provides a storage medium that non-temporarily stores computer-readable instructions, where when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in the first aspect are executed.
[0023] Fifth aspect, the present invention also provides a computer program product, including a computer program, which is used to implement the method described in the first aspect above when running on one or more processors.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] A multi-feature fusion remote sensing image change detection network structure based on a Transformer hybrid model is proposed, which overcomes the disadvantage of the change detection network lacking continuous information when using the Transformer, and has better accuracy for the edges of change information; compared with the existing methods, it can combine the global information concerned by the Transformers used in different dimensions, effectively enhancing the effect when using the Transformer network. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0027] Figure 1 It is a flowchart of the method for Embodiment 1;
[0028] Figure 2 It is an internal network structure diagram of the remote sensing image change detection model for Embodiment 1;
[0029] Figure 3 It is a structural diagram of the Transformer module for Embodiment 1;
[0030] Figure 4 It is a structural diagram of the multi-feature fusion module for Embodiment 1;
[0031] Figure 5 It is a schematic diagram of the structure of the predictor for Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0033] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0034] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0035] All data acquisition in this embodiment is based on compliance with laws, regulations and user consent, and is a legal application of data.
[0036] For a remote sensing image change detection model using Transformer, combining feature maps with different resolutions to supplement fuzzy edge information and achieving the judgment of change regions has become a simple and effective method. By fusing feature maps with different resolutions during the upsampling process and using a predictor with an attention module, a feature map with a more accurate segmentation effect can be obtained, effectively learning the semantic relationships and detailed edge information between adjacent pixels that are lost, and finally helping to segment more accurate change regions.
[0037] The Transformer hybrid model is a kind of Transformer model mentioned in the ViT model. Different from directly using the Transformer layer to replace the convolutional layer used in the original deep learning, the Transformer hybrid model achieves a similar effect to directly using the Transformer layer by connecting the convolutional layer and the Transformer layer in series. And due to the addition of a standard convolutional structure, the scalability of the model is enhanced.
[0038] Embodiment 1
[0039] This embodiment provides a multi-feature fusion remote sensing image change detection method;
[0040] As Figure 1 shown, the multi-feature fusion remote sensing image change detection method includes:
[0041] S101: Obtain the bi-temporal remote sensing images to be detected;
[0042] S102: Preprocess the bi-temporal remote sensing images to be detected;
[0043] S103: Input the preprocessed dual-temporal remote sensing image into the trained remote sensing image change detection model, and output the detection result;
[0044] Among them, the trained remote sensing image change detection model is used to perform preliminary feature extraction on the input dual-temporal remote sensing image to be detected, extract several different preliminary features, further extract feature maps for each preliminary feature, then perform multi-feature fusion on different feature maps to obtain a fused feature, and finally perform prediction on the fused feature to output a change map.
[0045] Further, the S101: Obtain the dual-temporal remote sensing image to be detected is obtained by collecting remote sensing image data of the same area at two different time points through the photographing unit of a remote sensing satellite and registering the data.
[0046] Further, the S102: Preprocess the dual-temporal remote sensing image to be detected, including: dividing the image into several images with the same size.
[0047] Further, as Figure 2 shown, the trained remote sensing image change detection model includes: a ResNet18 network, and the ResNet18 network includes a conv1 layer, a conv2_x layer, a conv3_x layer, a conv4_x layer, and a conv5_x layer connected in sequence;
[0048] The output end of the conv2_x layer is connected to the input end of the first Transformer layer, the output end of the first Transformer layer is connected to the input end of the first concatenation splicer, the output end of the first concatenation splicer is connected to the input end of the first multi-feature fusion module, and the output end of the first multi-feature fusion module is connected to the first predictor;
[0049] The output end of the conv3_x layer is connected to the input end of the second Transformer layer, the output end of the second Transformer layer is connected to the input end of the second concatenation splicer, the output end of the second concatenation splicer is connected to the input end of the second multi-feature fusion module, and the output end of the second multi-feature fusion module is connected to the input end of the first concatenation splicer;
[0050] The output end of the conv4_x layer is connected to the input end of the third Transformer layer, the output end of the third Transformer layer is connected to the input end of the third concatenation splicer, the output end of the third concatenation splicer is connected to the input end of the third multi-feature fusion module; the output end of the third multi-feature fusion module is connected to the input end of the second concatenation splicer;
[0051] The output end of the conv5_x layer is connected to the input end of the fourth Transformer layer, and the output end of the fourth Transformer layer is connected to the input end of the third concatenator.
[0052] Further, the conv2_x layer includes two convolutional blocks. Each convolutional block includes two sequentially connected 3×3 convolutional layers. After each convolutional layer, there are a batch normalization layer and an activation function layer directly connected to the convolutional layer. Then, through a cross-layer path, these two convolutional layers are skipped, and the result is added before the last activation function layer through a 1×1 convolutional layer.
[0053] The conv3_x layer includes two convolutional blocks. Each convolutional block includes two sequentially connected 3×3 convolutional layers. After each convolutional layer, there are a batch normalization layer and an activation function layer directly connected to the convolutional layer. Then, through a cross-layer path, these two convolutional layers are skipped, and the result is added before the last activation function layer through a 1×1 convolutional layer.
[0054] The conv4_x layer includes two convolutional blocks. Each convolutional block includes two sequentially connected 3×3 convolutional layers. After each convolutional layer, there are a batch normalization layer and an activation function layer directly connected to the convolutional layer. Then, through a cross-layer path, these two convolutional layers are skipped, and the result is added before the last activation function layer through a 1×1 convolutional layer.
[0055] The conv5_x layer includes two convolutional blocks. Each convolutional block includes two sequentially connected 3×3 convolutional layers. After each convolutional layer, there are a batch normalization layer and an activation function layer directly connected to the convolutional layer. Then, through a cross-layer path, these two convolutional layers are skipped, and the result is added before the last activation function layer through a 1×1 convolutional layer.
[0056] Further, the initial feature extraction of the input dual-temporal remote sensing image to be detected, extracting a number of different initial features, is implemented through the conv2_x layer, conv3_x layer, conv4_x layer, and conv5_x layer of the ResNet18 network; the conv1 layer of the ResNet18 network inputs the dual-temporal remote sensing image to be detected, the conv2_x layer outputs the initial feature D1, the conv3_x layer outputs the initial feature D2, the conv4_x layer outputs the initial feature D3, and the conv5_x layer outputs the initial feature D4.
[0057] It should be understood that the dual-temporal images are fed into a pre-trained ResNet-18 network to obtain the initial features of the dual-temporal remote sensing images in different dimensions. In the present invention, the modified ResNet-18 is used as the encoder structure of the Transformer feature extraction network model, and the initialized model parameters obtained by pre-training in the ImageNet dataset are used. The original ResNet-18 contains 5 stages, as Figure 2 shown, and in each stage, the length and width of the image are downsampled to 1 / 2 of the original size. The stride of the convolutional layer in the fifth stage of ResNet-18 is modified to 1, so that the fifth stage no longer downsamples the image.
[0058] Two dual-temporal remote sensing images of the same area with a size of 256*256*3 are processed through the ResNet-18 network, and finally two feature maps with a size of 16*16*512 are obtained as the initial feature information obtained by the feature extraction network.
[0059] Therefore, the remote sensing image data passes through the conv1 layer, conv2_x layer, conv3_x layer, conv4_x layer, and conv5_x layer of the ResNet-18 feature extraction network. The sizes of the output feature maps of each layer are 、 、 、 and , respectively, where H and W represent the height and width of the input original image. Then, the dual-temporal image feature maps extracted by the conv2_x layer, conv3_x layer, conv4_x layer, and conv5_x layer in the ResNet18 feature extraction network are concatenated in the channel dimension into feature maps D1, D2, D3, and D4 with sizes of 64*64*128, 32*32*256, 16*16*512, and 32*32*1024, respectively.
[0060] Furthermore, the further extraction of feature maps for each initial feature is achieved through the first, second, third, and fourth Transformer layers; the first Transformer layer inputs the initial feature D1, and the first Transformer layer outputs the feature map D1 ; the second Transformer layer inputs the initial feature D2, and the second Transformer layer outputs the feature map D2 ; the third Transformer layer inputs the initial feature D3, and the third Transformer layer outputs the feature map D3 ; the fourth Transformer layer inputs the initial feature D4, and the fourth Transformer layer outputs the feature map D4 .
[0061] Further, the internal structures of the first, second, third, and fourth Transformer layers are the same.
[0062] Further, as Figure 3 shown, the first Transformer layer includes eight Transformer feature extraction modules connected in series in sequence; each Transformer feature extraction module includes: an input layer, a first normalization module, a multi-head attention mechanism module, a first adder, a second normalization module, a multi-layer perceptron, a second adder, and an output layer; the output end of the input layer is also connected to the input end of the first adder; the output end of the first adder is also connected to the input end of the second adder.
[0063] It should be understood that each group of feature maps obtained in the ResNet18 network is respectively sent into the Transformer layer to obtain feature maps; the feature maps D1, D2, D3, and D4 obtained in the ResNet-18 network are respectively sent into the Transformer layer for feature map extraction to capture the global semantic information of the corresponding dimensions, and feature maps D1 , D2 , D3 , D4 are obtained.
[0064] The Transformer feature extraction layer is composed of 8 Transformer feature extraction modules, and each Transformer feature extraction module is composed of Figure 3 as shown.
[0065] The sizes of the feature maps output after passing through the Transformer layer are the same as those at the input, which are respectively , , and , where H and W respectively represent the height and width of the input original image.
[0066] Further, the multi-feature fusion of different feature maps to obtain the fused features is realized through the first, second, and third multi-feature fusion modules. The third multi-feature fusion module performs feature fusion on the concatenated result of feature map D3 and feature map D4 to obtain feature map D3 ;
[0067] The second multi-feature fusion module performs feature fusion on feature map D3 and feature map D2 Perform feature fusion on the concatenation result to obtain feature map D2 ;
[0068] The first multi-feature fusion module performs feature fusion on the concatenation result of feature map D2 and feature map D1 to obtain feature map D1 .
[0069] Furthermore, the internal structures of the first, second, and third multi-feature fusion modules are the same.
[0070] Furthermore, as Figure 4 shown, the first multi-feature fusion module includes: an input layer;
[0071] The input end of the input layer is used to input the output result of the corresponding concatenator;
[0072] The output end of the input layer is respectively connected to the input ports of the first, second, and third branches;
[0073] The first branch is a 1*1 convolutional layer; the second branch is an average pooling layer; the third branch is a max pooling layer; the output ends of the second branch and the third branch are both connected to the input end of the third adder;
[0074] The output end of the third adder is connected to the input end of the activation function layer; the output end of the activation function layer and the output end of the first branch are both connected to the input end of the first multiplier, and the output end of the first multiplier outputs the final feature fusion result.
[0075] It should be understood that each group of feature maps obtained by the Transformer feature extraction network is sent to the feature fusion module, and the feature maps obtained in the Transformer feature extraction network are gradually upsampled. Each operation concatenates the feature maps with the same size in feature maps D1 , D2 , D3 , D4 , and fuses the feature maps through the multi-feature fusion module, and finally obtains a higher-size feature map through the upsampling operation.
[0076] Furthermore, the working processes of the first, second, and third multi-feature fusion modules are the same;
[0077] The working process of the third multi-feature fusion module is as follows:
[0078] For and two feature maps, concatenate these two feature maps into a feature map with a size of Feature maps of size are used as the input of the multi-feature fusion module; as Figure 4 shown, the input feature map is respectively passed through an average pooling layer and a max pooling layer to obtain two feature maps of size , and these two feature maps are added together and passed through a sigmoid activation function to obtain a feature map E1 with spatial attention;
[0079] At the same time, the input feature map is passed through a 1×1 convolutional layer to reduce the number of channels, and a feature map E2 with the number of channels being the number of input channels is obtained. E2 is a feature map of size ;
[0080] Finally, E1 is multiplied by E2 to obtain the output of the multi-feature fusion module. The output has the same size as E2, which is . This output is subjected to a linear interpolation upsampling operation to obtain a new feature map of size , where H and W respectively represent the height and width of the original input image.
[0081] Exemplarily, for the first, second, and third multi-feature fusion modules, the overall process is that the input feature maps D3 , D4 are used to obtain a new feature map of size ; ;
[0082] Then, D3 and the feature map D2 are sent into the multi-feature fusion module to obtain a feature map D2 of size ; ;
[0083] Finally, D2 and the feature map D1 are sent into the multi-feature fusion module to obtain a feature map D1 of size ; .
[0084] D1 is subjected to a linear interpolation upsampling operation and passed through a 1×1 convolutional layer to reduce the number of channels to obtain the output D of the upsampling feature fusion module.
[0085] The size of the feature map D is , where H and W respectively represent the height and width of the original input image.
[0086] Furthermore, for the prediction of the output change map of the fused features, a predictor is used to predict the feature map D1 to output the change map.
[0087] Exemplarily, two groups of target resolution feature maps are sent into a predictor to obtain a change map, and the feature map with the size of is sent into a fully convolutional network predictor with attention for change recognition. As Figure 5 shown, the predictor consists of an attention module and a fully convolutional network layer, and the output size is , where H and W respectively represent the height and width of the input original image.
[0088] In the attention module, the input image obtains the channel attention information of the feature map through a global average pooling layer and a convolutional layer with a 1*1 convolutional kernel, and then controls the output feature between 0 and 1 through a sigmoid activation function to obtain a feature map with the size of channel feature map F. By multiplying the channel feature map F with the input image, a feature map with channel attention information is obtained, and then directly added to the input image to avoid a certain degree of information loss caused by the attention module, so as to obtain a feature map with the same size as the input image .
[0089] The fully convolutional network layer consists of a 3*3 convolutional layer, a batch normalization layer, a ReLU activation function, and a 3*3 convolutional layer. Among them, the number of output channels of the first convolutional layer is the same as the number of input channels, while the number of output channels of the second convolutional layer is 2, and the output size is , where H and W respectively represent the height and width of the input original image. The obtained output is the change map. The two channels of each pixel in the change map respectively represent the probabilities of the unchanged situation and the changed situation of the pixel, and the larger one of the two is used as the actual state of the pixel for output.
[0090] Furthermore, as Figure 5 shown, the predictor includes an attention module and a fully convolutional network layer;
[0091] The input end of the attention module is the input end of the predictor;
[0092] The attention module includes an average pooling layer, a 1*1 convolutional layer, and an activation function layer connected in sequence;
[0093] The input ends of the input end of the predictor and the output end of the activation function layer are both connected to the input end of the second multiplier;
[0094] The output end of the second multiplier and the input end of the predictor are both connected to the input end of the fourth adder;
[0095] The output end of the fourth adder is connected to the input end of the fully convolutional network layer;
[0096] The fully convolutional network layer includes: a 3×3 convolutional layer C1, a batch normalization layer, an activation function layer, and a 3×3 convolutional layer C2 connected in sequence; the 3×3 convolutional layer C1 serves as the input end of the fully convolutional network layer, the 3×3 convolutional layer C2 serves as the output end of the fully convolutional network layer, and the 3×3 convolutional layer C2 serves as the output end of the predictor.
[0097] Further, for the trained remote sensing image change detection model, the training process includes:
[0098] Construct a data set, and divide the data set into a training set, a validation set, and a test set according to a ratio. The training set is a pair of temporal remote sensing images with known image change results.
[0099] Input the training set into the remote sensing image change detection model to train the remote sensing image change detection model. When the loss function value of the remote sensing image change detection model no longer decreases or the number of iterations exceeds a set threshold, stop the training to obtain the trained remote sensing image change detection model.
[0100] Use the validation set to validate the trained remote sensing image change detection model in each iteration, and retain the model parameters with the highest score during validation.
[0101] Use the test set to test the remote sensing image change detection model after all iterations are completed.
[0102] Further, for the loss function, a cross-entropy loss function is used. The loss value is calculated by comparing the output result of the predictor with the true label of the pixel, and the parameters of the model are updated by backpropagation. The cross-entropy loss function is as follows:
[0103]
[0104] where H and W respectively represent the height and width of the input original image, represents the predicted label of the pixel at position (h, w), represents the true label of the pixel at position (h, w).
[0105] In the present invention, a pair of temporal images are sent into a pre-trained ResNet18 network to obtain preliminary features of the pair of temporal remote sensing images in different dimensions; each group of feature maps obtained in the ResNet18 network is respectively sent into a Transformer feature extraction module to obtain feature maps; the groups of feature maps obtained by the Transformer feature extraction network are sent into a feature fusion module; the feature maps after feature fusion are sent into a predictor to obtain a change map.
[0106] Exemplarily, dividing the data set into a training set, a validation set, and a test set according to a ratio means dividing the original data set into a training set, a validation set, and a test set according to a ratio of 7:1:2.
[0107] Exemplarily, during the process of constructing the dataset, the data is preprocessed. The preprocessing includes: due to the limitation of computing resources, the original remote sensing image with a size of 1024 * 1024 pixels per piece is divided into 16 pieces of images with a size of 256 * 256 pixels. Random cropping, rotation, flipping, Gaussian blurring and other data augmentation processes are performed on each divided image, and the processed images are used as the dataset for learning. The final obtained image size is 256 * 256 * 3 (length * width * number of channels).
[0108] Embodiment 2
[0109] This embodiment provides a multi-feature fusion remote sensing image change detection system;
[0110] The multi-feature fusion remote sensing image change detection system includes:
[0111] An acquisition module, which is configured to: acquire dual-temporal remote sensing images to be detected;
[0112] A preprocessing module, which is configured to: preprocess the dual-temporal remote sensing images to be detected;
[0113] An output module, which is configured to: input the preprocessed dual-temporal remote sensing images into the trained remote sensing image change detection model and output the detection results;
[0114] Among them, the trained remote sensing image change detection model is used to perform preliminary feature extraction on the input dual-temporal remote sensing images to be detected, extract several different preliminary features, further extract feature maps for each preliminary feature, then perform multi-feature fusion on different feature maps to obtain fusion features, and finally perform prediction on the fusion features to output a change map.
[0115] It should be noted here that the above acquisition module, preprocessing module and output module correspond to steps S101 to S103 in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.
[0116] In the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0117] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0118] Embodiment 3
[0119] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in Embodiment 1 above.
[0120] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0121] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.
[0122] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.
[0123] The method in Embodiment 1 can be directly embodied as being executed by the hardware processor, or executed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0124] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0125] Embodiment 4
[0126] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in Embodiment 1 is completed.
[0127] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-feature fusion remote sensing image change detection method, characterized in that Including: Obtain the dual-temporal remote sensing image to be detected; Preprocess the dual-temporal remote sensing image to be detected; Input the preprocessed dual-temporal remote sensing image into the trained remote sensing image change detection model, and output the detection result; Among them, the trained remote sensing image change detection model is used to perform preliminary feature extraction on the input dual-temporal remote sensing image to be detected, extract several different preliminary features, further extract feature maps for each preliminary feature, then perform multi-feature fusion on different feature maps to obtain a fusion feature, and finally perform prediction on the fusion feature to output a change map; The trained remote sensing image change detection model includes: a ResNet18 network, and the ResNet18 network includes a conv1 layer, a conv2_x layer, a conv3_x layer, a conv4_x layer, and a conv5_x layer connected in sequence; The output end of the conv2_x layer is connected to the input end of the first Transformer layer, the output end of the first Transformer layer is connected to the input end of the first concatenation splicer, the output end of the first concatenation splicer is connected to the input end of the first multi-feature fusion module, and the output end of the first multi-feature fusion module is connected to the first predictor; The output end of the conv3_x layer is connected to the input end of the second Transformer layer, the output end of the second Transformer layer is connected to the input end of the second concatenation splicer, the output end of the second concatenation splicer is connected to the input end of the second multi-feature fusion module, and the output end of the second multi-feature fusion module is connected to the input end of the first concatenation splicer; The output end of the conv4_x layer is connected to the input end of the third Transformer layer, the output end of the third Transformer layer is connected to the input end of the third concatenation splicer, the output end of the third concatenation splicer is connected to the input end of the third multi-feature fusion module; the output end of the third multi-feature fusion module is connected to the input end of the second concatenation splicer; The output end of the conv5_x layer is connected to the input end of the fourth Transformer layer, and the output end of the fourth Transformer layer is connected to the input end of the third concatenation splicer.
2. The multi-feature fusion remote sensing image change detection method according to claim 1, characterized in that, The preliminary feature extraction of the input dual-temporal remote sensing image to be detected to extract several different preliminary features is implemented by the conv2_x layer, conv3_x layer, conv4_x layer, and conv5_x layer of the ResNet18 network; the conv1 layer of the ResNet18 network inputs the dual-temporal remote sensing image to be detected, the conv2_x layer outputs the preliminary feature D1, the conv3_x layer outputs the preliminary feature D2, the conv4_x layer outputs the preliminary feature D3, and the conv5_x layer outputs the preliminary feature D4.
3. The multi-feature fusion remote sensing image change detection method according to claim 2, characterized in that, The further extraction of the feature map for each preliminary feature is achieved through the first, second, third, and fourth Transformer layers; the first Transformer layer inputs the preliminary feature D1, and the first Transformer layer outputs the feature map D1 ; The second Transformer layer inputs the preliminary feature D2 and outputs the feature map D2 ; The third Transformer layer inputs the preliminary feature D3 and outputs the feature map D3 ; The fourth Transformer layer inputs the preliminary feature D4 and outputs the feature map D4 .
4. The multi-feature fusion remote sensing image change detection method according to claim 3, characterized in that, The first Transformer layer includes eight Transformer feature extraction modules connected in series in sequence; each Transformer feature extraction module includes: an input layer, a first normalization module, a multi-head attention mechanism module, a first adder, a second normalization module, a multi-layer perceptron, a second adder, and an output layer; the output end of the input layer is further connected to the input end of the first adder; the output end of the first adder is further connected to the input end of the second adder.
5. The multi-feature fusion remote sensing image change detection method according to claim 3, characterized in that, The multi-feature fusion of different feature maps to obtain the fused feature is achieved through the first, second, and third multi-feature fusion modules. The third multi-feature fusion module performs feature fusion on the concatenation result of feature map D3 and feature map D4 to obtain feature map D3 ; The second multi-feature fusion module performs feature fusion on the concatenation result of the feature map D3 and the feature map D2 to obtain the feature map D2 ; The first multi-feature fusion module performs feature fusion on the concatenation result of the feature map D2 and the feature map D1 to obtain the feature map D1 ; The first multi-feature fusion module includes: an input layer; The input end of the input layer is used to input the output result of the corresponding concatenator; The output end of the input layer is respectively connected to the input ports of the first, second, and third branches; The first branch is a 1*1 convolutional layer; the second branch is an average pooling layer; the third branch is a max pooling layer; the output ends of the second branch and the third branch are both connected to the input end of the third adder; The output end of the third adder is connected to the input end of the activation function layer; the output end of the activation function layer and the output end of the first branch are both connected to the input end of the first multiplier, and the output end of the first multiplier outputs the result of the final feature fusion.
6. The multi-feature fusion remote sensing image change detection method according to claim 5, characterized in that, The working processes of the first, second, and third multi-feature fusion modules are the same; The working process of the third multi-feature fusion module is as follows: For and two feature maps, concatenate these two feature maps into a feature map with a size of as the input of the current multi-feature fusion module; separately pass the input feature map through an average pooling layer and a max pooling layer to obtain two feature maps with a size of size, add these two feature maps, and pass through a sigmoid activation function to obtain a feature map E1 with spatial attention; Meanwhile, the input feature map is passed through a 1×1 convolutional layer to reduce the number of channels, obtaining a feature map E2 with the number of channels equal to the number of input channels ; E2 is a feature map with a size of ; Finally, multiply E1 by E2 to obtain the output of the multi-feature fusion module. The output has the same size as E2, which is . Perform a linear interpolation upsampling operation on this output to obtain a new feature map with a size of , where H and W represent the height and width of the original input image respectively.
7. Remote sensing image change detection system with multi-feature fusion, characterized in that It includes: An acquisition module, which is configured to: acquire the dual-temporal remote sensing image to be detected; A preprocessing module, which is configured to: preprocess the dual-temporal remote sensing image to be detected; An output module, which is configured to: input the preprocessed dual-temporal remote sensing image into the trained remote sensing image change detection model and output the detection result; Among them, the trained remote sensing image change detection model is used to perform preliminary feature extraction on the input dual-temporal remote sensing image to be detected, extract several different preliminary features, further extract feature maps for each preliminary feature, then perform multi-feature fusion on different feature maps to obtain fusion features, and finally perform prediction on the fusion features to output the change map; The trained remote sensing image change detection model includes: a ResNet18 network, and the ResNet18 network includes a conv1 layer, a conv2_x layer, a conv3_x layer, a conv4_x layer, and a conv5_x layer connected in sequence; The output end of the conv2_x layer is connected to the input end of the first Transformer layer, the output end of the first Transformer layer is connected to the input end of the first concatenator, the output end of the first concatenator is connected to the input end of the first multi-feature fusion module, and the output end of the first multi-feature fusion module is connected to the first predictor; The output end of the conv3_x layer is connected to the input end of the second Transformer layer, the output end of the second Transformer layer is connected to the input end of the second concatenator, the output end of the second concatenator is connected to the input end of the second multi-feature fusion module, and the output end of the second multi-feature fusion module is connected to the input end of the first concatenator; The output end of the conv4_x layer is connected to the input end of the third Transformer layer. The output end of the third Transformer layer is connected to the input end of the third concatenator. The output end of the third concatenator is connected to the input end of the third multi-feature fusion module. The output end of the third multi-feature fusion module is connected to the input end of the second concatenator; The output end of the conv5_x layer is connected to the input end of the fourth Transformer layer. The output end of the fourth Transformer layer is connected to the input end of the third concatenator.
8. An electronic device, characterized in that it comprises: A memory for non-temporarily storing computer-readable instructions; And A processor for running the computer-readable instructions, wherein when the computer-readable instructions are run by the processor, the method according to any one of claims 1-7 above is executed.
9. A storage medium, characterized in that, Non-temporarily store computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method according to any one of claims 1-7 are executed.
Citation Information
Patent Citations
Optical remote sensing image change detection method based on adaptive fusion NestedUNet
CN115393718A
Crop drought detection method based on remote sensing image
CN115760866A