Swin Transformer-based remote sensing image fusion method and system
By employing a remote sensing image fusion method based on Swin Transformer and utilizing global attention enhancement and transform window mechanisms, the problems of insufficient information utilization and increased redundant information in remote sensing image fusion are solved, achieving a more efficient remote sensing image fusion effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HUANDING ENVIRONMENTAL BIG DATA RES INST
- Filing Date
- 2023-03-03
- Publication Date
- 2026-07-21
Smart Images

Figure CN116228615B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of remote sensing data fusion technology, and in particular to a remote sensing image fusion method based on Swing Transformer. Background Technology
[0002] Currently, mainstream remote sensing data fusion methods can be categorized into five types: weight function-based, demixing-based, Bayesian-based, hybrid-based, and deep learning-based methods. These methods all require one or more pairs of coarse and fine images (referring to images with low or high spatial resolution) as input reference images, along with a coarse image for the predicted date, to generate a fine image for that date. The first four methods require extensive manual design of empirical functions and setting of hyperparameters, and can only perform linear fitting of the feature relationships between remote sensing images. Deep learning-based methods can effectively address these issues. In recent years, an increasing number of deep learning methods have been applied to the field of remote sensing image fusion. Since image super-resolution reconstruction in computer vision is very similar to remote sensing image fusion, many methods use image super-resolution networks to directly enhance the resolution of coarse images. However, the resolution differences between remote sensing images are very high (Landsat image resolution is 16 times that of Modis), and directly applying image super-resolution methods (where the resolution difference is generally 2-4 times) can lead to problems such as loss of detail. Therefore, some methods add a high-pass modulation module after the super-resolution network to supplement image information using coarse and fine images from the reference date. Additionally, some methods design end-to-end trained convolutional neural networks that concatenate three images together to predict the final fine image.
[0003] Most current deep learning-based remote sensing image fusion methods, including some traditional methods, first segment the input image into smaller images and then perform fusion analysis on each smaller image. This can easily lead to the loss of overall image feature information. Convolutional neural networks (CNNs) have been widely used in remote sensing image fusion algorithms, bringing significant performance improvements. One characteristic of CNNs is that they use fixed weights to aggregate local features from small windows segmented from the image. However, enhancing different texture features may lead to an increase in redundant information. Summary of the Invention
[0004] This application provides a remote sensing image fusion method based on Swing Transformer to solve the technical problems of insufficient utilization of global image information and increased redundant information.
[0005] In a first aspect, embodiments of this application provide a remote sensing image fusion method based on the Swin Transformer. This Swin Transformer-based remote sensing image fusion method is applied to a Swin Transformer-based remote sensing image fusion system. The remote sensing image fusion system includes a global attention enhancement module, a local information fusion module, and a result output module, comprising:
[0006] The global attention enhancement module performs image slicing, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date to generate global image feature blocks for the coarse image of the reference date, global image feature blocks for the fine image of the reference date, and global image feature blocks for the coarse image of the predicted date.
[0007] The local information fusion module utilizes the transform window mechanism and multi-head attention mechanism of the Swing Transformer to perform fusion calculations on image blocks at the same position in the coarse global image feature blocks of the reference date, the fine global image feature blocks of the reference date, and the coarse global image feature blocks of the predicted date, thereby generating fine image feature blocks for the predicted date.
[0008] The result output module recovers the fine image feature blocks of the predicted date to generate a fine image of the predicted date.
[0009] Furthermore, the global attention enhancement module performs image slicing, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, generating global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, including:
[0010] The global attention enhancement module uses a lightweight network to extract image features from the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, to obtain coarse image features of the reference date, fine image feature pairs of the reference date, and coarse image features of the predicted date.
[0011] The global attention enhancement module segments the coarse image features of the reference date, the fine image features of the reference date, and the coarse image features of the predicted date into blocks, resulting in coarse image feature blocks of the reference date, fine image feature blocks of the reference date, and coarse image feature blocks of the predicted date.
[0012] Furthermore, the global attention enhancement module performs image slicing, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, generating global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date. It also includes:
[0013] The global attention enhancement module uses a global average pooling layer to extract features from the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date, to obtain features from the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date.
[0014] The global attention enhancement module, based on a multi-head attention mechanism, enhances the features of the coarse image feature blocks of the reference date, the fine image feature blocks of the reference date, and the coarse image feature blocks of the predicted date, resulting in enhanced features of the coarse image feature blocks of the reference date, enhanced features of the fine image feature blocks of the reference date, and enhanced features of the coarse image feature blocks of the predicted date.
[0015] Furthermore, the global attention enhancement module performs image slicing, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, generating global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date. It also includes:
[0016] The global attention enhancement module uses a Sigmoid layer to process the coarse image feature block enhancement features of the reference date, the fine image feature block enhancement features of the reference date, and the coarse image feature block enhancement features of the predicted date, to obtain the coarse image feature block attention features of the reference date, the fine image feature block attention features of the reference date, and the coarse image feature block attention features of the predicted date.
[0017] The global attention enhancement module performs a dot product between the attention feature of the coarse image feature block of the reference date and the coarse image feature block of the reference date to obtain a dot product feature block of the coarse image of the reference date, and adds the dot product feature block of the coarse image of the reference date to the coarse image feature block of the reference date to obtain a global image feature block of the coarse image of the reference date.
[0018] The global attention enhancement module performs a dot product between the attention feature of the fine image feature block of the reference date and the fine image feature block of the reference date to obtain a fine image dot product feature block of the reference date, and adds the fine image dot product feature block of the reference date to the fine image feature block of the reference date to obtain the fine image global image feature block of the reference date.
[0019] The global attention enhancement module performs a dot product between the attention feature of the coarse image feature block of the predicted date and the coarse image feature block of the predicted date to obtain a coarse image dot product feature block of the predicted date. Then, it adds the coarse image dot product feature block of the predicted date to the coarse image feature block of the predicted date to obtain the global image feature block of the coarse image of the predicted date.
[0020] Furthermore, the local information fusion module utilizes the transform window mechanism and multi-head attention mechanism of the Swing Transformer to perform fusion calculations on image blocks at the same location in the coarse global image feature block of the reference date, the fine global image feature block of the reference date, and the coarse global image feature block of the predicted date, generating a fine image feature block for the predicted date, including:
[0021] The local information fusion module extracts an image block at the same position from the global image feature block of the coarse image of the reference date, the global image feature block of the fine image of the reference date, and the global image feature block of the coarse image of the predicted date, respectively, to obtain the local image feature block of the coarse image of the reference date, the local image feature block of the fine image of the reference date, and the local image feature block of the coarse image of the predicted date.
[0022] The local information fusion module calculates the local image feature blocks of the coarse image and the local image feature blocks of the fine image of the reference date based on the spectral mixing theory, and obtains the conversion coefficient between the coarse image satellite sensor and the fine image satellite sensor.
[0023] The local information fusion module calculates the conversion coefficients and the local image feature blocks of the coarse image of the predicted date to generate the local image feature blocks of the fine image of the predicted date.
[0024] The local information fusion module utilizes the transform window mechanism and multi-head attention mechanism of the Swing Transformer to perform feature fusion on the local image feature blocks of the fine image of the predicted date, thereby generating the fine image feature blocks of the predicted date.
[0025] Furthermore, the result output module recovers the fine image feature blocks of the predicted date to generate a fine image of the predicted date, including:
[0026] The result output module reshapes the fine image feature blocks of the predicted date, and uses the lightweight network to restore the reshaped fine image feature blocks of the predicted date into the fine image of the predicted date.
[0027] Secondly, this application also provides a remote sensing image fusion system based on Swing Transformer, including:
[0028] The global attention enhancement module is used to perform image slicing, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, generating global image feature blocks for the coarse image of the reference date, global image feature blocks for the fine image of the reference date, and global image feature blocks for the coarse image of the predicted date.
[0029] The local information fusion module is used to perform fusion calculations on the image blocks at the same position in the coarse image global image feature blocks of the reference date, the fine image global image feature blocks of the reference date, and the coarse image global image feature blocks of the predicted date using the transformation window mechanism and multi-head attention mechanism of the Swing Transformer, so as to generate the fine image feature blocks of the predicted date.
[0030] The result output module is used to recover the fine image feature blocks of the predicted date and generate a fine image of the predicted date.
[0031] Furthermore, the local information fusion module specifically includes:
[0032] The local image feature extraction unit is used to extract an image block at the same position from the global image feature block of the coarse image of the reference date, the global image feature block of the fine image of the reference date, and the global image feature block of the coarse image of the predicted date, respectively, to obtain the local image feature block of the coarse image of the reference date, the local image feature block of the fine image of the reference date, and the local image feature block of the coarse image of the predicted date.
[0033] The conversion coefficient calculation unit is used to calculate the conversion coefficient between the coarse image satellite sensor and the fine image satellite sensor based on the spectral mixing theory.
[0034] The local feature generation unit is used to calculate the transformation coefficients and the local image feature blocks of the coarse image of the predicted date to generate the local image feature blocks of the fine image of the predicted date.
[0035] The fine image generation unit is used to perform feature fusion on the local image feature blocks of the fine image of the predicted date using the transform window mechanism and multi-head attention mechanism of the Swing Transformer to generate the fine image feature blocks of the predicted date.
[0036] Thirdly, this application also provides a computer device, the computer device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the remote sensing image fusion method as described above.
[0037] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the remote sensing image fusion method as described above.
[0038] Compared to existing technologies, the remote sensing image fusion method based on Swin Transformer provided in this application is applied to a Swin Transformer-based remote sensing image fusion system. This system includes a global attention enhancement module, a local information fusion module, and a result output module. The global attention enhancement module performs image segmentation, global attention calculation, and feature fusion on a coarse image of a reference date, a fine image of a reference date, and a coarse image of a predicted date, generating global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date. The local information fusion module utilizes the Swin Transformer's transform window mechanism and multi-head attention mechanism to perform fusion calculations on image blocks at the same location within the global image feature blocks of the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, generating a fine image feature block for the predicted date. The result output module restores the fine image feature block for the predicted date, generating a fine image for the predicted date. Through the above methods, this application constructs three modules: global attention enhancement, local information fusion, and result output. It utilizes the global attention enhancement mechanism and the transform window mechanism of the Swing Transformer to achieve adaptive feature fusion, thus solving the problems of insufficient utilization of global image information and increased redundant information in current remote sensing image fusion.
[0039] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 A flowchart illustrating the remote sensing image fusion method provided for embodiments of this application;
[0042] Figure 2 A schematic block diagram of a remote sensing image fusion system provided for an embodiment of this application;
[0043] Figure 3 A schematic block diagram of the structure of a computer device provided for an embodiment of this application. Detailed Implementation
[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0045] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0046] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0047] It should be understood that, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish the same or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0048] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0049] The inventors of this application have discovered that most current deep learning-based remote sensing image fusion methods, including some traditional methods, first segment the input image into smaller images and then perform fusion analysis on each smaller image. However, some surface features, such as rivers, are often distributed throughout the entire image. Segmenting into smaller images can easily lead to the loss of overall image feature information, and surface changes appearing in some smaller images cannot be captured well, resulting in insufficient utilization of global image information. Convolutional neural networks (CNNs) have been widely used in remote sensing image fusion algorithms, bringing significant performance improvements. One characteristic of CNNs is that they use fixed weights to aggregate local features from small windows segmented from the image. For homogeneous images, this static calculation method can effectively obtain local shape information, but enhancing different texture features may introduce redundant information. Traditional methods in the field of spatiotemporal fusion of remote sensing images focus on three aspects: reducing sensor bias, adapting to surface changes, and handling mixed pixels. In the past, deep learning-based algorithms have devoted a lot of effort to the first two aspects, but none have been able to propose a solution for mixed pixels; thus, they cannot achieve the remote sensing image fusion requirements of fully utilizing global image information and minimizing redundant information.
[0050] To address the aforementioned issues, this application provides a remote sensing image fusion method based on Swing Transformer.
[0051] See Figure 1 , Figure 1 This is a flowchart illustrating the remote sensing image fusion method provided in the embodiments of this application. The remote sensing image fusion method is applied to a remote sensing image fusion system based on Swin Transformer and includes steps S101-S103.
[0052] Step S101: The global attention enhancement module performs image slicing, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date to generate global image feature blocks for the coarse image of the reference date, global image feature blocks for the fine image of the reference date, and global image feature blocks for the coarse image of the predicted date.
[0053] First, the fusion of remote sensing images requires two different images: a low spatial resolution coarse image (hereinafter referred to as C) and a high spatial resolution fine image (hereinafter referred to as F). Commonly used coarse and fine images are typically obtained from MODIS and Landsat satellite images, respectively. The system designed in this patent requires three images as input: a coarse image (C1) on the reference date, a fine image (F1) on the reference date, and a coarse image (C2) on the predicted date. All three input images are square matrices of the same size, and their side length must be 2. n×10 (3≤n≤7). For the three images of the input system, this embodiment uses image slicing, global attention calculation, and feature fusion to generate a coarse global image feature block for the reference date, a fine global image feature block for the reference date, and a coarse global image feature block for the predicted date, which are used for subsequent image generation.
[0054] Furthermore, step S101 specifically includes:
[0055] The global attention enhancement module uses a lightweight network to extract image features from the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, to obtain coarse image features of the reference date, fine image feature pairs of the reference date, and coarse image features of the predicted date.
[0056] The global attention enhancement module segments the coarse image features of the reference date, the fine image features of the reference date, and the coarse image features of the predicted date into blocks, resulting in coarse image feature blocks of the reference date, fine image feature blocks of the reference date, and coarse image feature blocks of the predicted date.
[0057] In this embodiment, for the three input images—C1, C2, and F1—a lightweight network is first used to extract features. Then, each image is divided into 10×10 equal-sized image blocks, resulting in coarse image features for the reference date, fine image feature pairs for the reference date, and coarse image features for the predicted date. The lightweight network may contain multiple convolutional layers and batch normalization layers. The output image of the lightweight network remains unchanged from the input image in terms of planar dimensions. The output feature dimension of the lightweight network for each image is M×H×W, where M is the number of channels, and H and W are the side lengths of the input image. Then, the coarse image features for the reference date, the fine image feature pairs for the reference date, and the coarse image features for the predicted date are further segmented to obtain images of size M×H×W. The image feature blocks are: a coarse image feature block for the reference date, a fine image feature block for the reference date, and a coarse image feature block for the predicted date.
[0058] Furthermore, step S101 specifically includes:
[0059] The global attention enhancement module uses a global average pooling layer to extract features from the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date, to obtain features from the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date.
[0060] The global attention enhancement module, based on a multi-head attention mechanism, enhances the features of the coarse image feature blocks of the reference date, the fine image feature blocks of the reference date, and the coarse image feature blocks of the predicted date, resulting in enhanced features of the coarse image feature blocks of the reference date, enhanced features of the fine image feature blocks of the reference date, and enhanced features of the coarse image feature blocks of the predicted date.
[0061] In this embodiment, for the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date, a global average pooling layer is first used to obtain a feature X of size 100×M, namely the coarse image feature block features of the reference date, the fine image feature block features of the reference date, and the coarse image feature block features of the predicted date. Then, feature enhancement is performed based on a multi-head attention mechanism to obtain the enhanced features of the coarse image feature block of the reference date, the enhanced features of the fine image feature block of the reference date, and the enhanced features of the coarse image feature block of the predicted date. The SoftMax function calculation formula in the multi-head attention mechanism is as follows:
[0062] Q = XW Q K = XW K V = XW V
[0063]
[0064] Among them, W Q W K W V It is a three-parameter matrix, and W Q W K W V ∈R M×d W Q W K W V These are used to multiply with feature X to obtain query vector Q, key vector K, and value vector V, respectively. Q and K are then multiplied by matrix multiplication and normalized using the feature dimension d to calculate the attention matrix. Finally, the normalized attention matrix is multiplied by V to obtain the augmented feature X of size 100×M. ’ That is, the coarse image feature block enhancement feature of the reference date, the fine image feature block enhancement feature of the reference date, and the coarse image feature block enhancement feature of the predicted date.
[0065] Furthermore, step S101 specifically includes:
[0066] The global attention enhancement module uses a Sigmoid layer to process the coarse image feature block enhancement features of the reference date, the fine image feature block enhancement features of the reference date, and the coarse image feature block enhancement features of the predicted date, to obtain the coarse image feature block attention features of the reference date, the fine image feature block attention features of the reference date, and the coarse image feature block attention features of the predicted date.
[0067] The global attention enhancement module performs a dot product between the attention feature of the coarse image feature block of the reference date and the coarse image feature block of the reference date to obtain a dot product feature block of the coarse image of the reference date, and adds the dot product feature block of the coarse image of the reference date to the coarse image feature block of the reference date to obtain a global image feature block of the coarse image of the reference date.
[0068] The global attention enhancement module performs a dot product between the attention feature of the fine image feature block of the reference date and the fine image feature block of the reference date to obtain a fine image dot product feature block of the reference date, and adds the fine image dot product feature block of the reference date to the fine image feature block of the reference date to obtain the fine image global image feature block of the reference date.
[0069] The global attention enhancement module performs a dot product between the attention feature of the coarse image feature block of the predicted date and the coarse image feature block of the predicted date to obtain a coarse image dot product feature block of the predicted date. Then, it adds the coarse image dot product feature block of the predicted date to the coarse image feature block of the predicted date to obtain the global image feature block of the coarse image of the predicted date.
[0070] In this embodiment, attention features of the enhanced feature X' (enhanced features of the coarse image feature block of the reference date, enhanced features of the fine image feature block of the reference date, and enhanced features of the coarse image feature block of the predicted date) are extracted through a Sigmoid layer, resulting in three attention features of size M×100×1×1: the coarse image feature block attention feature of the reference date, the fine image feature block attention feature of the reference date, and the coarse image feature block attention feature of the predicted date. The coarse image feature block attention feature of the reference date, the fine image feature block attention feature of the reference date, and the coarse image feature block attention feature of the predicted date are then multiplied by the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date, respectively. The resulting feature block is then added to the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date, respectively, to obtain three image feature blocks f that incorporate global information. augThat is, the global image feature blocks of the coarse image of the reference date, the global image feature blocks of the fine image of the reference date, and the global image feature blocks of the coarse image of the predicted date, each global image feature block having a size of [missing information].
[0071] Step S102: The local information fusion module uses the transform window mechanism and multi-head attention mechanism of Swing Transformer to perform fusion calculation on the image blocks at the same position in the coarse image global image feature block of the reference date, the fine image global image feature block of the reference date, and the coarse image global image feature block of the predicted date, to generate the fine image feature block of the predicted date.
[0072] Specifically, step S102 includes:
[0073] The local information fusion module extracts an image block at the same position from the global image feature block of the coarse image of the reference date, the global image feature block of the fine image of the reference date, and the global image feature block of the coarse image of the predicted date, respectively, to obtain the local image feature block of the coarse image of the reference date, the local image feature block of the fine image of the reference date, and the local image feature block of the coarse image of the predicted date.
[0074] The local information fusion module calculates the local image feature blocks of the coarse image and the local image feature blocks of the fine image of the reference date based on the spectral mixing theory, and obtains the conversion coefficient between the coarse image satellite sensor and the fine image satellite sensor.
[0075] The local information fusion module calculates the conversion coefficients and the local image feature blocks of the coarse image of the predicted date to generate the local image feature blocks of the fine image of the predicted date.
[0076] The local information fusion module utilizes the transform window mechanism and multi-head attention mechanism of the Swing Transformer to perform feature fusion on the local image feature blocks of the fine image of the predicted date, thereby generating the fine image feature blocks of the predicted date.
[0077] In this embodiment, the local information fusion module extracts an image block from the same location in each of the three input images, resulting in three image blocks: a coarse image local feature block for the reference date, a fine image local feature block for the reference date, and a coarse image local feature block for the predicted date. The size of each image block is [size missing].
[0078] Based on the theory of spectral mixing, the coarse image at coordinate point (x i ,y iThis can be represented as a combination of fine-grained image pixels of different categories. The specific formula is:
[0079]
[0080] l is the number of fine pixel categories, f c Let F(c) be the proportion of fine pixels of category c contained within a coarse pixel, and F(c) be the value of a fine pixel belonging to category c. When the number of categories l is infinitely large, each distinct fine pixel value can be considered as a category. Therefore, for coarse and fine images of the same size, each coarse pixel can be represented by a combination of surrounding fine pixels, and vice versa.
[0081] Extending this to image features, at coordinate points (x i ,y i The fine feature values on a surface can be obtained by combining coarse features. The specific formula is:
[0082]
[0083] Where C' and F' represent local image feature blocks in the coarse image and the fine image, respectively. To obtain the weight matrix value w... i (x j ,y j This requires using local image feature blocks C'1 of the coarse image and F'1 of the fine image from the reference date to analyze and obtain the conversion coefficients between the coarse and fine image satellite sensors. The fine image feature block F'2 for the predicted date is then calculated using the SoftMax function, with the specific formula as follows:
[0084] Q = C′1W Q K = F′1W K V=C′2W V
[0085]
[0086] After multiplying Q and K by matrices, normalize them using the feature dimension d to calculate the attention matrix. Add the normalized attention matrix to the relative position matrix B, and finally multiply it by V. Here, B is the relative position matrix, which represents the relative position of different pixels within each calculation window. C'2 represents the local image feature block of the coarse image for the predicted date. Calculate and fuse the local image feature block F'2 of the fine image for the predicted date using the above formula.
[0087] By combining the transform window mechanism of the Swin Transformer, feature fusion of image patches can be performed through multiple multi-attention mechanisms to finally obtain a fine image feature block of the predicted date, with a size of [missing information].
[0088] S103. The result output module recovers the fine image feature blocks of the predicted date and generates a fine image of the predicted date.
[0089] Specifically, step S103 includes:
[0090] The result output module reshapes the fine image feature blocks of the predicted date, and uses the lightweight network to restore the reshaped fine image feature blocks of the predicted date into the fine image of the predicted date.
[0091] In this embodiment, in the result output module, the fine image feature blocks of the predicted date are reshaped into M×H×W and restored to the same size as the input image through a lightweight network, thereby recovering the fine image of the predicted date.
[0092] Therefore, through the above methods:
[0093] 1. This application proposes a global attention enhancement mechanism to obtain more texture information from a larger global image and enhance image features.
[0094] 2. This application uses the transform window mechanism of Swing Transformer to replace the traditional convolutional neural network, which generates different weight matrices for different inputs, thereby achieving adaptive feature fusion.
[0095] 3. This application incorporates the demixing concept from traditional remote sensing image fusion methods into network design, combining excellent traditional theories with deep learning methods, thus achieving a major innovation in this field.
[0096] Furthermore, this invention also provides a remote sensing image fusion system based on Swin Transformer.
[0097] Please see Figure 2 , Figure 2 This application provides a schematic block diagram of a remote sensing image fusion system based on the Swin Transformer.
[0098] like Figure 2 As shown, this remote sensing image fusion system based on Swin Transformer includes:
[0099] The global attention enhancement module 10 is used to perform image slicing, global attention calculation and feature fusion on the coarse image of the reference date, the fine image of the reference date and the coarse image of the predicted date, to generate global image feature blocks of the coarse image of the reference date, global image feature blocks of the fine image of the reference date and global image feature blocks of the coarse image of the predicted date.
[0100] The local information fusion module 20 is used to perform fusion calculations on image blocks at the same position in the coarse image global image feature block of the reference date, the fine image global image feature block of the reference date, and the coarse image global image feature block of the predicted date using the transform window mechanism and multi-head attention mechanism of the Swing Transformer, to generate a fine image feature block for the predicted date.
[0101] The result output module 30 is used to recover the fine image feature blocks of the predicted date and generate a fine image of the predicted date.
[0102] Furthermore, the global attention enhancement module specifically includes:
[0103] The image feature extraction unit is used to extract image features of the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date using a lightweight network, to obtain coarse image features of the reference date, fine image feature pairs of the reference date, and coarse image features of the predicted date.
[0104] The image feature slicing unit is used to slice the coarse image features of the reference date, the fine image feature pair of the reference date, and the coarse image features of the predicted date into blocks, resulting in coarse image feature blocks of the reference date, fine image feature blocks of the reference date, and coarse image feature blocks of the predicted date.
[0105] A global average pooling unit is used to extract features of the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date using a global average pooling layer, to obtain features of the coarse image feature block of the reference date, features of the fine image feature block of the reference date, and features of the coarse image feature block of the predicted date.
[0106] The feature enhancement unit is used to enhance the coarse image feature block features of the reference date, the fine image feature block features of the reference date, and the coarse image feature block features of the predicted date based on a multi-head attention mechanism, so as to obtain the coarse image feature block enhanced features of the reference date, the fine image feature block enhanced features of the reference date, and the coarse image feature block enhanced features of the predicted date.
[0107] The attention feature extraction unit is used to process the coarse image feature block enhancement features of the reference date, the fine image feature block enhancement features of the reference date, and the coarse image feature block enhancement features of the predicted date using a Sigmoid layer, to obtain coarse image feature block attention features of the reference date, fine image feature block attention features of the reference date, and coarse image feature block attention features of the predicted date.
[0108] A global image feature extraction unit is configured to perform a dot product between the attention feature of the coarse image feature block of the reference date and the coarse image feature block of the reference date to obtain a coarse image dot product feature block of the reference date, and add the coarse image dot product feature block of the reference date to the coarse image feature block of the reference date to obtain a coarse image global image feature block of the reference date; perform a dot product between the attention feature of the fine image feature block of the reference date and the fine image feature block of the reference date to obtain a fine image dot product feature block of the reference date, and add the fine image dot product feature block of the reference date to the fine image feature block of the reference date to obtain a fine image global image feature block of the reference date; perform a dot product between the attention feature of the coarse image feature block of the predicted date and the coarse image feature block of the predicted date to obtain a coarse image dot product feature block of the predicted date, and add the coarse image dot product feature block of the predicted date to the coarse image feature block of the predicted date to obtain a coarse image global image feature block of the predicted date.
[0109] Furthermore, the local information fusion module specifically includes:
[0110] The local image feature extraction unit is used to extract an image block at the same position from the global image feature block of the coarse image of the reference date, the global image feature block of the fine image of the reference date, and the global image feature block of the coarse image of the predicted date, respectively, to obtain the local image feature block of the coarse image of the reference date, the local image feature block of the fine image of the reference date, and the local image feature block of the coarse image of the predicted date.
[0111] The conversion coefficient calculation unit is used to calculate the conversion coefficient between the coarse image satellite sensor and the fine image satellite sensor based on the spectral mixing theory.
[0112] The local feature generation unit is used to calculate the transformation coefficients and the local image feature blocks of the coarse image of the predicted date to generate the local image feature blocks of the fine image of the predicted date.
[0113] The fine image generation unit is used to perform feature fusion on the local image feature blocks of the fine image of the predicted date using the transform window mechanism and multi-head attention mechanism of the Swing Transformer to generate the fine image feature blocks of the predicted date.
[0114] Furthermore, the result output module specifically includes:
[0115] A fine image reshaping unit is used to reshape the fine image feature blocks of the predicted date, and use the lightweight network to restore the reshaped fine image feature blocks of the predicted date into the fine image of the predicted date.
[0116] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described device and each module can be referred to the corresponding processes in the aforementioned control method embodiments, and will not be repeated here.
[0117] Please see Figure 3 , Figure 3 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.
[0118] See Figure 3 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0119] The non-volatile storage medium can store the operating system and computer program. The computer program includes program instructions that, when executed, cause the processor to perform any remote sensing image fusion method based on the Swing Transformer.
[0120] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0121] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any remote sensing image fusion method based on the Swin Transformer.
[0122] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0123] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0124] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the remote sensing image fusion methods based on Swin Transformer provided in the embodiments of this application.
[0125] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0126] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A remote sensing image fusion method based on Swing Transformer, characterized in that, The remote sensing image fusion method based on SwinTransformer is applied to a remote sensing image fusion system based on SwinTransformer. The remote sensing image fusion system includes a global attention enhancement module, a local information fusion module, and a result output module. The remote sensing image fusion method includes: The global attention enhancement module performs image segmentation on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date to obtain multiple image blocks. Global average pooling is performed on each image block to obtain its feature vector. Based on a multi-head attention mechanism, global attention weights are calculated among the feature vectors of all image blocks to obtain enhanced features. A sigmoid layer is then used to process the enhanced features to obtain attention features. These attention features are multiplied by the corresponding original feature blocks and then added to the original feature blocks to generate global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date. The local information fusion module extracts one image block at the same location from the three global image feature blocks to obtain three local image feature blocks. It calculates the conversion coefficient between the coarse image and the fine image based on the spectral mixing theory, and uses the transform window mechanism of the Swing Transformer and the multi-head attention mechanism to perform fusion calculation on the local image feature blocks to generate a fine image feature block for the predicted date. The result output module reshapes the fine image feature blocks of the predicted date and uses a lightweight network to restore the reshaped fine image feature blocks of the predicted date into a fine image of the predicted date; wherein, the lightweight network contains multiple convolutional layers and normalization layers, and the image output by the lightweight network and the image input to the lightweight network remain unchanged in the planar dimension.
2. The remote sensing image fusion method according to claim 1, characterized in that, The global attention enhancement module performs image segmentation, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, generating global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, including: The global attention enhancement module uses a lightweight network to extract image features from the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, to obtain coarse image features of the reference date, fine image feature pairs of the reference date, and coarse image features of the predicted date. The global attention enhancement module segments the coarse image features of the reference date, the fine image features of the reference date, and the coarse image features of the predicted date into blocks, resulting in coarse image feature blocks of the reference date, fine image feature blocks of the reference date, and coarse image feature blocks of the predicted date.
3. The remote sensing image fusion method according to claim 2, characterized in that, The global attention enhancement module performs image slicing, global attention calculation, and feature fusion on the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date, generating global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date. It also includes: The global attention enhancement module uses a global average pooling layer to extract features from the coarse image feature block of the reference date, the fine image feature block of the reference date, and the coarse image feature block of the predicted date, thus obtaining the coarse image feature block features of the reference date, the fine image feature block features of the reference date, and the coarse image feature block features of the predicted date. The global attention enhancement module, based on a multi-head attention mechanism, enhances the features of the coarse image feature blocks of the reference date, the fine image feature blocks of the reference date, and the coarse image feature blocks of the predicted date, resulting in enhanced features of the coarse image feature blocks of the reference date, enhanced features of the fine image feature blocks of the reference date, and enhanced features of the coarse image feature blocks of the predicted date.
4. A remote sensing image fusion system based on Swing Transformer, characterized in that, The remote sensing image fusion system includes: The global attention enhancement module is used to segment the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date into multiple image blocks. Global average pooling is performed on each image block to obtain its feature vector. Based on a multi-head attention mechanism, global attention weights are calculated among the feature vectors of all image blocks to obtain enhanced features. A sigmoid layer is then used to process the enhanced features to obtain attention features. These attention features are then multiplied by the corresponding original feature blocks and added to the original feature blocks to generate global image feature blocks for the coarse image of the reference date, the fine image of the reference date, and the coarse image of the predicted date. The local information fusion module is used to extract an image block at the same location from the three global image feature blocks to obtain three local image feature blocks. The conversion coefficient between the coarse image and the fine image is calculated according to the spectral mixing theory. The local image feature blocks are fused using the transform window mechanism of the Swing Transformer and the multi-head attention mechanism to generate a fine image feature block for the predicted date. The result output module is used to reshape the fine image feature blocks of the predicted date, and use a lightweight network to restore the reshaped fine image feature blocks of the predicted date into a fine image of the predicted date; wherein, the lightweight network contains multiple convolutional layers and normalization layers, and the image output by the lightweight network and the image input to the lightweight network remain unchanged in the planar dimension.
5. The remote sensing image fusion system according to claim 4, characterized in that, The local information fusion module specifically includes: The local image feature extraction unit is used to extract an image block at the same position from the global image feature block of the coarse image of the reference date, the global image feature block of the fine image of the reference date, and the global image feature block of the coarse image of the predicted date, respectively, to obtain the local image feature block of the coarse image of the reference date, the local image feature block of the fine image of the reference date, and the local image feature block of the coarse image of the predicted date; The conversion coefficient calculation unit is used to calculate the conversion coefficient between the coarse image satellite sensor and the fine image satellite sensor based on the spectral mixing theory. A local feature generation unit is used to calculate the transformation coefficients and the local image feature blocks of the coarse image of the predicted date to generate the local image feature blocks of the fine image of the predicted date. The fine image generation unit is used to perform feature fusion on the local image feature blocks of the fine image of the predicted date using the transform window mechanism and multi-head attention mechanism of the Swing Transformer, so as to generate the fine image feature blocks of the predicted date.
6. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the remote sensing image fusion method as described in any one of claims 1 to 3.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the remote sensing image fusion method as described in any one of claims 1 to 3.