Multi-data-source remote sensing image space-spectrum fusion method, device, equipment and medium
Through the multi-data source remote sensing image spatial-spectral fusion network, using gradient guidance and self-supervised loss function, the spatial deformation and spectral distortion problems in remote sensing image spatial-spectral fusion are solved, high-quality image fusion effect is achieved, and the applicability and robustness of the network are improved.
Patent Information
- Application Number
- CN202410875959.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-07-02
AI Technical Summary
Existing remote sensing image spatial-spectral fusion technology has problems of spatial deformation and spectral distortion. Traditional methods are difficult to effectively fit the image fusion process. Deep learning methods lack ideal labels and are prone to information distortion. In addition, the training data is single, making it difficult to achieve a balance between spectral and spatial quality.
A multi-data source remote sensing image spatial-spectral fusion method is adopted. Images of different specifications are obtained through multiple sensors. The multi-data source remote sensing image spatial-spectral fusion network is trained, including gradient guidance, U-shaped multi-channel stitching and information interaction fusion. The training is combined with a self-supervised loss function to enhance the robustness and applicability of the network.
It improves the quality and robustness of remote sensing images, enhances the generalization ability of the network, can better process a variety of satellite image features, and improves the spectral and spatial quality of the fusion results.
Smart Images

Figure CN118735798B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite remote sensing technology, and in particular to a method, device, equipment and medium for spatial-spectral fusion of remote sensing images from multiple data sources. Background Art
[0002] Currently, remote sensing image spatial-spectral fusion techniques encompass two main streams: traditional methods and deep learning approaches. Traditional fusion methods, including component replacement, multi-resolution analysis, and variational optimization, primarily fuse PAN and MS image information using algorithms such as IHS and PCA. These methods are limited in their ability to represent complex landforms using nonlinear models. Furthermore, most traditional methods rely on linear models, making it difficult to effectively fit the fusion process between the two images and prone to spatial and spectral distortion. Conventional supervised deep learning methods lack ideal fusion results as reference labels. The creation of simulated reference labels is cumbersome and can easily lead to loss of spatial information in the panchromatic image. Furthermore, the fusion process ignores the scale difference between simulated and real data, often resulting in a certain degree of information distortion. It is also difficult to achieve a good balance between the spectral and spatial quality of the fusion results. Furthermore, training and testing models are often limited to data from a single specific sensor type, such as using only GF-2 data for training and testing. This approach results in a relatively limited data set, making the trained network models difficult to generalize. Summary of the Invention
[0003] The present invention provides a method, device, equipment and medium for spatial-spectral fusion of remote sensing images from multiple data sources, which is used to solve the problems in related technologies such as the tendency for spatial deformation and spectral distortion to occur, the single source of training data, the cumbersome process steps, the easy loss of spatial information of the full-color image, the existence of a certain degree of information distortion, and the difficulty in achieving a good balance between the spectral quality and spatial quality of the fusion results.
[0004] The present invention provides a spatial-spectral fusion method for remote sensing images from multiple data sources, comprising the following steps.
[0005] Obtain the original image for image spatial-spectral fusion;
[0006] Input the original image into the multi-data source remote sensing image spatial-spectral fusion network to obtain the fused image;
[0007] Among them, the spatial-spectral fusion of remote sensing images from multiple data sources is trained using the following method:
[0008] Acquire multiple sets of panchromatic images PAN and multispectral images MS of different specifications through multiple sensor groups;
[0009] The multiple groups of PAN and the multiple groups of MS are input into an initial multi-data-source remote sensing image space spectrum fusion network for training to obtain the multi-data-source remote sensing image space spectrum fusion network.
[0010] According to the multi-data-source remote sensing image space spectrum fusion method provided in the application, the multi-data-source remote sensing image space spectrum fusion network comprises multiple double-flow fusion layers connected in sequence, and the multiple groups of PAN and the multiple groups of MS are input into the initial multi-data-source remote sensing image space spectrum fusion network for training to obtain the multi-data-source remote sensing image space spectrum fusion network, which comprises:
[0011] The multiple groups of MS are up-sampled to obtain multiple groups of up-sampled MS;
[0012] The multiple groups of PAN, the multiple groups of MS and the multiple groups of up-sampled MS are sequentially input into the multiple double-flow fusion layers for encoding, information interaction fusion and decoding in sequence;
[0013] The multiple groups of PAN, the multiple groups of MS and the multiple groups of up-sampled MS, and the multiple groups of target PAN and the multiple groups of target MS obtained after decoding are input into the next double-flow fusion layer until the multi-data-source remote sensing image space spectrum fusion network is obtained.
[0014] According to the multi-data-source remote sensing image space spectrum fusion method provided in the application, each double-flow fusion layer comprises a first encoding module, a first decoding module, a U-shaped multi-channel splicing module, an information interaction fusion module and a second decoding module, and the multiple groups of PAN, the multiple groups of MS and the multiple groups of up-sampled MS are sequentially input into the multiple double-flow fusion layers for encoding, information interaction fusion and decoding in sequence, which comprises:
[0015] In each double-flow fusion layer, the multiple groups of PAN are input into the first encoding module for encoding to obtain first feature information;
[0016] The multiple groups of MS and the multiple groups of up-sampled MS are input into the U-shaped multi-channel splicing module for encoding to obtain second feature information;
[0017] The first feature information and the second feature information are input into the information interaction fusion module for information interaction fusion to obtain third feature information;
[0018] The multiple groups of PAN and the first feature information are decoded in the first decoding module to obtain target PAN;
[0019] The third feature information is decoded in the second decoding module to obtain target MS.
[0020] According to the multi-data-source remote sensing image space spectrum fusion method provided in the application, the first feature information and the second feature information are input into the information interaction fusion module for information interaction fusion to obtain third feature information, which comprises:
[0021] The third feature information is obtained by interactively fusing the first feature information and the second feature information using the following formula:
[0022] ;
[0023] ;
[0024] ;
[0025] ;
[0026] in, and Respectively expressed in The first feature information and the second feature information extracted from the two-stream fusion layer, S represents the Sigmoid function, represents element-wise multiplication, Indicates the The third feature information obtained from the two-stream fusion layer.
[0027] According to a multi-data source remote sensing image spatial-spectral fusion method provided by the present invention, after obtaining multiple sets of panchromatic images PAN and multiple sets of multispectral images MS of different specifications through multiple sensor groups, the method further includes:
[0028] Gradient guidance of multiple groups of PANs.
[0029] According to a multi-data source remote sensing image spatial-spectral fusion method provided by the present invention, gradient guidance is performed on multiple groups of PANs, including:
[0030] The Laplace operator is used to perform gradient guidance on multiple groups of PANs using the following formula:
[0031] ;
[0032] ;
[0033] ;
[0034] in, Indicates PAN.
[0035] According to a multi-data source remote sensing image spatial-spectral fusion method provided by the present invention, a multi-data source remote sensing image spatial-spectral fusion network is obtained by self-supervised training based on a target loss function, the target loss function is constructed based on a spatial metric function and a spectral metric function, the spatial metric function is used to evaluate the difference between a first residual probability distribution and a second residual probability distribution, the first residual probability distribution is determined based on multiple groups of PAN and target fused images, and the second residual probability distribution is determined based on multiple groups of MS and grayscale images corresponding to the multispectral image MS;
[0036] The spectral metric function is used to evaluate the spectral angle differences between multiple groups of MS and the target fusion image.
[0037] The present invention also provides a multi-data source remote sensing image spatial-spectral fusion device, comprising the following modules:
[0038] An acquisition module, used for acquiring the original image for image spatial-spectral fusion;
[0039] A fusion module is used to input the original image into the multi-data source remote sensing image spatial-spectral fusion network to obtain a fused image;
[0040] A training module, used for acquiring multiple groups of panchromatic images PAN and multiple groups of multispectral images MS of different specifications through multiple sensor groups;
[0041] The training module is also used to input multiple groups of PANs and multiple groups of MSs into the multi-data source remote sensing image spatial-spectral fusion network to obtain the multi-data source remote sensing image spatial-spectral fusion network output.
[0042] According to a multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the multi-data source remote sensing image spatial-spectral fusion network includes a plurality of dual-stream fusion layers connected in sequence, and the training module is specifically used for:
[0043] Upsampling the multiple groups of MSs to obtain multiple groups of upsampled MSs;
[0044] Multiple groups of PANs, multiple groups of MSs, and multiple groups of upsampled MSs are sequentially input into multiple dual-stream fusion layers for encoding, information interaction fusion, and decoding;
[0045] Multiple groups of PANs, multiple groups of MSs, multiple groups of upsampled MSs, and multiple groups of target PANs and multiple groups of target MSs obtained after decoding are input into the next dual-stream fusion layer until a multi-data source remote sensing image spatial-spectral fusion network is obtained.
[0046] According to a multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, each dual-stream fusion layer includes a first encoding module, a first decoding module, a U-shaped multi-channel splicing module, an information interaction fusion module, and a second decoding module, and the training module is specifically used for:
[0047] In each dual-stream fusion layer, multiple groups of PANs are input into the first encoding module for encoding to obtain first feature information;
[0048] Inputting the multiple groups of MSs and the multiple groups of upsampled MSs into a U-shaped multi-channel splicing module for encoding to obtain second feature information;
[0049] Interactively fusing the first feature information and the second feature information in an information interaction fusion module to obtain third feature information;
[0050] Decoding the multiple sets of PANs and the first feature information in a first decoding module to obtain a target PAN;
[0051] The third characteristic information is decoded in the second decoding module to obtain the target MS.
[0052] According to the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the training module is specifically used for:
[0053] The third feature information is obtained by interactively fusing the first feature information and the second feature information using the following formula:
[0054] ;
[0055] ;
[0056] ;
[0057] ;
[0058] in, and Respectively expressed in The first feature information and the second feature information extracted from the two-stream fusion layer, S represents the Sigmoid function, represents element-wise multiplication, Indicates the The third feature information obtained from the two-stream fusion layer.
[0059] According to the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the training module is further used for:
[0060] Gradient guidance of multiple groups of PANs.
[0061] According to the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the training module is specifically used for:
[0062] The Laplace operator is used to perform gradient guidance on multiple groups of PANs using the following formula:
[0063] ;
[0064] ;
[0065] ;
[0066] in, Indicates PAN.
[0067] According to a multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, in a training module, a multi-data source remote sensing image spatial-spectral fusion network is obtained through self-supervision training based on a target loss function, the target loss function is constructed based on a spatial metric function and a spectral metric function, the spatial metric function is used to evaluate the difference between a first residual probability distribution and a second residual probability distribution, the first residual probability distribution is determined based on multiple groups of PANs and target fused images, and the second residual probability distribution is determined based on multiple groups of MSs and grayscale images corresponding to the multispectral image MS;
[0068] The spectral metric function is used to evaluate the spectral angle differences between multiple groups of MS and the target fusion image.
[0069] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, any of the above-mentioned methods for spatial-spectral fusion of remote sensing images from multiple data sources is implemented.
[0070] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for spatial-spectral fusion of remote sensing images from multiple data sources is implemented.
[0071] The present invention also provides a computer program product, comprising a computer program, which implements any of the above-mentioned methods for spatial-spectral fusion of remote sensing images from multiple data sources when executed by a processor.
[0072] The multi-data source remote sensing image spatial-spectral fusion method provided by the present invention first obtains an original image for image spatial-spectral fusion, then inputs the original image into a multi-data source remote sensing image spatial-spectral fusion network to obtain a fused image. The training process of the multi-data source remote sensing image spatial-spectral fusion network is as follows: first, multiple sets of panchromatic images PAN and multiple sets of multispectral images MS of different specifications are obtained through multiple sensor groups, and then the multiple sets of PAN and MS are input into an initial multi-data source remote sensing image spatial-spectral fusion network to obtain a multi-data source remote sensing image spatial-spectral fusion network. Using the method provided by the embodiment of the present invention, spatial features and spectral features can be interactively fused based on the characteristics of the original image using a trained multi-data source remote sensing image spatial-spectral fusion network, thereby more accurately and specifically expressing and reconstructing features, improving the quality of satellite images output by the multi-data source remote sensing image spatial-spectral fusion network. At the same time, by training the network using multiple sets of data from different sensors, the richness of samples for training the network is increased, thereby improving the usability and reliability in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0074] Figure 1 It is a flow chart of the multi-data source remote sensing image spatial-spectral fusion method provided by the present invention.
[0075] Figure 2 It is a flow chart of the multi-data source remote sensing image spatial-spectral fusion network training method provided by the present invention.
[0076] Figure 3 This is a schematic diagram of the operator structure and calculation principle provided by the present invention.
[0077] Figure 4 It is a structural diagram of the multi-data source remote sensing image space-spectrum fusion network provided by the present invention.
[0078] Figure 5 It is a structural schematic diagram of the U-shaped multi-channel splicing module provided by the present invention.
[0079] Figure 6 This is a schematic diagram of the principle of the pixel reorganization upsampling process provided by the present invention.
[0080] Figure 7 Schematic diagram of the structure of the efficient channel attention module provided by the present invention.
[0081] Figure 8 It is a structural diagram of the information interaction and fusion module provided by the present invention.
[0082] Figure 9 It is a schematic diagram of the construction principle of the target loss function provided by the present invention.
[0083] Figure 10 It is a structural schematic diagram of the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention.
[0084] Figure 11 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0085] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0086] The current remote sensing image pansharpening technology mainly includes traditional fusion methods and deep learning fusion methods. The traditional fusion methods mainly include component substitution method, multi-resolution analysis method and variational optimization method. Among them, the CS and MRA methods usually consist of two parts: one is to extract spatial details from the PAN image, and the other is to inject the extracted information into the up-sampled MS image. Intensity-Hue-Saturation (IHS) and Principal Component Analysis (PCA) methods are the initial CS algorithms. Adaptive GS (GSA) realizes the fusion of PAN and MS information through the combination of a guided filter and a Gram-Schmidt (GS) transform. Then, by modeling the pixel values of the PAN and MS channels, Partial Replacement Adaptive CS (PRACS) is proposed. However, most of the current traditional methods are linear models, which are difficult to effectively fit the fusion process between the two images, and are prone to spatial deformation and spectral distortion problems. The general supervised deep learning method lacks ideal fusion results as reference labels, and needs to make a simulated data set according to the Wald protocol, which is a tedious process and easy to cause the loss of spatial information of the panchromatic image. Moreover, the scale difference between the simulated data and the real data is ignored in the fusion process, which often causes a certain degree of information distortion. In addition, the general single-stream network does not fully consider the information characteristics of MS and PAN and directly splices them, which causes the spectral and spatial features to be mixed and learned together. The double-stream network lacks the interaction between spatial and spectral features, and the feature extraction method is also relatively single, which is not conducive to feature extraction and reconstruction. It is difficult to achieve a good balance between the spectral quality and spatial quality of the fusion results. At present, in the field of pansharpening, the training and testing of the model are usually limited to data of a single specific sensor type, such as training and testing using only GF-2 data. This way of data is relatively single, and the trained network model is difficult to have universality.
[0087] Therefore, there is an urgent need for a multi-data-source remote sensing image pansharpening method to solve the above problems.
[0088] The following will be combined Figures 1-11 The multi-data-source remote sensing image pansharpening method, device, equipment and medium of the present application are described.
[0089] Figure 1 is one of the flowcharts of the multi-data-source remote sensing image pansharpening method provided by the present application, as Figure 1 shown, the method comprises the following steps:
[0090] Step 101, obtaining an original image for image hyperspectral fusion.
[0091] In specific implementation, the original image to be fused is obtained, and the original image usually includes a PAN band image and n MS band images.
[0092] Step 102, inputting the original image into a multi-data-source remote sensing image hyperspectral fusion network to obtain a fused image.
[0093] In specific implementation, the original image is input into the trained model, and the output fusion result is an n-band fused image.
[0094] By using the multi-data-source remote sensing image hyperspectral fusion method provided in the embodiments of the present application,
[0095] Figure 2 is one of the flowcharts of the multi-data-source remote sensing image hyperspectral fusion training method provided in the present application, as shown in the figure, the method includes the following: Figure 2
[0096] Step 201, obtaining multiple sets of panchromatic images PAN of different specifications and multiple sets of multispectral images MS of different specifications through multiple sensors.
[0097] In specific implementation, the PAN and MS are obtained from multiple sensors, aiming to train a multi-data-source remote sensing image hyperspectral fusion network with high quality and high robustness, and then obtain a high-quality satellite image. Specifically, the basis for selection is that the image data should cover different geographical regions, different seasonal times, and different climate conditions as much as possible. Compared with using one kind of satellite data, the information amount can be effectively increased, because each kind of satellite data provides different geographical ranges, different perspectives, and different time changes of ground features, more comprehensive and rich information can be obtained, which helps to improve the performance of the model. It can also reduce the risk of overfitting. When only a single data is used, the model may be too dependent on the specific features of the data, leading to overfitting. Using multiple data together for training can reduce this risk, because the model needs to consider multiple features and data distributions at the same time. At the same time, it can also enhance the robustness, and multiple data combinations can help the model better cope with data noise, because different data sources may be affected by different types of interference or noise, and combining multiple data can improve the robustness and anti-interference of the model.
[0098] In one example, a high-resolution sensor group is used to obtain raw PAN and MS, such as the GaoFen-1 (GF-1), GaoFen-1B (GF-1B), GaoFen-1C (GF-1C), GaoFen-1D (GF-1D), and GaoFen-6 (GF-6) satellites. In specific use, the satellite data of GF-1, GF-1B, GF-1C, GF-1D, and GF-6 can be mixed in a ratio of 1:1:1:1:1 to train the fusion model, or other ratios can be used for training, which is not limited by the embodiments of the present application. By using data from different satellites for training, the richness of the samples can be increased, so that the model can better understand and process the image features of different satellites. Further, the trained network model will be able to more effectively improve the quality of GaoFen satellite images through fusion technology, and improve its usability and reliability in practical applications.
[0099] In this step, the original PAN can also be gradient guided and the MS can be up-sampled. Specifically, the gradient information of the original PAN is extracted through high-pass filtering, and this gradient information is added as prior knowledge to the training of the network, thereby guiding the learning direction of the network in the feature space. When processing spectral information, the original MS retains the integrity of the spectral data, providing more detailed information. The up-sampled MS image can obtain larger scale spectral information by increasing the resolution of the image, so in this step, the MS is up-sampled to retain detailed information while obtaining larger scale spectral information.
[0100] For gradient guidance, high-pass filtering is used to convolve the original image with a specific high-pass filter kernel to effectively extract the high-frequency spatial information of the image. This high-pass filter kernel is commonly referred to as an operator, and commonly used high-pass filter operators include the Laplace operator, the Sobel operator, and the Prewitt detection operator. These three operators can effectively enhance the high-frequency information in the image, thereby highlighting the edge details of the image, and play an important role in image spatial feature extraction, edge detection, and image enhancement tasks.
[0101] Since the Laplace operator can fully extract the high-frequency spatial details of the image. Therefore, the embodiment of the present disclosure preferably uses the Laplace operator to perform gradient filtering on the original PAN image to obtain the gradient map of the original image, and uses it as a priori condition to enhance the network's ability to extract spatial information. Of course, other forms of operators can also be selected for gradient filtering, and the embodiment of the present disclosure does not limit this. The operator structure used and the calculation process are as follows: Figure 3 As shown in , the Laplace operator has a negative center and is surrounded by 1s and 0s. It is a linear operator. It highlights the high-frequency changes in the image by performing second-order differentials in both the horizontal and vertical directions, thereby enhancing the edge information in the image. Assuming the image to be filtered is , its expression is:
[0102] ;
[0103] ;
[0104] .
[0105] This filter kernel can highlight edges and details in an image because the convolution operation produces a larger response value when the pixel values in the image change rapidly. During the filtering process, the Laplacian high-pass filter suppresses low-frequency components in the image and emphasizes high-frequency components. As a result, image details and edges become clearer and more prominent, while reducing the intensity of smooth areas in the image.
[0106] By extracting prior gradient information from the PAN image using Laplace high-pass filtering, the network can be effectively guided to focus on the spatial high-frequency details of the image, thereby improving the network's perception of image edges and details, and helping to produce clearer and more accurate fusion results. Secondly, guided by the prior gradient information, the network can more accurately learn the spatial features in the image, allowing the network to better retain and enhance the structural information of the image during the fusion process, thereby improving spatial extraction capabilities. In addition, the prior gradient information extracted using Laplace high-pass filtering can serve as an additional constraint to help the network more stably learn the correlation between images, thereby improving the network's generalization and robustness. Therefore, the gradient guidance strategy can effectively improve network performance and spatial extraction capabilities, thereby producing higher-quality fusion results.
[0107] Step 202: Input multiple groups of PANs and multiple groups of MSs into a multi-data source remote sensing image spatial-spectral fusion network to obtain a multi-data source remote sensing image spatial-spectral fusion network.
[0108] In specific implementation, multiple groups of PANs and multiple groups of MSs are input into Figure 4In the illustrated multi-data-source remote sensing image space spectrum fusion network, the multi-data-source remote sensing image space spectrum fusion network comprises a plurality of double-flow fusion layers connected in turn, each double-flow fusion layer comprising a first encoding module, a first decoding module, a U-shaped multi-channel splicing module, an information interaction fusion module and a second decoding module. The U-shaped multi-channel splicing module can input the original MS and the up-sampled MS together, integrate the spectral information of different scale resolutions, and complete channel splicing, thereby fully retaining the spectral information of the input MS image. During up-sampling, the up-sampling can be 4 times or 8 times, and the specific ratio can be set as required, which is not limited in the embodiment of the application. In addition, the feature transmission in the same level is performed through the jump connection, which helps to reduce the loss of information in the network transmission process and ensures that the network retains and utilizes important spectral information better in the fusion task. The structure of the U-shaped multi-channel splicing module is as shown in Figure 5 As shown, the pixel reorganization (Pixel Shuffle) operation is adopted to up-sample the low-scale spectral features to the same size as the high-resolution features by pixel rearrangement. The pixel reorganization up-sampling process is as shown in Figure 6 As an example of the up-sampling ratio being 4, the pixel reorganization realizes up-sampling by rearranging pixel values. Compared with the traditional interpolation method, the calculation efficiency is higher, so that the feature up-sampling operation in the U-shaped multi-channel splicing module can be performed more quickly. In addition, the pixel reorganization does not have additional parameters in the up-sampling process, so it will not introduce additional information loss or distortion, nor will it increase the complexity and computational burden of the network. It can effectively retain the information in the low-resolution feature map and rearrange it into a high-resolution feature map, avoiding information loss.
[0109] For the U-shaped multi-channel splicing module, the pooling operation that may cause scale difference and information loss is discarded, and the up-sampling operation is adopted to construct a unique U-shaped network, retaining multi-scale features. Through the combination of multi-scale information, it is helpful to improve the perception ability of the network to spectral features and the fusion quality.
[0110] In addition, the channel attention mechanism is beneficial to learning the dependency relationship between bands and can fully retain the spectral information of the MS image in the fusion process. In order to further optimize the attention mechanism, an efficient channel attention module (Efficient Channel Attention, ECA) is introduced. The structure diagram is as shown in Figure 7As shown in the figure, the efficient channel attention module obtains aggregated features for each channel through global average pooling. It then uses fast one-dimensional convolution to calculate k feature points near each channel feature point and regenerates channel weights to capture local cross-channel interaction information. While ensuring network computational efficiency, this module helps improve the discrimination and correlation between channel features through weight redistribution, thereby reducing spectral distortion in the fusion result. In the U-shaped multi-channel splicing module, the spectral features obtained by splicing different levels are re-aggregated using the ECA module to enhance the closeness between features.
[0111] For the information interaction fusion module, the structure is as follows Figure 8 As shown in Figure 2, this module is more conducive to adaptively selecting the features of panchromatic and multispectral images for fusion, and plays a better interactive fusion role. In order to focus on more important information, the module first reconstructs the spatial and spectral features of the input, and its expression is:
[0112] ; ; ;
[0113] in and Respectively represent The features of PAN and MS extracted from the layer network, S represents the Sigmoid function, Represents element-wise multiplication. By introducing the attention mechanism, it is possible to automatically learn the relationship between input features and dynamically adjust the weights according to their importance, thereby increasing the network's attention to important information. This mechanism helps solve the problems of information imbalance and information redundancy. Secondly, a 1×1 convolution kernel is used in this module. While keeping the size of the feature map unchanged, the depth of the feature map can be adjusted, thereby increasing the nonlinear expression ability of the network. Compared with ordinary 3×3 convolution, it can effectively reduce the number of parameters, while reducing computational complexity, accelerating the network training and feature transfer process. Finally, this module interactively fuses the two reconstructed features to enhance the fusion effect of the network. Its expression is:
[0114] .
[0115] in, Indicates the The fused information obtained by the interactive fusion module is then fed into the feature restoration module for reconstruction. The interactive fusion module combines multiple techniques, such as the attention mechanism and low-parameter convolution, to effectively extract and fuse the features of PAN and MS images, thereby improving the effectiveness and performance of spatial-spectral fusion.
[0116] In this embodiment, the multi-data source remote sensing image spatial-spectral fusion network is obtained by self-supervised training based on a target loss function, which is constructed based on a spatial metric function and a spectral metric function. The spatial metric function is used to evaluate the difference between the first residual probability distribution and the second residual probability distribution. The first residual probability distribution is determined based on multiple groups of PAN and target fused images, and the second residual probability distribution is determined based on multiple groups of MS and grayscale images corresponding to the multispectral image MS; the spectral metric function is used to evaluate the spectral angle difference between the multiple groups of MS and the target fused image, such as Figure 9 As shown in the figure, up-MS represents the upsampled MS image, and stack represents replicating the PAN image along the channel dimension to ensure that the size is consistent with the fusion result. The constraints and implementation methods of the function are introduced below.
[0117] SAM is a spectral evaluation index function that can be used to represent the spectral angle difference between HRMS and MS, and measure the spectral loss between the two. It can be expressed by the following formula: ;
[0118] Where the SAM value calculated in the formula is the spectral angle of a single pixel point in the image, <·,·> is the vector inner product symbol, ||·||2 represents the l2 norm, arccos(·) represents the inverse cosine function, z1∈R 1×1×c , z2∈R 1×1×c Represents the spectral vector at the specified pixel coordinates. KL divergence is an indicator function used to calculate the similarity. The closer the two probability distributions are, the smaller the KL divergence. It can be expressed by the following formula:
[0119] ;
[0120] Here, p(x) is the probability distribution function of the true information, and q(x) is the probability distribution function of the fitted information. By calculating the information entropy of the two probability distributions, the difference in information is quantitatively measured. Spectral metric function (or spectral loss function): The upsampled MS and the original MS can be considered to contain the same spectral features. The spectral features are mapped to the vector direction between the image bands. Therefore, by minimizing the SAM value between the HRMS and the upsampled MS, the spectral loss of the network can be controlled. This spectral loss function can be expressed as follows:
[0121] ;
[0122] Where SAM(·) represents the spectral angle mapping function, M↑ represents the image obtained by upsampling the multispectral image MS, Represents the target fusion image. Spatial metric function (or spatial loss function): Unlike the process of establishing spectral constraints, it is impossible to directly establish the connection between the predicted result HRMS and the original PAN image. The spatial loss can be established based on three conventions: (1) The PAN image can be degraded from the MS image under the same spatial resolution conditions; (2) The degraded grayscale image contains the spatial texture information of the source image; (3) The differences between the MS image and the PAN image at different spatial resolutions should have similar distributions. The MS is degraded to a grayscale image, and the softmax function is used to establish the residual probability distribution map between the MS image and the grayscale image. The residual probability distribution between HRMS and PAN is then fitted, and the KL divergence value between the two is minimized to achieve the purpose of constraining the spatial loss, thereby realizing unsupervised training without the need for manually labeled data.
[0123] The multi-data source remote sensing image spatial-spectral fusion device provided by the present invention is described below. The multi-data source remote sensing image spatial-spectral fusion device described below and the multi-data source remote sensing image spatial-spectral fusion method described above can refer to each other.
[0124] By using the method provided by the embodiment of the present disclosure, a self-supervised fusion framework is constructed, and the original image to be fused itself is used as a label. The training of the model is based on real data that does not require any downscaling preprocessing, thereby bridging the gap between simulated data and real data, and ensuring that the fusion result is more consistent with the actual scene. In terms of spatial flow, the present invention specifically introduces gradient guidance, which provides the network with prior knowledge of spatial information by capturing gradient information, thereby accurately guiding the optimization direction of the network. In the spectral flow, a U-shaped multi-channel splicing module is proposed, which inputs the original multi-spectral and upsampled multi-spectral images into the network together, thereby more comprehensively processing spectral information of different scales. Subsequently, by designing an interactive fusion module with an attention mechanism, the network can adaptively fuse spatial and spectral information.
[0125] In terms of data acquisition, a joint training and testing dataset containing multiple satellite data types was constructed. Its rich feature information provides strong data support for deep learning model training, effectively improving the model's fusion performance in various scenarios. Training with multiple different satellite data types enriches the types of training samples, helping to explore the fusion potential of deep learning methods and avoid network overfitting caused by a single data type. This also alleviates the current lack of universal performance of deep learning methods, achieving the effect of "train once, apply everywhere."
[0126] like Figure 10 As shown, the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention includes the following modules:
[0127] An acquisition module 1001 is used to acquire an original image for image spatial-spectral fusion;
[0128] Fusion module 1002, used for inputting the original image into the multi-data source remote sensing image spatial-spectral fusion network to obtain a fused image;
[0129] A training module 1003 is configured to acquire multiple sets of panchromatic images PAN and multiple sets of multispectral images MS of different specifications through multiple sensor groups;
[0130] The training module 1003 is further configured to input multiple groups of PANs and multiple groups of MSs into a multi-data source remote sensing image spatial-spectral fusion network to obtain a multi-data source remote sensing image spatial-spectral fusion network.
[0131] According to a multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the multi-data source remote sensing image spatial-spectral fusion network includes a plurality of dual-stream fusion layers connected in sequence, and the training module 1003 is specifically used for:
[0132] Upsampling the multiple groups of MSs to obtain multiple groups of upsampled MSs;
[0133] Multiple groups of PANs, multiple groups of MSs, and multiple groups of upsampled MSs are sequentially input into multiple dual-stream fusion layers for encoding, information interaction fusion, and decoding;
[0134] Multiple groups of PANs, multiple groups of MSs, multiple groups of upsampled MSs, and multiple groups of target PANs and multiple groups of target MSs obtained after decoding are input into the next dual-stream fusion layer until the multi-data source remote sensing image spatial-spectral fusion network.
[0135] According to a multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, each dual-stream fusion layer includes a first encoding module, a first decoding module, a U-shaped multi-channel splicing module, an information interaction fusion module, and a second decoding module. The training module 1003 is specifically used to:
[0136] In each dual-stream fusion layer, multiple groups of PANs are input into the first encoding module for encoding to obtain first feature information;
[0137] Inputting the multiple groups of MSs and the multiple groups of upsampled MSs into a U-shaped multi-channel splicing module for encoding to obtain second feature information;
[0138] Interactively fusing the first feature information and the second feature information in an information interaction fusion module to obtain third feature information;
[0139] Decoding the multiple sets of PANs and the first feature information in a first decoding module to obtain a target PAN;
[0140] The third characteristic information is decoded in the second decoding module to obtain the target MS.
[0141] According to the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the training module 1003 is specifically used for:
[0142] The third feature information is obtained by interactively fusing the first feature information and the second feature information using the following formula:
[0143] ;
[0144] ;
[0145] ;
[0146] ;
[0147] in, and Respectively expressed in The first feature information and the second feature information extracted from the two-stream fusion layer, S represents the Sigmoid function, represents element-wise multiplication, Indicates the The third feature information obtained from the two-stream fusion layer.
[0148] According to the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the training module 1003 is further used for:
[0149] Gradient guidance of multiple groups of PANs.
[0150] According to the multi-data source remote sensing image spatial-spectral fusion device provided by the present invention, the training module 1003 is specifically used for:
[0151] The Laplace operator is used to perform gradient guidance on multiple groups of PANs using the following formula:
[0152] ;
[0153] ;
[0154] ;
[0155] in, Indicates PAN.
[0156] According to the multi-data-source remote sensing image space-spectrum fusion device provided by the application, in the training module 1003, the multi-data-source remote sensing image space-spectrum fusion network is obtained through self-supervised training based on a target loss function, and the target loss function is constructed based on a space measurement function and a spectrum measurement function. The space measurement function is used to evaluate the difference between the first residual probability distribution and the second residual probability distribution, the first residual probability distribution is determined based on a plurality of PAN and target fusion images, and the second residual probability distribution is determined based on a plurality of MS and a plurality of corresponding gray images of the multispectral image MS.
[0157] The spectrum measurement function is used to evaluate the spectral angle difference between the plurality of MS and the target fusion image.
[0158] Figure 11 An example of an entity structure diagram of an electronic device is shown in Figure 11 As shown, the electronic device can include a processor 1110, a communications interface 1120, a memory 1130, and a communications bus 1140, wherein the processor 1110, the communications interface 1120, and the memory 1130 communicate with each other through the communications bus 1140. The processor 1110 can invoke the logical instructions in the memory 1130 to execute the multi-data-source remote sensing image space-spectrum fusion method, which includes:
[0159] Obtaining an original image for image space-spectrum fusion;
[0160] Inputting the original image into a multi-data-source remote sensing image space-spectrum fusion network to obtain a fusion image;
[0161] Wherein, the multi-data-source remote sensing image space-spectrum fusion is trained by the following method:
[0162] Obtaining a plurality of different specifications of PAN and a plurality of different specifications of MS through a plurality of sensors;
[0163] Inputting the plurality of PAN and the plurality of MS into an initial multi-data-source remote sensing image space-spectrum fusion network for training to obtain the multi-data-source remote sensing image space-spectrum fusion network.
[0164] Further, the logic instructions in the memory 1130 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0165] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the multi-data-source remote sensing image hyperspectral fusion method provided by the above-mentioned methods, the method comprising:
[0166] obtaining original images for image hyperspectral fusion;
[0167] inputting the original images into a multi-data-source remote sensing image hyperspectral fusion network to obtain a fusion image;
[0168] wherein the multi-data-source remote sensing image hyperspectral fusion is trained by the following method:
[0169] obtaining multiple groups of panchromatic images PAN of different specifications and multiple groups of multispectral images MS of different specifications through multiple sensor groups;
[0170] inputting the multiple groups of PAN and the multiple groups of MS into an initial multi-data-source remote sensing image hyperspectral fusion network for training to obtain the multi-data-source remote sensing image hyperspectral fusion network.
[0171] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, the computer program is executed by a processor to implement the multi-data-source remote sensing image hyperspectral fusion method provided by the above-mentioned methods, the method comprising:
[0172] obtaining original images for image hyperspectral fusion;
[0173] inputting the original images into a multi-data-source remote sensing image hyperspectral fusion network to obtain a fusion image;
[0174] The multi-data-source remote sensing image space-spectrum fusion is trained by the following method:
[0175] Different specifications of multiple sets of panchromatic images PAN and different specifications of multiple sets of multispectral images MS are acquired through multiple sensor groups.
[0176] The multiple sets of PAN and the multiple sets of MS are input into the initial multi-data-source remote sensing image space-spectrum fusion network for training, to obtain the multi-data-source remote sensing image space-spectrum fusion network.
[0177] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course, can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0179] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for spatial-spectral fusion of remote sensing images from multiple data sources, characterized in that: The method comprises: Obtain the original image for image spatial-spectral fusion; Inputting the original image into a multi-data source remote sensing image spatial-spectral fusion network to obtain a fused image; The spatial-spectral fusion of remote sensing images from multiple data sources is trained by the following method: Acquire multiple sets of panchromatic images PAN and multiple sets of multispectral images MS of different specifications by using multiple sensor groups, wherein each of the panchromatic image PAN and each of the multispectral images MS covers different geographical areas, different seasons and different climatic conditions; Inputting the multiple groups of PANs and the multiple groups of MSs into an initial multi-data source remote sensing image spatial-spectral fusion network for training to obtain the multi-data source remote sensing image spatial-spectral fusion network; The multi-data source remote sensing image spatial-spectral fusion network includes a plurality of dual-stream fusion layers connected in sequence; Each of the dual-stream fusion layers includes a first encoding module, a first decoding module, a U-shaped multi-channel splicing module, an information interaction fusion module, and a second decoding module, and the multiple groups of PANs, the multiple groups of MSs, and the multiple groups of upsampled MSs are sequentially input into the multiple dual-stream fusion layers for encoding, information interaction fusion, and decoding, including: In each dual-stream fusion layer, the multiple groups of PANs are input into the first encoding module for encoding to obtain first feature information; Inputting the multiple groups of MSs and the multiple groups of upsampled MSs into the U-shaped multi-channel splicing module for encoding to obtain second feature information; The first feature information and the second feature information are interactively integrated using the following formula to obtain the third feature information: ; ; ; ; in, and Respectively expressed in The first feature information and the second feature information extracted from the two-stream fusion layer, S represents the Sigmoid function, represents element-wise multiplication, Indicates the The third feature information obtained in the two-stream fusion layer; Decoding the multiple groups of PANs and the first feature information in the first decoding module to obtain a target PAN; The third characteristic information is decoded in the second decoding module to obtain a target MS.
2. The method for spatial-spectral fusion of remote sensing images from multiple data sources according to claim 1, characterized in that: Inputting the multiple groups of PANs and the multiple groups of MSs into an initial multi-data source remote sensing image spatial-spectral fusion network for training to obtain the multi-data source remote sensing image spatial-spectral fusion network includes: Upsampling the multiple groups of MSs to obtain multiple groups of upsampled MSs; Inputting the multiple groups of PANs, the multiple groups of MSs and the multiple groups of upsampled MSs into the multiple dual-stream fusion layers in sequence for encoding, information interaction fusion and decoding; The multiple groups of PANs, the multiple groups of MSs and the multiple groups of upsampled MSs, as well as the multiple groups of target PANs and multiple groups of target MSs obtained after decoding are input into the next dual-stream fusion layer until the multi-data source remote sensing image spatial-spectral fusion network is obtained.
3. The method for spatial-spectral fusion of remote sensing images from multiple data sources according to any one of claims 1 to 2, characterized in that: After acquiring multiple groups of panchromatic images PAN and multiple groups of multispectral images MS of different specifications by using multiple sensor groups, the method further includes: Gradient guidance is performed on the multiple groups of PANs.
4. The method for spatial-spectral fusion of remote sensing images from multiple data sources according to claim 3, characterized in that: The step of gradient-guiding the plurality of PAN groups includes: The Laplace operator is used to perform gradient guidance on the multiple groups of PANs using the following formula: ; ; ; in, Indicates the PAN.
5. The method for spatial-spectral fusion of remote sensing images from multiple data sources according to claim 4, characterized in that: The multi-data source remote sensing image spatial-spectral fusion network is obtained by self-supervised training based on a target loss function, the target loss function is constructed based on a spatial metric function and a spectral metric function, the spatial metric function is used to evaluate the difference between a first residual probability distribution and a second residual probability distribution, the first residual probability distribution is determined based on the multiple groups of PANs and target fused images, and the second residual probability distribution is determined based on the multiple groups of MSs and grayscale images corresponding to the multispectral image MS; The spectral metric function is used to evaluate the spectral angle difference between the multiple groups of MSs and the target fused image.
6. A multi-data source remote sensing image spatial-spectral fusion device, characterized in that: include: An acquisition module, used for acquiring the original image for image spatial-spectral fusion; A fusion module, configured to input the original image into a multi-data source remote sensing image spatial-spectral fusion network to obtain a fused image; a training module for acquiring, through a plurality of sensor groups, a plurality of sets of panchromatic images PAN and a plurality of sets of multispectral images MS of different specifications, wherein each of the panchromatic image PAN and each of the multispectral images MS covers a different geographical area, a different season and a different climatic condition; The training module is further configured to input the multiple groups of PANs and the multiple groups of MSs into a multi-data source remote sensing image spatial-spectral fusion network to obtain the multi-data source remote sensing image spatial-spectral fusion network; The multi-data source remote sensing image spatial-spectral fusion network includes multiple dual-stream fusion layers connected in sequence; each of the dual-stream fusion layers includes a first encoding module, a first decoding module, a U-shaped multi-channel splicing module, an information interaction fusion module, and a second decoding module. The training module sequentially inputs the multiple groups of PANs, the multiple groups of MSs, and the multiple groups of upsampled MSs into the multiple dual-stream fusion layers for encoding, information interaction fusion, and decoding, including: In each dual-stream fusion layer, the multiple groups of PANs are input into the first encoding module for encoding to obtain first feature information; Inputting the multiple groups of MSs and the multiple groups of upsampled MSs into the U-shaped multi-channel splicing module for encoding to obtain second feature information; The first feature information and the second feature information are interactively integrated using the following formula to obtain the third feature information: ; ; ; ; in, and Respectively expressed in The first feature information and the second feature information extracted from the two-stream fusion layer, S represents the Sigmoid function, represents element-wise multiplication, Indicates the The third feature information obtained in the two-stream fusion layer; Decoding the multiple groups of PANs and the first feature information in the first decoding module to obtain a target PAN; The third characteristic information is decoded in the second decoding module to obtain a target MS.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the multi-data source remote sensing image spatial-spectral fusion method as described in any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for spatial-spectral fusion of remote sensing images from multiple data sources as claimed in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Optical remote sensing spatial spectrum fusion method, device and equipment without reference image and medium
CN114581347A
Ship level identification technology based on high-resolution optical satellite remote sensing image under condition of few samples
CN115471759A
Remote sensing image space-spectrum fusion method and device, electronic equipment and storage medium
CN117079105A
Double-encoder pavement crack detection method based on Sobel operator bridging
CN117689897A