Remote sensing multispectral image panchromatic fusion method and device based on auto-encoder
Through the pixel-level mask autoencoder model of the autoencoder, combined with the cross-attention mechanism and self-attention mechanism, the problems of insufficient fusion accuracy and generalization ability in remote sensing multispectral image processing are solved, and high-quality reconstruction of high-resolution multispectral images and improved computational efficiency are achieved.
Patent Information
- Application Number
- CN202510624247.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-10-17
AI Technical Summary
Existing remote sensing multispectral image processing technology has problems such as insufficient image fusion accuracy and insufficient model generalization ability when processing complex scenes, especially in the nonlinear mapping processing of high spatial resolution and spectral resolution, which is prone to distortion.
A pixel-level mask autoencoder (PEMAE) model based on autoencoders is adopted, combined with the cross-attention mechanism and the self-attention mechanism. Multiple low-resolution multispectral image versions are generated through pixel-level masks, and feature extraction and reconstruction are performed. Finally, a high-resolution multispectral image is generated through scattering and integration.
It significantly improves the accuracy of image fusion and the generalization application capability of the model, can retain spectral information and spatial details with high quality in complex scenes, reduce computational complexity, and enhance the adaptability and robustness of the model.
Smart Images

Figure CN120807303A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of remote sensing image processing, and particularly relates to a remote sensing multispectral image panchromatic fusion method and device based on an autoencoder. BACKGROUND
[0002] Multispectral image fusion technology aims to combine low spatial resolution multispectral images and high spatial resolution panchromatic images to generate multispectral images with high spatial resolution. Traditional panchromatic sharpening methods rely on prior knowledge to build models and are widely used in practical applications, but have significant limitations when dealing with highly nonlinear mapping problems.
[0003] For example, methods based on component replacement and multi-resolution analysis are limited in their nonlinear capabilities of feature representation and are prone to spatial or spectral distortion. In addition, the assumptions of these methods deviate from the actual situation at the physical level of remote sensing.
[0004] Therefore, the traditional technology has certain limitations when processing remote sensing multispectral images, and it is difficult to meet the high-quality image fusion requirements in complex scenarios. SUMMARY
[0005] The present application aims to propose a remote sensing multispectral image sharpening method based on autoencoder technology and a matching device to overcome the limitations of existing remote sensing multispectral image processing technology and significantly improve the accuracy of image fusion and the generalization application ability of the model.
[0006] The present application provides a remote sensing multispectral image sharpening method based on an autoencoder, comprising the following steps.
[0007] Obtain a low-resolution multispectral image and a high-resolution panchromatic image of a target geographic location; input the low-resolution multispectral image and the high-resolution panchromatic image into a trained pixel-level integrated mask autoencoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location, wherein it includes: performing pixel-level masking on the low-resolution multispectral image to obtain multiple masked versions of the low-resolution multispectral image; based on the multiple masked versions of the low-resolution multispectral image and the high-resolution panchromatic image, performing feature extraction and reconstruction according to a cross-attention mechanism and a self-attention mechanism to obtain multiple high-resolution reconstruction results; based on a pre-set relationship modeling, performing scattering and integration on the multiple high-resolution reconstruction results to obtain a reconstructed high-resolution multispectral image.
[0008] According to a remote sensing multispectral image sharpening method based on an autoencoder provided by the present invention, the method further includes: obtaining a training set of data sample pairs from different satellites, wherein the data sample pair training set includes: simulated low-resolution multispectral image samples, high-resolution panchromatic image samples, and high-resolution multispectral image labels; using a target loss function, a preset pixel-level integrated mask autoencoder model is trained based on the data sample pair training set to obtain a trained mask autoencoder; wherein the target loss function includes: mean square error and mean absolute error.
[0009] According to a remote sensing multispectral image sharpening method based on an autoencoder provided by the present invention, the method of obtaining a training set of data sample pairs from different satellites includes: obtaining low-resolution multispectral image samples and high-resolution panchromatic image samples with the same geographic coordinates; performing an element-by-element masking operation on the low-resolution multispectral image samples to obtain simulated low-resolution multispectral image samples; using the simulated low-resolution multispectral image samples and the high-resolution panchromatic image samples as data sample pairs; and using the low-resolution multispectral image samples as high-resolution multispectral image labels.
[0010] According to a remote sensing multispectral image sharpening method based on an autoencoder provided by the present invention, the target loss function includes: in, represents the target loss function, MAE represents the mean absolute error, and MSE represents the mean square error, where: in, represents the mean absolute error, represents the mean square error, The first The first band Rank The actual value of the column pixel; Represents the first output of the preset pixel-level integrated mask autoencoder model The first band Rank The predicted value of the column pixel; Indicates the number of image bands; Indicates the image height; Indicates the image width; is an adjustable weight parameter used to balance and The proportion of contribution to the total loss.
[0011] According to the self-encoder-based remote sensing multispectral image sharpening method provided by the application, the cross attention mechanism and the self attention mechanism are used for feature extraction and reconstruction based on the plurality of mask versions of low-resolution multispectral images and the high-resolution panchromatic image, to obtain a plurality of high-resolution reconstruction results, including: The embedding of the plurality of mask versions of low-resolution multispectral images is taken as a query, and the embedding of the high-resolution panchromatic image is taken as a key and a value: Among them, The cross attention weight result is represented as, The self attention weight result is represented as, The query is represented as, The key is represented as, The value is represented as, The dimension of the key is represented as, The normalization function is represented as, The transpose matrix of the key is represented as, The focusing function is represented as, which is used for approximating the normalization function and maintaining linear time complexity.
[0012] According to the self-encoder-based remote sensing multispectral image sharpening method provided by the application, the focusing function includes: Among them, The focusing function is represented as, The input feature matrix is represented as, The direction adjustment mapping is represented as, The nonlinear activation function is represented as; Among them, The input vector is represented as, The norm of the input vector is represented as, The focusing factor is represented as.
[0013] The application further provides a self-encoder-based remote sensing multispectral image panchromatic fusion device, comprising the following modules: an acquisition module, configured to acquire a low-resolution multispectral image and a high-resolution panchromatic image of a target geographical position; and a pixel-level integrated mask automatic coding module, configured to input the low-resolution multispectral image and the high-resolution panchromatic image into a trained pixel-level integrated mask automatic encoder model to obtain a reconstructed high-resolution multispectral image of the target geographical position, wherein the pixel-level integrated mask automatic coding module comprises: performing pixel-level mask coding on the low-resolution multispectral image to obtain a plurality of low-resolution multispectral images in mask versions; performing feature extraction and reconstruction based on the plurality of low-resolution multispectral images in mask versions and the high-resolution panchromatic image according to a cross-attention mechanism and a self-attention mechanism to obtain a plurality of high-resolution reconstruction results; and performing scattering and integration on the plurality of high-resolution reconstruction results based on a preset relationship modeling to obtain the reconstructed high-resolution multispectral image.
[0014] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the self-encoder-based remote sensing multispectral image panchromatic fusion method according to any one of the above when executing the program.
[0015] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the self-encoder-based remote sensing multispectral image panchromatic fusion method according to any one of the above.
[0016] The application further provides a computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the self-encoder-based remote sensing multispectral image panchromatic fusion method according to any one of the above.
[0017] The self-encoder-based remote sensing multispectral image panchromatic fusion method and device provided by the application can acquire a low-resolution multispectral image and a high-resolution panchromatic image of a target geographical position, generate a plurality of mask versions of the low-resolution multispectral image through pixel-level mask coding, fuse the spatial detail features of the high-resolution panchromatic image and the spectral features of the multispectral image by using a cross-attention mechanism, and enhance the relevance of local and global features by using a self-attention mechanism; on this basis, the plurality of high-resolution reconstruction results are scattered and integrated by using a preset relationship modeling, and finally the spatial resolution of the multispectral image is significantly improved and the high-fidelity reconstruction of the spectral characteristics is realized. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description one by one. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0019] Figure 1 is the flowchart of the remote sensing multispectral image panchromatic fusion method based on the autoencoder provided by the present application.
[0020] Figure 2 is the visualization comparison diagram of the data distribution in the remote sensing multispectral image panchromatic fusion process using the tSNE technology.
[0021] Figure 3 is the training paradigm diagram of the network model provided by the present application.
[0022] Figure 4 is the model training process diagram provided by the present application.
[0023] Figure 5 is the scattering diagram of the low-resolution multispectral image provided by the present application.
[0024] Figure 6 is the reconstruction result instance diagram in the real scene provided by the present application.
[0025] Figure 7 is the flowchart of the cross-attention mechanism and the self-attention mechanism provided by the present application.
[0026] Figure 8 is the flowchart of the focused linear attention method provided by the present application.
[0027] Figure 9 is the structural diagram of the remote sensing multispectral image panchromatic fusion device based on the autoencoder provided by the present application.
[0028] Figure 10 is the entity structure diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0029] In the specific implementation process, in order to make the purpose, technical solutions and advantages of the present application more clear, the technical solutions of the present application will be described in detail below in combination with the drawings. It should be noted that the described embodiments are only some embodiments of the present application, not all. Based on these embodiments, other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0030] Multi-spectral image fusion technology aims to combine low spatial resolution multi-spectral images with high spatial resolution panchromatic images to generate high spatial resolution multi-spectral images. Current techniques mainly rely on convolutional neural networks and traditional observation models, such as the attention mechanism-based dual-branch convolutional neural network (AIDB-Net). This model combines convolution and self-attention mechanisms to some extent to improve sharpening effects. However, the complex structure of AIDB-Net increases the computational burden and has a high dependence on the quality and diversity of training data, limiting its application potential on resource-constrained devices.
[0031] In addition, Squeeze-and-Excitation Networks (SE-Net) enhances feature representation through channel attention mechanisms, but has limitations in processing spatial dimension information, and is not sensitive enough to changes in local spatial structures such as texture and edges. These limitations indicate that existing techniques still need to be broken through in handling complex scenes and achieving high-quality image reconstruction.
[0032] To address the above problems, the present application proposes a multi-spectral image sharpening method based on a pixel-level mask autoencoder (PEMAE). The present application introduces a pixel-level mask mechanism and an efficient linear cross-attention mechanism, effectively improving the sharpening effect and the generalization ability of the model. By generating multiple mask versions of high-resolution multi-spectral images through random masking and scattering operations, and using an ensemble strategy to fuse the results of these versions, the present application preserves spectral information while enhancing spatial details, thereby achieving high-quality multi-spectral image sharpening.
[0033] In previous panchromatic sharpening technology applications, most methods rely on established prior knowledge to build relatively stable model architectures. This type of method has been widely used in practical operational scenarios, but its limitations are exposed when dealing with complex and highly nonlinear mapping relationships. Methods based on component replacement and multi-resolution analysis often struggle with nonlinear representation of features, easily causing spatial or spectral distortions. Furthermore, from the perspective of remote sensing physics, the assumption conditions relied on by such methods have significant limitations and deviate from real-world scenarios.
[0034] While the emerging deep learning based panchromatic sharpening methods have achieved significant performance improvement by learning the optimal mapping from the observed data to the ideal fused high-resolution multispectral image from massive sample data, they have to rely on Wald's protocol to simulate the data for training due to the lack of real high-resolution multispectral image data support. Although they can achieve good results in the low-resolution evaluation scenario, their performance drops significantly in the full-resolution evaluation. At the same time, the current deep learning methods perform poorly in the promotion of new satellite data, reflecting the lack of satellite-independent ability, i.e., the domain shift problem.
[0035] Based on the above series of problems, the present application innovatively proposes a brand new observation model architecture and the matching sharpening technology. We regard the low-resolution multispectral image as the product of the pixel-level mask operation on the high-resolution multispectral image, and then use the pixel-level integrated mask autoencoder (PEMAE) to realize the effective restoration of the high-resolution multispectral image. In actual operation, PEMA can generate a large number of high-resolution multispectral image versions with mask features (which can simulate low-resolution multispectral images to some extent) through random mask and scattering operations, and then fuse the results of each version through an integrated strategy. The final output of the high-resolution multispectral image can not only cover the spectral information carried by the low-resolution multispectral image, but also integrate the fine spatial details possessed by the high-resolution panchromatic image. The present application discards the traditional self-attention mechanism and adopts a linear cross-attention mechanism, which reduces the computational complexity to linear time level, improves the model calculation efficiency, and enhances its general application ability.
[0036] Numerous experimental data and results show that PEMA has significant advantages over existing advanced methods in quantitative evaluation, visual effect presentation and other dimensions, whether in low-resolution evaluation scenarios or in full-resolution evaluation situations. The technical solution of the present application not only greatly optimizes the spatial quality and spectral quality of the fused image, but also shows excellent general application ability in actual use.
[0037] It should be noted that the remote sensing multispectral image panchromatic fusion method based on the autoencoder in the embodiments covered by the present application has flexibility in specific implementation, and can be executed by a server alone, or by a terminal device independently, or by a server and a terminal device in cooperation. Here, the server executing the remote sensing multispectral image panchromatic fusion method based on the autoencoder in the present embodiment is taken as an example for detailed description.
[0038] Figure 1is a flowchart of a remote sensing multispectral image panchromatic fusion method based on an autoencoder provided by the present application, as shown in Figure 1 The method comprises the following steps.
[0039] Step 101, obtaining a low-resolution multispectral image and a high-resolution panchromatic image of a target geographic location.
[0040] The present application innovatively proposes a remote sensing multispectral image sharpening scheme based on a pixel-level mask autoencoder (PEMAE). This scheme aims to significantly improve the accuracy and generalization application ability of image fusion through a unique observation model and an efficient calculation strategy. The method mainly processes low spatial-resolution multispectral images (LRMS) and high spatial-resolution panchromatic images (PAN), and finally outputs high-quality high spatial-resolution multispectral images (HRMS).
[0041] In some embodiments, the low-resolution multispectral image of the target area is collected by satellite or aerial remote sensing platforms such as Landsat, Sentinel-2, etc. LRMS usually covers multiple spectral bands, but the spatial resolution is relatively low. At the same time, the high-resolution panchromatic image of the same area is obtained by means of high-resolution remote sensing satellites, and PAN is generally a single band, but the spatial resolution is high.
[0042] In this process, it is necessary to ensure that the two types of images are relatively close in time (so as to avoid interference caused by surface changes) and that the geographical coordinates are accurately registered. Remote sensing data processing software can be used to carry out geometric correction, radiation correction and resampling operations, so as to eliminate the offset phenomenon caused by the difference of sensors or imaging conditions.
[0043] For example, for a data pair of a certain target geographic location, it contains a LRMS image with a size of 256x256 pixels and a PAN image with a size of 1024x1024 pixels.
[0044] Step 102, inputting the low-resolution multispectral image and the high-resolution panchromatic image into the trained pixel-level integrated mask autoencoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location.
[0045] Before inputting the low-resolution multispectral image (LRMS) and high-resolution panchromatic image (PAN) into the pixel-level ensemble mask autoencoder model, the pixel values are transformed in distribution to reduce the distribution difference between different datasets. The specific steps are as follows: Compute empirical quantiles: For the pixel values X={x1,x2,…,xn} sampled from the spectral band, compute m empirical quantiles Q={q0,q1,…,qm−1} to discretize the empirical cumulative distribution function (eCDF) of the source data.
[0046] Interpolate to estimate eCDF values: For each pixel value x, estimate its value u(x) in the eCDF by the bidirectional interpolation method to deal with the non-unique or overlapping quantile boundary problem caused by repeated values or sparse regions in the source distribution.
[0047] Apply the inverse CDF of the target distribution: Use the inverse cumulative distribution function (inverse CDF) of the target distribution (such as the standard normal distribution N(0,1) or the standard uniform distribution U(0,1)) to transform the pixel values. The transformation formula is: where, is the transformation function, is the inverse cumulative distribution function of the target distribution, is the cumulative probability of the source distribution at is equivalent to the cumulative probability value of the source distribution at .
[0048] For the standard normal distribution: For the standard uniform distribution: where, is the mapping to the standard normal distribution N(0,1), is the mapping to the standard uniform distribution U(0,1), is the inverse cumulative distribution function of the standard normal distribution, is the cumulative distribution function of the source distribution.
[0049] By distribution transformation, the LRMS and PAN data are unified into a common distribution space, thereby reducing the distribution difference between different satellite data and enhancing the generalization ability of the model.
[0050] Limited discrete distribution transformation: We provide a practical implementation for one-dimensional limited discrete distribution based on the quantile-based distribution transformation algorithm for UniPAN.
[0051] Given the pixel values sampled from the spectral band We first compute m empirical quantiles where: where, denotes the i-th quantile, denotes the smallest x value that satisfies the condition, denotes the proportion of samples in the source data that are less than or equal to x (i.e., the eCDF value), denotes the target cumulative probability threshold, where i is the quantile index (starting from 0).
[0052] These quantiles discretize the empirical cumulative distribution function (eCDF) of the source data. The number of quantiles m controls the fidelity of the approximation.
[0053] Referring to FIG. 2, Figure 2 a visualization comparison of data distribution in the process of remote sensing multispectral image panchromatic fusion using tSNE technology is shown, which specifically presents the differences in fusion results before and after applying uniform distribution.
[0054] Before applying uniform distribution, Figure 2 the visualization result on the left shows that data points from different satellites form distinct independent clusters, indicating significant distribution differences between satellite data. Such differences can lead to spectral distortion and loss of spatial details during image fusion.
[0055] After applying uniform distribution, Figure 2 the right shows that these originally separate clusters become highly mixed, with data points distributed more evenly, indicating that the distribution differences between satellite data have been significantly alleviated. This change means that image data after uniform distribution processing can be better fused, and the generated high-resolution multispectral image is better in preserving spectral and spatial information. Through this visualization comparison, the effectiveness of uniform distribution in improving image fusion quality is intuitively demonstrated, proving that it can effectively reduce the distribution differences between different satellite data and enhance the generalization ability of the fusion model.
[0056] For any data point that needs to be transformed x we use bidirectional interpolation to estimate its value in the eCDF: where, denotes the estimated eCDF value, denotes the forward interpolation, denotes the backward interpolation, denotes the source quantile set, denotes the reference cumulative probability, , is the reversed quantile, denotes the reference probability.
[0057] The function interp denotes the linear interpolation between the quantile point Q and the reference point U: Linear interpolation function is defined as: When x≤Q[0]: When Q[i]<x≤Q[i+1]: When : where, denotes the original data point (e.g., pixel value) to be converted, denotes the set of quantile points, denotes the set of cumulative probabilities corresponding to Q (reference cumulative probabilities), denotes the quantile index, and m denotes the number of quantile points (quantile numbers).
[0058] where, The mapping from the quantile point Q to the reference point U is calculated as: Finally, the conversion result T(x) is obtained by converting u(x) through the inverse cumulative distribution function Ft−1 of the target distribution ρt: The above step 102 specifically includes the following steps.
[0059] Step 1021, performing pixel-level masking on the low-resolution multispectral image to obtain a plurality of masked versions of the low-resolution multispectral image.
[0060] In the operation of the low-resolution multispectral image, the pixel-level masking process is intervened in advance, and a series of low-resolution multispectral image versions with different mask characteristics are output by means of various mask modes, thereby widening the data dimension and laying a solid foundation for subsequent processing. At the same time, the high-resolution panchromatic image (PAN) is matched into an embedding space with a size that matches the low-resolution multispectral image by linear mapping operation, which fully guarantees the collaborative correlation of the two in the feature level, and facilitates subsequent deep fusion processing. The core formula of the mapping operation is: where,Z represent the low-resolution multispectral images after mapping, whose size and channel number meet the experimental requirements. Q is the obtained low-resolution multispectral image (LRMS), which covers several channel information, and the specific channel number and length-width size depend on the actual situation. represents the corresponding high-resolution panchromatic image (PAN), which exists in the form of a single channel and has a higher spatial resolution. is a mapping operation, which is responsible for feature fusion of the low-resolution multispectral image and the high-resolution panchromatic image. represents an embedding operation, which embeds the fused features into the target space. The mask matrix formula is as follows: The elements of the mask matrix T take values of 0 or 1, and need to meet the restriction condition of the sum of the elements, and the specific formula is as follows: wherein, is a constant between 0 and . and represent the height and width of the image, respectively.
[0061] Step 1022, according to the cross-attention mechanism and the self-attention mechanism, feature extraction and reconstruction are performed based on the low-resolution multispectral images and the high-resolution panchromatic images of multiple mask versions, to obtain multiple high-resolution reconstruction results.
[0062] In specific implementation, the specific implementation method of the mask autoencoder is as follows: input the LRMS generated under different mask modes into the pixel-level integrated mask autoencoder, perform information fusion and reconstruction by using the cross-attention and self-attention mechanisms, and obtain multiple high-resolution reconstruction results.
[0063] Step 1023, based on the preset relationship modeling, scatter and integrate the multiple high-resolution reconstruction results to obtain a reconstructed high-resolution multispectral image.
[0064] In the embodiments of the present application, the obtained multiple high-resolution reconstruction results are integrated. Through different spatial relationship modeling schemes, scattering and integration are performed to maximize the preservation of spectral information and spatial details, so as to generate the final high-resolution multispectral image. This process not only enhances the detail performance of the image, but also ensures the integrity of the spectral information, so that the generated image is more valuable in actual application.
[0065] To ensure that the model can adapt to any changes in the mask matrix M when the data distribution changes, the present application proposes an innovative solution, namely the integration strategy. Specifically, by randomly sampling multiple instances of M and aggregating the reconstruction results obtained under each sampling mode, the robustness and generalization ability of the model are improved. This process can be represented as wherein, represents the reconstructed high-resolution multispectral image, represents the aggregation function responsible for fusing multiple reconstruction results. represents the process of scattering y under a randomly sampled mask matrix (Mi) to obtain the covered high-resolution multispectral image. represents the trained autoencoder model, N is the number of mask patterns, represents the parameters related to the mask pattern.
[0066] Reference Figure 3 , Figure 3 is the training paradigm schematic diagram of the network model provided by the present application. It includes high-resolution multispectral images, masked high-resolution multispectral images, and low-resolution multispectral images. The forward process includes pixel-level masking and aggregation operations, while the reverse process includes scattering and encoding operations. ).
[0067] Reference Figure 4 , Figure 4 is the model training process schematic diagram provided by the present application. It includes: low-resolution multispectral images, scattering, mask-high-resolution multispectral images, pixel-level integrated mask autoencoder, panchromatic image, integration, and estimated high-resolution multispectral image.
[0068] Through the embodiments of the present application, the pixel-level mask mechanism and multiple mask patterns are introduced, realizing the multi-angle feature extraction of low spatial resolution multispectral images (LRMS). This multi-angle feature extraction method can more comprehensively capture the detailed information in the image, avoiding the common spectral and spatial distortion problems in traditional high-resolution reconstruction methods.
[0069] Compared with traditional methods that rely on a single mask or fixed model, the present application significantly improves the accuracy of image reconstruction and the preservation ability of spectral information through the integration of multiple mask patterns. For example, traditional methods often lose details or amplify noise when dealing with complex scenes, while the present application effectively solves these problems through the synergistic effect of multiple mask patterns. In addition, the model of the present application, through the comprehensive application of multiple data sources and multiple mask patterns in the training process, enhances the robustness and generalization ability of the model.
[0070] In some embodiments, during the feature extraction and reconstruction process, the distribution-transformed features are fused, with the following specific steps: Feature extraction: Cross-attention and self-attention mechanisms are used to extract features from the distribution-transformed LRMS and PAN images, respectively.
[0071] Feature interaction: Deep convolution (DWC) is used to enhance the effective rank of the attention matrix, strengthen the interaction between different modal features, and ensure the richness of the output features.
[0072] Feature fusion: The extracted features are fused to generate multiple high-resolution reconstruction results. This process fully preserves the spectral information and spatial details, providing a foundation for subsequent result fusion.
[0073] Model training and optimization specifically includes the following steps: During model training, a distribution alignment loss function is introduced to ensure that the distribution of the high-resolution multispectral image (HRMS) output by the model is consistent with the target distribution. The specific steps are as follows: Constructing the loss function: On the basis of the original mean square error and structural similarity index, a distribution alignment loss term is added. This loss term measures the difference between the distribution of the model output HRMS and the target distribution (such as the standard normal distribution or uniform distribution).
[0074] Optimizing model parameters: By minimizing the comprehensive loss function, the model parameters are optimized. This not only improves the accuracy of image reconstruction, but also enhances the model's adaptability to different data sets.
[0075] Data augmentation and regularization: During training, data augmentation techniques (such as random flipping and rotation) and regularization methods (such as L2 regularization) are applied to prevent overfitting and further improve the model's generalization performance.
[0076] According to the remote sensing multispectral image panchromatic fusion method based on the autoencoder provided by the present application, the above method further comprises: Obtain data sample pairs from different satellites to train the training set, wherein the data sample pair training set includes: simulated low-resolution multispectral image samples, high-resolution panchromatic image samples, and high-resolution multispectral image labels; Train the preset pixel-level integrated mask autoencoder model based on the data sample pair training set through the target loss function to obtain the trained mask autoencoder; Wherein, the target loss function includes: mean square error and mean absolute error.
[0077] In the embodiment of the present application, the training process of the pixel-level integrated mask autoencoder model specifically includes the following steps: Step 1: Select low spatial resolution multispectral images (LRMS) and high spatial resolution panchromatic images (PAN) from different satellites in the same geographical location to construct a dataset. Perform distribution analysis on pixel values in each spectral band, calculate empirical quantiles, and discretize the empirical cumulative distribution function (eCDF) of the source data. Estimate the value of each pixel in the eCDF through interpolation to prepare for subsequent distribution conversion. The data pair consists of a 256x256 pixel LRMS image and a 1024x1024 pixel PAN image. Divide the dataset into training and test sets for network training and testing, respectively.
[0078] Step 2: Design a pixel-level integrated mask autoencoder architecture, including an encoder, a mask generator, a decoder, and an integration module. The encoder aims to extract feature information from the LRMS; the mask generator is responsible for generating multiple random masks; the decoder combines the masks with the features to reconstruct the HRMS; and the integration module integrates multiple reconstruction results. Integrate a distribution conversion module at the input end of the model to apply the inverse cumulative distribution function of the target distribution to the pixel values of the LRMS and PAN images, converting them to a unified target distribution. When designing the pixel-level integrated mask autoencoder model, combine the distribution conversion strategy in the paper and integrate the distribution conversion process into the model architecture. This ensures the consistency of the input data distribution and enhances the model's adaptability to different data distributions during model training and inference.
[0079] Step 3: Train the pixel-level integrated mask autoencoder model using a large-scale training set of multispectral and panchromatic image pairs. Select mean squared error (MSE) and structural similarity index (SSIM) as the loss function to optimize model parameters. To improve the model's generalization ability, introduce data augmentation and regularization techniques during training. In addition to using mean squared error (MSE) and structural similarity index (SSIM) as loss functions, introduce a distribution alignment loss during model training. This loss term measures the difference between the model's output HRMS and the target distribution, guiding the model to learn a feature representation aligned with the target distribution during training, thereby further improving the model's generalization ability and image reconstruction quality.
[0080] Step 4: Test the trained pixel-level integrated mask autoencoder model using the test set. Evaluate the model's performance in both low-resolution and full-resolution dimensions, focusing on verifying the model's reconstruction accuracy and generalization ability. Compare and analyze the model's performance using different target distributions (such as standard normal distribution and uniform distribution). By comparing the performance of different target distributions during model testing, the effectiveness of the distribution conversion and alignment strategy can be further verified, providing a basis for selecting target distributions in practical applications.
[0081] Reference Figure 5 , Figure 5 is a scattering diagram of a low-resolution multispectral image provided by the application. Among them, including: low-resolution multispectral image, scattering scheme (1-N), mask-low-resolution multispectral image and mask-high-resolution multispectral image.
[0082] In the embodiment of the application, by collecting the paired data of the low-resolution multispectral image (LRMS) and the corresponding high-resolution panchromatic image (PAN) of different satellites, a cross-satellite training set is constructed.
[0083] Simulated LRMS generation: pixel-level masking and downsampling (such as 4 times downsampling) are performed on the real low-resolution multispectral image to generate simulated low-resolution multispectral image samples matched with different satellite sensor characteristics.
[0084] Label construction: directly using the original low-resolution multispectral image as the high-resolution multispectral image label ensures the spectral and spatial authenticity.
[0085] The present achievement uses an unsupervised learning mode, skillfully integrates the essence advantages of self-supervised mechanism and mask autoencoder, and carries out integrated training for images generated under different mask modes. Such training architecture retains key spectral information and fine spatial detail features in the image in all directions. Thanks to this training form, the model constructed can achieve more stable and accurate image reconstruction results when dealing with various different scenes and various data sets. At the same time, the regularization means and data enhancement strategy integrated in the training process further improve the generalization performance of the model, so that it can realize more fine and stable reconstruction quality output when processing multispectral images in different scenes.
[0086] Through the embodiment of the application, low-resolution (simulated downsampling) and full-resolution (original PAN) evaluation indicators are optimized in the training, ensuring that the model is effective in both scenarios. MSE dominates the overall structure (such as large-area ground object spectral consistency), and MAE suppresses salt and pepper noise (such as edge sharpening artifacts).
[0087] According to the remote sensing multispectral image panchromatic fusion method based on the autoencoder provided by the application, the data sample pairs from different satellites are obtained to train the set, including: Obtain low-resolution multispectral image samples and high-resolution panchromatic image samples with the same geographic coordinates; Perform element-by-element masking operation on the low-resolution multispectral image samples to obtain simulated low-resolution multispectral image samples; The simulated low-resolution multispectral image samples and the high-resolution panchromatic image samples are used as data sample pairs; The low-resolution multispectral image sample is taken as a high-resolution multispectral image label.
[0088] In the embodiment of the present application, the low-resolution multispectral image sample and the high-resolution panchromatic image sample image pair with the same geographical position and corresponding to each other are collected. The low-resolution multispectral image sample is taken as an operation object, and an element-by-element mask process is performed, so as to obtain a simulated low-resolution multispectral image sample. The simulated low-resolution multispectral image sample and the corresponding high-resolution panchromatic image sample are combined to form a data sample pair. The low-resolution multispectral image sample is explicitly taken as a label of the high-resolution multispectral image.
[0089] In the implementation case of the present solution, in the initial stage, the image pair composed of the low-resolution multispectral image sample and the high-resolution panchromatic image sample with the same geographical coordinate attribute is accurately acquired. It is determined that the low-resolution multispectral image sample is equivalent to the label of the high-resolution multispectral image. Then, the mask matrix with the same size and the high-resolution multispectral image label is used to perform the element-by-element mask operation on the high-resolution multispectral image label, and the simulated low-resolution multispectral image (LRMS) is successfully generated. Finally, the processed LRMS and the PAN image are taken as the input source, and the original low-resolution multispectral image sample is taken as the label of the high-resolution multispectral image. According to the predetermined ratio, the overall data is divided into a training set and a test set, so as to lay a data foundation for the subsequent image sharpening process.
[0090] According to the embodiment of the present application, due to the lack of real HRMS data, the present application regards the LRMS as the result after the pixel-level mask on the HRMS.
[0091] According to the present application, a remote sensing multispectral image panchromatic fusion method based on an autoencoder is provided, and a target loss function includes: Wherein, The target loss function is represented by MAE, and the average absolute error is represented by MAE. The mean square error is represented by MSE, and the mean square error is represented by MSE. Wherein, The average absolute error is represented by MAE, and the average absolute error is represented by MAE. The mean square error is represented by MSE, and the mean square error is represented by MSE. The actual value of the pixel in the i-th row and the j-th column of the i-th band of the high-resolution multispectral image label is represented by Y i,j. The actual value of the pixel in the i-th row and the j-th column of the i-th band of the high-resolution multispectral image label is represented by Y i,j. The actual value of the pixel in the i-th row and the j-th column of the i-th band of the high-resolution multispectral image label is represented by Y i,j. The actual value of the pixel in the i-th row and the j-th column of the i-th band of the high-resolution multispectral image label is represented by Y i,j. The actual value of the pixel in the i-th row and the j-th column of the i-th band of the high-resolution multispectral image label is represented by Y i,j. the first row of the column of pixels; denotes the number of image bands; denotes the image height; denotes the image width; is an adjustable weight parameter for balancing and the contribution ratio in the total loss.
[0092] In the implementation of the present application, the AdamW optimizer is used in the training link, and the specific parameter settings are: β1=0.9, β2=0.999, and the initial learning rate is 0.001. The training process continues until the model converges, thereby obtaining a model with high-quality image reconstruction capability.
[0093] Reference Figure 6 shows examples of reconstruction results achieved by the present scheme in real scenes, covering low-resolution multispectral images, high-resolution multispectral images, and panchromatic images. In the model training stage, the comprehensive performance of the model can be optimized by adjusting the number of mask modes (N). Specifically, in the training process, a plurality of different mask modes are used to generate a corresponding number of high-resolution multispectral images (HRMS) processed by masks, and then the image restoration results obtained under these different mask modes are integrated to effectively enhance the robustness and reconstruction accuracy of the model. In addition, the necessity of each component module is rigorously verified by means of ablation experiment, fully confirming the rationality of the model design and the effectiveness of the actual application. Figure 6 The test results show that the remote sensing multispectral image panchromatic fusion method based on the autoencoder proposed by the present scheme performs well in significantly improving the image fusion accuracy and generalization ability. Whether in a low-resolution scene or a full-resolution scene, high-quality image reconstruction results can be achieved, thereby greatly improving the overall quality of multispectral remote sensing image data and its value in practical applications.
[0094] According to the remote sensing multispectral image panchromatic fusion method based on the autoencoder designed by the present scheme, according to the cross-attention mechanism and the self-attention mechanism, feature extraction and reconstruction are performed based on a plurality of mask versions of low-resolution multispectral images and high-resolution panchromatic images, to obtain a plurality of high-resolution reconstruction results, including: The embedding of the plurality of mask versions of the low-resolution multispectral images is taken as the query, and the embedding of the high-resolution panchromatic image is taken as the key and the value: wherein, denotes the cross-attention weight result, denotes the self-attention weight result, denotes the query, denotes the key, denotes the value, denotes the dimension of the key, denotes the normalization function, denotes the transpose matrix of the key, denotes the focusing function, which is used to approximate the normalization function and maintain linear time complexity.
[0095] The core process of using the cross-attention mechanism and the self-attention mechanism to perform feature extraction and reconstruction based on multiple mask versions of low-resolution multispectral images and high-resolution panchromatic images to obtain multiple high-resolution reconstruction results is as follows: The embedding vectors of multiple mask-processed different versions of low-resolution multispectral images are taken as the query (Q), and the embedding vectors of the high-resolution panchromatic image are taken as the key (K) and the value (V) respectively, in this way, the interaction and correlation between the features of the two different modal images are fully promoted, and the mathematical definition form is: wherein, denotes the cross-attention weight result, denotes the self-attention weight result, denotes the query, denotes the key, denotes the value, denotes the dimension of the key, denotes the normalization function, denotes the transpose matrix of the key, denotes the focusing function, which is used to approximate the normalization function and maintain linear time complexity.
[0096] Reference Figure 7 , Figure 7 is the flowchart of the cross-attention mechanism and the self-attention mechanism provided by the application, which comprises: query (Q), key (K) and value (V), cross-attention, addition and normalization, multi-layer perception, addition and normalization, query (Q), key (K) and value (V), self-attention, addition and normalization, multi-layer perception, and addition and normalization.
[0097] Linear cross-attention calculation is performed on the input feature map and the mask. First, the embedding of the LRMS is used as the query (Q), and the embedding of the PAN is used as the key (K) and the value (V) to enhance the interaction between the two modalities.
[0098] wherein, denotes a cross-attention weight result, denotes a query, denotes a key, denotes a value, denotes a dimension of the key, denotes a normalization function, denotes a transpose matrix of the key.
[0099] Reference Figure 8 , Figure 8 is a flowchart of the focusing linear attention method provided by the application.
[0100] Figure 8 The flow of the focusing linear attention method used in the scheme is clearly illustrated, and the entire flow covers key steps such as feature projection, query embedding, background embedding, deep convolution (DWC), matrix multiplication operation, element-wise addition operation, feature normalization processing, activation function application, and element-wise p power calculation in turn.
[0101] In order to realize the significant improvement of the calculation efficiency, the traditional self-attention mechanism with quadratic time complexity is optimized and improved to convert it into a linear time complexity calculation form. Specifically, the focusing linear attention method is used (for details, see FIG. 7), so that the original self-attention calculation process is converted into a calculation method that meets the linear time complexity requirement, and the core steps are summarized as follows: wherein, is a focusing function, which is used to approximate the Softmax function while maintaining linear time complexity.
[0102] According to the remote sensing multispectral image panchromatic fusion method based on the autoencoder provided by the application, the focusing function comprises: wherein, denotes a focusing function, denotes an input feature matrix, denotes a direction adjustment mapping, denotes a nonlinear activation function; wherein, denotes an input vector, denotes a norm of the input vector, denotes a focusing factor.
[0103] Here, the focusing factor is used to enlarge the difference between similar query key pairs.
[0104] In the linear attention mechanism used in the present solution, the LeakyReLU activation function is particularly selected to enhance the ability of the model to capture complex data patterns. LeakyReLU (Leaky ReLU) activation function is a widely used activation function in the field of neural networks. Its core advantage is to effectively avoid the gradient vanishing problem of traditional ReLU (Rectified Linear Unit) activation function when the input value is less than zero by introducing a small slope value. Its specific mathematical definition formula is: Where α is a small constant, usually taking the value of 0.01. Through this formula, A certain gradient can be maintained on the negative input value, which can better update the parameters during training.
[0105] In addition, in order to increase the effective rank of the attention matrix, deep convolution (DWC) is applied to the value (V). The additional DWC layer increases the effective rank of the attention matrix, thereby restoring the diversity of the output features. The final attention calculation is: Where the formula of DWC (Depthwise Convolution) combined with projection can be expressed as: : the value of the output feature map at position ( ) and channel ; : the value of the input feature map at position ( ) and channel ; : the weight of the convolution kernel at position ( ) and channel ; : represents the value of the projection matrix at position ( ) and channel ; : the height and width of the convolution kernel.
[0106] Through the above steps, the linear cross-attention mechanism can improve the computational efficiency while maintaining the diversity of the attention matrix and the richness of the output features.
[0107] The linear cross-attention mechanism introduced in this solution significantly optimizes the computational complexity of traditional self-attention mechanisms, reducing the original quadratic time complexity to linear time complexity. This optimization not only greatly reduces the occupation of computing resources, but also performs well in maintaining efficient feature interaction. By using the embedding of LRMS (low-resolution multi-spectral image) as the query (Q), the embedding of PAN (high-resolution panchromatic image) as the key (K) and value (V), this mechanism effectively enhances the interaction between the two modalities and significantly improves the computational efficiency. Traditional self-attention mechanisms often result in long training and inference times due to high computational complexity when dealing with large-scale image data. By introducing the linear cross-attention mechanism, this solution not only improves computational efficiency but also exhibits stronger adaptability. Especially in resource-constrained scenarios, the advantages are particularly obvious, significantly expanding the application scenarios.
[0108] The following are the example steps of this solution in practical application: Step 1: Data set preparation Collect low-resolution multi-spectral images (LRMS) and high-resolution panchromatic images (PAN) from different satellites, each data pair consisting of a 256x256 pixel LRMS and a 1024x1024 pixel PAN. Divide the data set into training and validation sets to prepare for subsequent model training and evaluation.
[0109] Step 2: Pixel-level integrated mask autoencoder design Construct a pixel-level integrated mask autoencoder architecture, covering the encoder, mask generator, decoder, and integration module. The encoder is responsible for extracting features from LRMS; the mask generator is used to create multiple random masks; the decoder reconstructs HRMS (high-resolution multi-spectral image) combined with the mask and features; and the integration module integrates multiple reconstruction results to generate the final high-resolution multi-spectral image.
[0110] Step 2-1: Mask processing Perform pixel-level mask operations on LRMS to generate multiple mask versions of LRMS, providing diverse inputs for subsequent feature extraction and reconstruction.
[0111] Step 2-2: Feature extraction and reconstruction Input the mask version of LRMS into the autoencoder to extract features and reconstruct HRMS using cross-attention and self-attention mechanisms, obtaining multiple high-resolution reconstruction results to provide a basis for result fusion.
[0112] Step 2-3: Result fusion By modeling different spatial relationships, scattering and integrating multiple reconstruction results, the spectral information and spatial details are maximized to generate high-quality high-resolution multispectral images.
[0113] Step 2-4: Data distribution analysis Perform data distribution analysis on the input LRMS and PAN, calculate the empirical cumulative distribution function (eCDF), and understand the statistical characteristics of the image data.
[0114] Step 2-5: Distribution conversion Adopt a similar distribution conversion strategy as in the paper to convert the pixel values of LRMS and PAN images to a unified target distribution (such as Gaussian distribution or uniform distribution), reducing the distribution differences between different data sets.
[0115] Step 3: Model training Use a large-scale multispectral and panchromatic image training set to train the pixel-level integrated mask autoencoder model. During training, use mean square error and structural similarity index as loss functions to optimize model parameters. At the same time, apply data augmentation and regularization techniques to improve model generalization. Through the integration of different mask modes, multiple recovery results are obtained and fused to further strengthen the model's ability to preserve spectral information and spatial details.
[0116] Step 3-1: Distribution alignment in model training Introduce distribution alignment loss in model training to make the distribution of HRMS output by the model consistent with the target distribution, further improving the model's generalization ability.
[0117] Step 4: Use the validation set to evaluate the performance of the trained model, focusing on the accuracy of image reconstruction at different resolutions and the generalization performance of the model.
[0118] The pixel-level integrated mask autoencoder proposed in the scheme includes a mask generator, a decoder, and an integration module. By performing random mask and scattering operations, multiple versions of mask-processed high-resolution multispectral images (HRMS) can be generated, and then the results of these versions are fused using an integration strategy to generate HRMS that not only retains the spectral features of low-resolution multispectral images (LRMS) but also fuses the spatial details of high-resolution panchromatic images (PAN), effectively improving the accuracy of image reconstruction and the level of spectral information retention. The linear cross-attention mechanism used in this scheme combines the advantages of cross-attention and self-attention mechanisms, optimizing the quadratic time complexity of traditional self-attention mechanisms to linear time complexity. Specifically, the focus linear attention method is used, which not only improves computational efficiency but also strengthens the interaction between different modal features, further optimizing the reconstruction effect. In the value calculation process, deep convolution (DWC) is applied to improve the effective rank of the attention matrix and ensure that the richness of the output features is effectively preserved. In addition, based on the integration training strategy of multiple mask modes, loss functions such as mean square error and structural similarity index are used to optimize model parameters, significantly enhancing the robustness and reconstruction accuracy of the model.
[0119] The scheme has significant application prospects in many fields, including but not limited to agriculture, forestry, environmental protection, and urban planning. By achieving high-quality reconstruction of multispectral images, the scheme can provide more accurate and complete data support for these fields, thereby improving the practical application value of remote sensing data.
[0120] Overall, the scheme uses innovative technical means such as pixel-level integrated mask autoencoder models, linear cross-attention mechanisms, and unsupervised training methods to successfully solve the problems of spectral and spatial distortion, high computational complexity, and limited generalization ability in existing technologies, significantly improving the sharpening effect and processing efficiency of multispectral images. These technical innovations have brought outstanding performance and wide applicability to the scheme in practical applications.
[0121] The following is a description of the self-encoder-based remote sensing multispectral image panchromatic fusion device provided by the scheme, which corresponds to the aforementioned self-encoder-based sharpening method and can be understood with reference to it.
[0122] Reference Figure 9 The figure shows the structural composition of the self-encoder-based remote sensing multispectral image panchromatic fusion device involved in the scheme.
[0123] The acquisition module 901 is used to acquire low-resolution multispectral images and high-resolution panchromatic images of a target geographic location. The pixel-level integrated mask automatic coding module 902 is configured to input the low-resolution multispectral image and the high-resolution panchromatic image into the trained pixel-level integrated mask automatic encoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location, and the method comprises the following steps: performing pixel-level mask on the low-resolution multispectral image to obtain a plurality of masked versions of the low-resolution multispectral image; performing feature extraction and reconstruction based on the plurality of masked versions of the low-resolution multispectral image and the high-resolution panchromatic image according to the cross-attention mechanism and the self-attention mechanism to obtain a plurality of high-resolution reconstruction results; performing scattering and integration on the plurality of high-resolution reconstruction results based on the preset relationship modeling to obtain the reconstructed high-resolution multispectral image.
[0124] The remote sensing multispectral image panchromatic fusion device based on the autoencoder provided in the scheme can completely implement the operation steps in the method embodiments and achieve the corresponding technical effects. To avoid repeated description, the same part of the method embodiments and the beneficial effects thereof will not be described here.
[0125] Figure 10 is the entity structure schematic diagram of the electronic device provided by the present application, as Figure 10 shown, the electronic device can include: a processor (processor) 1010, a communications interface (communications interface) 1020, a memory (memory) 1030 and a communications bus 1040, wherein the processor 1010, the communications interface 1020, the memory 1030 complete the communication between each other through the communications bus 1040. The processor 1010 can call the logic instruction in the memory 1030 to execute the remote sensing multispectral image panchromatic fusion method based on the autoencoder, and the method comprises: acquiring a low-resolution multispectral image and a high-resolution panchromatic image of a target geographic location; inputting the low-resolution multispectral image and the high-resolution panchromatic image into the trained pixel-level integrated mask automatic encoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location, wherein the method comprises: performing pixel-level mask on the low-resolution multispectral image to obtain a plurality of masked versions of the low-resolution multispectral image; performing feature extraction and reconstruction based on the plurality of masked versions of the low-resolution multispectral image and the high-resolution panchromatic image according to the cross-attention mechanism and the self-attention mechanism to obtain a plurality of high-resolution reconstruction results; performing scattering and integration on the plurality of high-resolution reconstruction results based on the preset relationship modeling to obtain the reconstructed high-resolution multispectral image.
[0126] Furthermore, the logic instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0127] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the remote sensing multispectral image panchromatic fusion method based on an autoencoder provided by the above methods, the method including: obtaining a low-resolution multispectral image and a high-resolution panchromatic image of the target geographic location; inputting the low-resolution multispectral image and the high-resolution panchromatic image into a trained pixel-level integrated mask autoencoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location, which includes: pixel-level masking of the low-resolution multispectral image to obtain multiple masked versions of the low-resolution multispectral image; according to the cross-attention mechanism and the self-attention mechanism, feature extraction and reconstruction are performed based on the multiple masked versions of the low-resolution multispectral image and the high-resolution panchromatic image to obtain multiple high-resolution reconstruction results; scattering and integrating the multiple high-resolution reconstruction results based on preset relationship modeling to obtain a reconstructed high-resolution multispectral image.
[0128] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the self-encoder-based panchromatic fusion method for remote sensing multispectral images provided by the above method, the method comprising: obtaining a low-resolution multispectral image and a high-resolution panchromatic image of a target geographic location; inputting the low-resolution multispectral image and the high-resolution panchromatic image into a trained pixel-level integrated mask auto-encoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location, wherein the method comprises: performing pixel-level masking on the low-resolution multispectral image to obtain a plurality of masked versions of the low-resolution multispectral image; performing feature extraction and reconstruction based on the plurality of masked versions of the low-resolution multispectral image and the high-resolution panchromatic image according to a cross-attention mechanism and a self-attention mechanism to obtain a plurality of high-resolution reconstruction results; and performing scattering and integration on the plurality of high-resolution reconstruction results based on a pre-set relationship modeling to obtain the reconstructed high-resolution multispectral image.
[0129] The device embodiments described above are merely illustrative, wherein the units shown as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0130] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software and the necessary general hardware platform, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in each embodiment or some part of the embodiment.
[0131] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A remote sensing multispectral image panchromatic fusion method based on autoencoder, characterized in that: include: Acquire low-resolution multispectral images and high-resolution panchromatic images of the target geographic location; Inputting the low-resolution multispectral image and the high-resolution panchromatic image into a trained pixel-level integrated mask autoencoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location, wherein: Performing pixel-level masking on the low-resolution multispectral image to obtain multiple masked versions of the low-resolution multispectral image; performing feature extraction and reconstruction based on the multiple masked versions of the low-resolution multispectral images and the high-resolution panchromatic image according to a cross-attention mechanism and a self-attention mechanism to obtain multiple high-resolution reconstruction results; The multiple high-resolution reconstruction results are scattered and integrated based on a preset relationship model to obtain a reconstructed high-resolution multispectral image.
2. The remote sensing multispectral image panchromatic fusion method based on autoencoder according to claim 1, characterized in that: The method further comprises: Acquire a training set of data sample pairs from different satellites, wherein the training set of data sample pairs includes: simulated low-resolution multispectral image samples, high-resolution panchromatic image samples, and high-resolution multispectral image labels; Training a preset pixel-level integrated mask autoencoder model based on the training set of the data samples using a target loss function to obtain a trained mask autoencoder; The target loss function includes mean square error and mean absolute error.
3. The remote sensing multispectral image panchromatic fusion method based on autoencoder according to claim 2, characterized in that: The step of obtaining data samples from different satellites for a training set includes: Obtain low-resolution multispectral image samples and high-resolution panchromatic image samples with the same geographic coordinates; performing an element-by-element masking operation on the low-resolution multispectral image samples to obtain simulated low-resolution multispectral image samples; taking the simulated low-resolution multispectral image sample and the high-resolution panchromatic image sample as a data sample pair; The low-resolution multispectral image samples are used as high-resolution multispectral image labels.
4. The remote sensing multispectral image panchromatic fusion method based on autoencoder according to claim 2, characterized in that: The objective loss function includes: in, represents the target loss function, MAE represents the mean absolute error, and MSE represents the mean square error, where: in, represents the mean absolute error, represents the mean square error, The first The first band Rank The actual value of the column pixel; Represents the first output of the preset pixel-level integrated mask autoencoder model The first band Rank The predicted value of the column pixel; Indicates the number of image bands; Indicates the image height; Indicates the image width; is an adjustable weight parameter used to balance and The proportion of contribution to the total loss.
5. The remote sensing multispectral image panchromatic fusion method based on autoencoder according to claim 1, characterized in that: The method further comprises performing feature extraction and reconstruction based on the multiple masked versions of the low-resolution multispectral images and the high-resolution panchromatic image according to the cross-attention mechanism and the self-attention mechanism to obtain multiple high-resolution reconstruction results, including: Using the embeddings of the multiple masked versions of the low-resolution multispectral image as queries and the embeddings of the high-resolution panchromatic image as keys and values: in, represents the cross attention weight result, represents the self-attention weight result, Indicates a query, Indicates the key, Represents a value, represents the dimension of the key, represents the normalization function, represents the transposed matrix of the key, Represents the focusing function, which is used to approximate the normalization function and maintain linear time complexity.
6. The remote sensing multispectral image panchromatic fusion method based on autoencoder according to claim 5, characterized in that: The focusing function includes: in, represents the focusing function, represents the input feature matrix, Indicates the direction adjustment mapping, represents a nonlinear activation function; in, represents the input vector, represents the norm of the input vector, Represents the focusing factor.
7. A remote sensing multispectral image full color fusion device based on autoencoder, characterized in that: include: An acquisition module is used to acquire low-resolution multispectral images and high-resolution panchromatic images of the target geographic location; A pixel-level integrated mask auto-encoding module is configured to input the low-resolution multispectral image and the high-resolution panchromatic image into a trained pixel-level integrated mask auto-encoder model to obtain a reconstructed high-resolution multispectral image of the target geographic location, including: Performing pixel-level masking on the low-resolution multispectral image to obtain multiple masked versions of the low-resolution multispectral image; performing feature extraction and reconstruction based on the multiple masked versions of the low-resolution multispectral images and the high-resolution panchromatic image according to a cross-attention mechanism and a self-attention mechanism to obtain multiple high-resolution reconstruction results; The multiple high-resolution reconstruction results are scattered and integrated based on a preset relationship model to obtain a reconstructed high-resolution multispectral image.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the remote sensing multispectral image sharpening method based on the autoencoder is implemented as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the remote sensing multispectral image sharpening method based on an autoencoder as described in any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the remote sensing multispectral image sharpening method based on an autoencoder as described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Remote sensing large model pre-training method and device, electronic equipment and storage medium
CN121482632A
Multi-modal correction panchromatic sharpening system and method based on task allocation method
CN121937330A