Hyperspectral and multispectral image fusion method based on wavelet feature fusion and comparative learning

By constructing a hyperspectral and multispectral image fusion network based on wavelet feature fusion and contrast learning, the problems of insufficient detail texture restoration and modal consistency in hyperspectral image fusion are solved, efficient image fusion effect is achieved, and high-quality high-resolution hyperspectral images are generated.

CN120672589APending Publication Date: 2025-09-19DONGHUA UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510760596.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing hyperspectral and multispectral image fusion technologies have shortcomings in detail texture restoration and modal consistency maintenance, especially the poor effects of high-frequency information extraction and cross-modal feature alignment, resulting in unsatisfactory image fusion effects.

Method used

The wavelet feature fusion and contrastive learning methods are used to construct a fusion network model, including a wavelet transform module, a cross-modal wavelet feature fusion module and a high-frequency contrastive learning module. End-to-end supervised training is performed through a composite loss function to improve the spatial resolution and spectral consistency of the image.

Benefits of technology

It effectively enhances the image's high-frequency texture and cross-modal feature alignment capabilities, generates fused images with high spatial resolution and high spectral fidelity, solves the problems of detail blur and cross-modal mismatch in existing methods, and improves detail reconstruction and spectral information preservation in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672589A_ABST
    Figure CN120672589A_ABST
Patent Text Reader

Abstract

The invention discloses a high-resolution hyperspectral image reconstruction method based on wavelet domain feature fusion and contrast learning, and belongs to the technical field of image fusion and super-resolution reconstruction. The method comprises the following steps: constructing a fusion network model comprising a wavelet transformation module, a cross-modal feature fusion module, a high-frequency contrast learning module and an image reconstruction module; performing end-to-end supervised training by using a training data set containing the low-resolution hyperspectral image and the high-resolution multispectral image; and after training is completed, inputting a test image pair to realize image reconstruction. According to the method, the detail retention capability is improved by combining wavelet decomposition and a directional fusion mechanism, the cross-modal high-frequency feature alignment capability is enhanced through comparative learning, a fusion image with high spatial resolution and high spectral consistency is finally generated, and the method is suitable for multi-modal image reconstruction tasks such as remote sensing, medical and natural images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and artificial intelligence technology, and in particular to a hyperspectral and multispectral image fusion method based on wavelet feature fusion and contrast learning. Background Art

[0002] In recent years, hyperspectral and multispectral image fusion technology has developed rapidly in the fields of computer vision and remote sensing, becoming a hot topic among researchers. The core of this technology lies in the effective integration of information obtained by different sensors, giving full play to the high spectral resolution advantages of hyperspectral images in material identification and the high resolution advantages of multispectral images in spatial detail representation, so as to generate images with both rich spectral information and fine spatial details, thereby more accurately restoring the scene content. This fusion not only improves the overall quality of the image, but also provides more comprehensive data support for application fields such as precision agriculture, land cover monitoring, and medical image analysis. It is widely used in tasks such as environmental monitoring, urban planning, and disease detection. With the advancement of imaging technology, more and more systems are able to simultaneously acquire low-resolution hyperspectral images (LR-HSI) and high-resolution multispectral images (HR-MSI), providing a practical foundation for the development of image fusion technology. Overall, hyperspectral and multispectral image fusion technology has shown great potential in enriching scene understanding and improving the performance of downstream visual tasks by making up for the shortcomings of a single imaging mode, and has become an important research direction in the field of image super-resolution.

[0003] In practical applications, due to the limitations of imaging hardware, hyperspectral images often suffer from insufficient spatial resolution, while multispectral images suffer from limited spectral information. To address this issue, researchers have proposed a variety of fusion methods, primarily those based on transform domains, sparse representations, and deep learning. Transform domain-based methods, such as wavelet transforms and principal component analysis, achieve feature decoupling and reorganization by converting image features to the frequency domain, enhancing image detail to a certain extent. However, due to their reliance on linear transformations, these methods have significant limitations in capturing complex nonlinear image features and can easily lead to loss and distortion of spectral information. Sparse representation methods achieve image reconstruction through dictionary learning and feature selection, achieving a balance between spatial resolution and spectral fidelity to a certain extent. However, these methods rely heavily on prior knowledge of the image and have limited generalization performance. In recent years, with the development of deep learning technology, methods such as convolutional neural networks (CNNs) have been widely used in hyperspectral and multispectral image fusion tasks. The CNN model can automatically extract multi-level features and significantly improve the fusion effect. However, it is limited by the receptive field of the convolution kernel, which can easily lead to insufficient performance of the fused image in detail texture restoration. At the same time, there are significant modal differences between hyperspectral and multispectral, which makes feature alignment and fusion challenging and easily leads to spectral distortion and spatial structure degradation. Although some studies have attempted to introduce frequency domain analysis and contrast learning mechanisms to enhance feature expression capabilities and cross-modal consistency, existing methods still have deficiencies in high-frequency information extraction, modal consistency maintenance, and fine-grained feature reconstruction. How to achieve effective extraction and fusion of multimodal high-frequency details is still an important issue that needs to be solved in the field of hyperspectral image super-resolution reconstruction.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0005] The purpose of the present invention is to solve the technical problems existing in the background technology. To this end, a hyperspectral and multispectral image fusion method based on wavelet feature fusion and contrast learning is provided, which aims to effectively improve the spatial resolution and spectral consistency of the reconstructed hyperspectral image, improve detail preservation and heterogeneous modality alignment problems.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A hyperspectral and multispectral image fusion method based on wavelet feature fusion and contrast learning includes the following steps:

[0008] Step S1: Constructing a fusion network model, wherein the fusion network model includes a wavelet transform module, a cross-modal wavelet feature fusion module (Cross-Modal Wavelet Fusion Module, CMWF), a high-frequency contrastive learning module (High-Frequency Contrastive Learning Module, HFCL) and an image reconstruction module.

[0009] Specifically: before network input, the low-resolution hyperspectral image is preprocessed by band grouping. According to the spectral correlation between the bands, all the bands are divided into several band groups, and each band group serves as the input of an independent subnetwork; the wavelet transform module is used to perform stationary wavelet decomposition on each band group and the corresponding high-resolution multispectral image to generate low-frequency and high-frequency features; the subnetwork feature extraction module is used to extract the subspace features of each band group; the CMWF module is used to fuse the high-frequency details of hyperspectral and multispectral images in the frequency domain; the HFCL module is used for feature alignment optimization; the inverse wavelet transform unit is used to perform inverse wavelet transform on the fused features output by each subnetwork respectively, and finally the reconstruction results of each subnetwork are spliced ​​in the channel dimension to generate a complete high-resolution hyperspectral image (HR-HSI).

[0010] Step S2: Obtain a training dataset, which includes low-resolution hyperspectral images, high-resolution multispectral images, and their corresponding reference high-resolution hyperspectral images. Construct training pairs by combining images of different modalities, and perform end-to-end supervised training using a composite loss function.

[0011] Specifically: low-resolution hyperspectral images and corresponding high-resolution multispectral images are obtained to form a training data set. The fusion network model is trained using the training data set. During the training process, a composite loss function including reconstruction loss, spectral angle loss, contrastive learning loss and wavelet grouping loss is used for optimization to improve the spatial resolution, spectral consistency and texture detail recovery ability of the image.

[0012] Step S3: Input the low-resolution hyperspectral image to be fused and the corresponding high-resolution multispectral image into the trained fusion network model, and output the high-resolution hyperspectral image reconstruction result.

[0013] Specifically, the low-resolution hyperspectral image to be fused and the corresponding high-resolution multispectral image are input into the fusion network model trained in step S2, and the final high-resolution hyperspectral image output is generated through feature extraction, fusion, alignment and reconstruction.

[0014] The following is a technical solution further defined by the present invention. In step S1, the wavelet transform module uses stationary wavelet transform to perform multi-scale decomposition on the input image, extracting low-frequency sub-bands and high-frequency sub-bands in three directions respectively, and all sub-bands are kept consistent with the original Figure 1 Consistent space dimensions.

[0015] Specifically, the wavelet transform module uses the Stationary Wavelet Transform (SWT) to decompose the input low-resolution hyperspectral image and high-resolution multispectral image, generating low-frequency subbands that maintain a constant spatial size and high-frequency subbands in three directions. The wavelet transform module includes an SWT decomposition layer, which uses a fixed wavelet filter to extract multi-scale frequency-domain features from the image. The resulting subband features are used in subsequent subnetwork modules for feature extraction and fusion processing.

[0016] The following is a technical solution further defined by the present invention. In step S1, the CMWF module is composed of multiple parallel fusion units. Each fusion unit receives the sub-band features of the hyperspectral image and the multispectral image in a high-frequency direction, and adopts the attention mechanism to fuse the high-frequency features of the two modalities.

[0017] Specifically: The CMWF module includes a cross-attention calculation unit and a fusion convolution unit. The cross-attention calculation unit uses a convolution structure to generate query (Q), key (K) and value (V) features, and applies the scaled dot product attention mechanism to calculate the attention weight; the fusion convolution unit is used to fuse the weighted high-frequency features to enhance the expression of local spatial texture and high-frequency details.

[0018] The following is a technical solution further defined by the present invention. In step S1, the HFCL module includes a sample sampler, an embedding encoder and a loss calculation unit. The sample sampler is used to extract anchor samples, positive samples and negative samples from the fused high-frequency features. The embedding encoder is a multi-layer perceptron structure. The samples are projected into the latent space through the MLP layer. The MLP module includes two sequentially connected fully connected layers with a GELU activation function inserted in the middle. By optimizing the contrast loss, the model learns to distinguish textures in different directions and enhances the discriminative power and alignment effect of the features. The contrast loss is constructed based on the temperature-scaled dot product contrast loss function.

[0019] Specifically: The HFCL module includes a high-frequency feature sampling unit (i.e., sample sampler) and a contrastive loss calculation unit (i.e., loss calculation unit). The high-frequency feature sampling unit extracts anchor features, positive sample features, and negative sample features from the high-frequency sub-band of the fused features; the contrastive loss calculation unit uses the standardized vector dot product similarity and combines it with the temperature scaling parameter to perform similarity calculation, ultimately forming a contrastive loss term to optimize the feature space to aggregate homologous samples and separate heterogeneous samples.

[0020] The following is a technical solution further defined by the present invention. In step S1, the image reconstruction module includes an inverse stationary wavelet transform module, which inputs the fused low-frequency sub-band and the directional high-frequency sub-band. A high-resolution image is generated through inverse wavelet reconstruction, and the reconstructed multiple sub-network output channels are spliced ​​to obtain the final image.

[0021] Specifically, the inverse stationary wavelet transform module uses the inverse stationary wavelet transform (ISWT) filter to perform an inverse transform operation on the fused high-frequency sub-band and low-frequency sub-band features, directly restoring them into a high-resolution hyperspectral image in the spatial domain. Finally, the sub-images output by each subnetwork are spliced ​​in the channel dimension to form a complete high-resolution hyperspectral image.

[0022] In step S2, the training dataset consists of LR-HSI, HR-MSI, and the corresponding HR-HSI. The low-resolution hyperspectral image is obtained by Gaussian blurring the reference hyperspectral image and then downsampling it to simulate the physical process of spatial resolution degradation during imaging. During the training process, the generated LR-HSI and the corresponding HR-MSI are used as input to the fusion network model, and are processed in sequence by the wavelet transform module, CMWF module, HFCL module, and image reconstruction module to generate a reconstructed image. The training adopts end-to-end supervised optimization based on a composite loss function. The composite loss function simultaneously constrains the pixel error, spectral consistency, feature distribution alignment, and high-frequency texture preservation of the reconstructed image to improve the overall quality and detail restoration ability of the fused image.

[0023] The following is a technical solution further defined by the present invention. In step S2, the composite loss function includes reconstruction pixel error loss, spectral angle loss, high-frequency contrast learning loss, and wavelet domain high-frequency reconstruction loss, specifically:

[0024]

[0025] Among them, L rec is the pixel-level L1 reconstruction loss, which is used to constrain the pixel difference between the generated image and the reference image; L cl It is a high-frequency contrast learning loss that promotes the consistency of cross-modal feature space by optimizing the similarity of positive and negative samples; sam is the spectral angle error loss, which is used to measure the angle consistency between the generated image and the reference image in the spectral direction and reduce spectral distortion; L swt is the high-frequency reconstruction loss in the wavelet domain, which is based on the high-frequency reconstruction loss of wavelet grouping. It improves the quality of detail and boundary restoration by performing fine-grained supervision on each high-frequency subband in the wavelet domain. I represents the number of samples in a batch; α, β, and γ are weighting coefficients, which are adjusted according to different task requirements.

[0026] The following is a technical solution further defined by the present invention: pixel-level L1 reconstruction loss L rec It is used to constrain the consistency between the network output image and the reference image in the pixel space. It is defined by the L1 norm and the calculation formula is:

[0027]

[0028] in, is the reconstructed hyperspectral image of the i-th sample in a batch; Z i is the corresponding reference high-resolution hyperspectral image; g(·) represents the L1 loss calculation.

[0029] The following is a technical solution further defined by the present invention: high-frequency contrast learning loss L cl It is used to improve the consistency of cross-modal high-frequency features in the feature space and optimizes the similarity difference between anchor samples and positive and negative samples. Specifically:

[0030]

[0031] Among them, A i represents the i-th anchor point sample in a batch; P i is a positive sample; is a negative sample; sim(·) represents the dot product similarity after vector normalization; τ is the temperature scaling parameter; I represents the number of samples in a batch.

[0032] The following is a technical solution further defined by the present invention: spectral angle error loss L sam It is used to measure the angle consistency between the reconstructed image and the reference image in the spectral direction and reduce spectral distortion. The calculation formula is:

[0033]

[0034] in, Represents the predicted spectrum vector of the i-th sample in a batch on the band group k; represents the corresponding reference spectrum vector; h(·) represents the cosine similarity calculation.

[0035] The following is a technical solution further defined by the present invention: the wavelet domain high frequency reconstruction loss L swt It is used to supervise the high-frequency subbands of the reconstructed image in a fine-grained manner in the frequency domain to maintain texture details and boundary information. The definition formula is:

[0036]

[0037] Among them, m is the direction of the wavelet high-frequency subband, and its values ​​include horizontal direction (LH), vertical direction (HL) and diagonal direction (HH); Represents the feature representation of the reconstructed image of the i-th sample in a batch on the m-th wavelet subband; represents the feature representation of the reference image on the same wavelet subband; g(·) represents the L1 loss calculation.

[0038] Compared with the prior art, the present invention has the following technical effects:

[0039] The image fusion network model constructed by the present invention includes a subnetwork feature extraction module for extracting the subspace features of each band group respectively, a wavelet transform module for performing multi-scale frequency domain decomposition of low-resolution hyperspectral images and high-resolution multispectral images, a CMWF module for aligning and fusing heterogeneous modal high-frequency features in the wavelet domain, an HFCL module for further optimizing the consistency of cross-modal high-frequency features in the feature space, and an image reconstruction module for reconstructing the fused features and splicing them to output a complete hyperspectral image. During the training process of the image fusion network model, the training data consists of a low-resolution hyperspectral image after Gaussian blurring and downsampling and the corresponding high-resolution multispectral image. A composite loss function including pixel reconstruction loss, spectral angle error loss, high-frequency contrast learning loss and wavelet domain high-frequency reconstruction loss is used for end-to-end supervised optimization, so that the fusion network model can take into account both spatial detail restoration and spectral consistency, effectively enhance the image's high-frequency texture and cross-modal feature alignment capabilities, thereby generating a fused image with high spatial resolution and high spectral fidelity, improving the detail reconstruction and spectral information preservation effects in complex scenes, avoiding the problems of detail blurring, cross-modal mismatch and high-frequency information loss caused by only using convolutional neural networks in existing technologies, and solving the technical defects of existing methods in high-spectral image super-resolution tasks, such as insufficient accuracy, weak texture recovery and unstable fusion.

[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 It is a schematic diagram of the network structure of the present invention;

[0043] Figure 2 It is a structural principle diagram of the CMWF module of the present invention;

[0044] Figure 3 It is a structural principle diagram of the HFCL module of the present invention;

[0045] Figure 4 This is the comparison result of the image reconstruction visual effect of the present invention and various existing methods on the MDC dataset;

[0046] Figure 5 This is the comparison result of the image reconstruction visual effect of the present invention and various existing methods on the Pavia dataset;

[0047] Figure 6 This is the comparison result of the visual effect of image reconstruction performed by the present invention on the Harvard dataset with that of various existing methods;

[0048] Figure 7 The box plots of PSNR and SSIM indicators of various methods under different band grouping settings on the MDC dataset are compared, showing the performance distribution differences between the method of the present invention and other fusion methods under different scale conditions;

[0049] Figure 8 The comparison results of the average PSNR and SSIM index histograms of various methods under different scaling factor settings on the MDC dataset are used to evaluate the comprehensive performance of the method of the present invention in different resolution restoration tasks;

[0050] Figure 9 Figure 3 is a graph of spectral vector consistency analysis results of the method of the present invention and other methods at two different spatial locations selected in the Pavia dataset, which is used to compare the performance differences of each method in maintaining spectral characteristics. DETAILED DESCRIPTION

[0051] To make the above-mentioned objects, features, and advantages of the present invention more readily apparent, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. The following description sets forth numerous specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0052] A hyperspectral and multispectral image fusion method based on wavelet feature fusion and contrast learning includes the following steps:

[0053] Step S1: constructing a fusion network model, wherein the fusion network model includes a wavelet transform module, a CMWF module, a HFCL module and an image reconstruction module;

[0054] Step S2: Obtain the low-resolution hyperspectral image and the corresponding high-resolution multispectral image to form a training dataset. Use the training dataset to train the fusion network model. Construct a loss function to optimize the network training.

[0055] Step S3: Input the low-resolution hyperspectral image to be fused and the corresponding high-resolution multispectral image into the fusion network model trained in Step S2. After feature extraction, fusion, alignment, and reconstruction processing, generate the final high-resolution hyperspectral image output.

[0056] The above method steps are further described in detail as follows:

[0057] Step S1. As Figure 1 shown, construct a fusion network model. The fusion network model includes a wavelet transform module, a CMWF module, a HFCL module, and an image reconstruction module. Each module is connected in sequence to complete the reconstruction process from the input low-resolution hyperspectral image and high-resolution multispectral image to the output high-resolution hyperspectral image. In the present invention, the set of low-resolution hyperspectral images and high-resolution multispectral images in each batch input into the fusion network is denoted as where h << H, w << W, c << C; I is the number of samples in the current batch. For each sample in a batch, first perform bicubic interpolation upsampling on the low-resolution hyperspectral image X i to obtain i with the same spatial resolution as Y

[0058] Specifically, as Figure 1 shown, before inputting into the network, perform band grouping preprocessing on according to the spectral correlation between bands. Let the set after band grouping of be where K is the total number of groups, |Ψ k | is the number of bands in group k. Define the k-th subnet function as f k (·), whose input is the corresponding hyperspectral image sub-band and the high-resolution multispectral image Y i , and the output is the reconstructed high-resolution sub-band The final reconstructed image is obtained by stitching the outputs of each subnet:

[0059] As Figure 1 shown, each subnet processing unit processes the input and Y iProcessing. First, the wavelet transform module is used to decompose the input into low-frequency components (LL) and high-frequency components in three directions (LH, HL and HH). The CMWF module receives high-frequency components, especially for the directional high-frequency features of LH, HL and HH, for cross-modal fusion. The CMWF module consists of cascaded cross-modal wavelet feature fusion blocks (CMWFB), which enhances the fusion of texture and detail information through attention mechanisms and other methods. The HFCL module receives the fused features and part of the wavelet components output by the CMWF module, and uses contrastive learning to further refine and align the high-frequency features to enhance the collaborative representation capability of multimodal data. The image reconstruction module receives the processed high-frequency components and the low-frequency components obtained by SWT decomposition, inversely transforms them back to the spatial domain, and generates reconstructed high-resolution sub-bands corresponding to the band group.

[0060] Specifically, the CMWF module is used to fuse the high-frequency details of hyperspectral and multispectral images in the frequency domain, especially for the high-frequency features of LH, HL and HH directions. Figure 1 As shown in Figure 1, the network cascades multiple CMWF modules, which are composed of multiple CMWFBs in parallel. Each CMWFB processes the high-frequency features in one direction. The first CMWF module receives the high-frequency component decomposed by SWT as input, and the subsequent cascade modules receive the output of the previous CMWF module and the initial Y i High-frequency wavelet components. The CMWF module uses an attention-based cross-modal fusion mechanism to selectively integrate high-frequency details from multispectral and hyperspectral images, improving the ability to express spatial details and preserve spectral information.

[0061] Each CMWFB structure is as follows Figure 2 As shown in Figure 1, it receives high-frequency features as input and selectively integrates high-frequency details from different modalities through an attention-based cross-modal fusion mechanism. The input features are first processed by 3×3 convolution, batch normalization, and LeakyRelu activation function. Subsequently, the processed features are reshaped and used to generate Q, K, and V, respectively. A cross-attention mechanism is adopted in the CMWF module, Q and K are applied to features from different modalities, and V uses high-frequency features from high-resolution MSI to emphasize detail preservation. By calculating K·Q T Softmax is applied to obtain an attention map, which is then multiplied by V to obtain a weighted fused feature. In order to preserve the original information and facilitate gradient flow, the weighted fused feature is finally added to the input feature of CMWFB (residual connection) to obtain the output of CMWFB.

[0062] Specifically, the HFCL module is used to refine and align the fused high-frequency features output by the CMWF module to enhance the collaborative representation capability of cross-modal data. Figure 3As shown. The HFCL module receives the fused high-frequency features output by the CMWF module and Y i The initial HH component of is taken as input, and the positive and negative sample selection strategy based on the directionality of wavelet features is used for contrastive learning. For example, the HH component in the fusion output is used as the anchor point, Y i The HH component of the fused image is used as the positive sample, and the LH and HL components of the fused image are used as the negative sample. These samples are projected into the latent space through an MLP layer. The MLP module consists of two sequentially connected fully connected layers with a GELU activation function inserted in between. By optimizing the contrastive loss, the model learns to distinguish textures in different directions, enhancing the discriminative power and alignment of features.

[0063] Specifically, the image reconstruction module is used to inversely transform the processed frequency domain components, including the high-frequency components refined by the HFCL module and the low-frequency components decomposed by the SWT, back to the spatial domain to generate the reconstructed high-resolution subband corresponding to the band group. The final reconstructed image Obtained by splicing the outputs of each subnetwork.

[0064] Step S2: Obtain training data set and train network model. i With high-resolution multispectral image Y i and the corresponding high-resolution hyperspectral image Z as the reference image i The training data set consists of i Perform spatial downsampling and blurring to synthesize low-resolution hyperspectral image X i , to simulate the situation where the spatial resolution of the hyperspectral imaging sensor is low due to hardware limitations. In the present invention, Z is first blurred using a spatial degradation function (applying a 7×7 Gaussian filter with a standard deviation of 2), and then the spatial resolution is reduced by downsampling (applying d×d average pooling or nearest neighbor sampling, where d is the scaling factor). Specifically, in the present invention, the scaling factor d is set to 4 or 8 for the MDC and Pavia datasets; and 8 or 16 for the Harvard dataset. High-resolution multispectral image Y i There are two ways to obtain Z: one is to simulate the spectral response function of the multispectral sensor i Processing is done, Z i The spectral dimension information of Z is mapped to a small number of multispectral bands through linear combination, which simulates the spectral response in the multispectral image acquisition process; the other is to use i Multispectral images with high spatial resolution collected simultaneously or approximately simultaneously. In the present invention, for the Pavia dataset, Y i It is through Z iThe adjacent bands are averaged and simulated; for Harvard and MDC data sets, the corresponding color images are directly provided as Y i . The synthesized or acquired X i and Y i And the corresponding Z i To increase the number of training samples and improve the generalization ability of the model, the present invention crops these image pairs into small-sized image blocks (e.g., 64×64 or 128×128).

[0065] Using the training dataset, the constructed WaveCLN network model is trained end-to-end with a composite loss function consisting of reconstruction loss, spectral angle loss, contrastive learning loss, and wavelet domain loss to improve the spatial resolution, spectral consistency, and texture detail recovery of the image. The composite loss function is composed of reconstruction loss, spectral angle loss, contrastive learning loss, and wavelet domain loss, aiming to jointly optimize the model's performance. This composite loss function can be expressed as a weighted sum of the individual losses, and is defined as follows:

[0066]

[0067] Among them, L rec is the pixel-level L1 reconstruction loss, which is used to constrain the pixel difference between the generated image and the reference image; L cl It is a high-frequency contrast learning loss that promotes the consistency of cross-modal feature space by optimizing the similarity of positive and negative samples; L sam is the spectral angle error loss, which is used to measure the angle consistency between the generated image and the reference image in the spectral direction and reduce spectral distortion; L swt is the high-frequency reconstruction loss in the wavelet domain, which is based on the high-frequency reconstruction loss of wavelet grouping. It improves the quality of detail and boundary restoration by performing fine-grained supervision on each high-frequency subband in the wavelet domain. I represents the number of samples in a batch; α, β, and γ are weighting coefficients, which are adjusted according to different task requirements.

[0068] Pixel-level L1 reconstruction loss L rec It is used to constrain the consistency of the network output image and the reference image in pixel space, ensuring that the basic spatial structure and brightness information are restored. It is defined using the L1 norm, which calculates the sum of the L1 norm differences between the reconstruction results of all band groups and the corresponding reference image. The definition formula is:

[0069]

[0070] in, is the reconstructed hyperspectral image of the i-th sample in a batch; Z i is the corresponding reference high-resolution hyperspectral image; g(·) represents the L1 loss calculation.

[0071] High-frequency contrastive learning loss L cl It is used to enhance the alignment of high-frequency features across modalities. By shortening the distance between similar samples (anchors and positive samples) in the latent space and pushing the distance between dissimilar samples (anchors and negative samples) further apart, it optimizes the discriminative power of feature representation and reduces the radiation and spatial differences between multimodal data. Its calculation method is based on the anchor samples, positive samples, and negative samples generated by the HFCL module (refer to step S1 above). The similarity sim(·) is calculated using the normalized vector dot product and is scaled with the temperature hyperparameter τ. For samples in a batch, the contrast loss is defined as:

[0072]

[0073] Among them, A i represents the i-th anchor point sample in a batch; P i is a positive sample; is a negative sample; sim(·) represents the dot product similarity after vector normalization; τ is the temperature scaling parameter; I represents the number of samples in a batch.

[0074] Spectral angle error loss L sam Used to measure the reconstructed image With the reference image Z i The spectral angle difference between them ensures spectral fidelity and is used to penalize the angular deviation between the reconstructed spectral vector and the reference spectral vector. It calculates the spectral angle between the spectral vector of each pixel in all band groups and the corresponding reference spectral vector and averages them. The definition formula is:

[0075]

[0076] in, Represents the predicted spectrum vector of the i-th sample in a batch on the band group k; represents the corresponding reference spectrum vector; h(·) represents the cosine similarity calculation.

[0077] Wavelet domain high frequency group reconstruction loss L swt Focus on preserving the reconstructed image With the reference image Z i The high-frequency component details of the image are measured by L1 differences in high-frequency subbands. This is used to provide fine-grained supervision of the details of the reconstructed image in the frequency domain, encouraging the model to restore clear textures and boundaries while avoiding the introduction of high-frequency artifacts. It calculates the sum of the L1 norm differences between the wavelet coefficients of all band groups in the three directions of LH, HL, and HH and the corresponding reference wavelet coefficients. The definition formula is:

[0078]

[0079] Among them, m is the direction of the wavelet high-frequency subband, and its values ​​include horizontal direction (LH), vertical direction (HL) and diagonal direction (HH); Represents the feature representation of the reconstructed image of the i-th sample in a batch on the m-th wavelet subband; represents the feature representation of the reference image on the same wavelet subband; g(·) represents the L1 loss calculation.

[0080] In this paper, the model is implemented based on the PyTorch framework and trained on an Nvidia RTX 4060 GPU with a batch size of 16 for 200 epochs. The training process is divided into two stages: in the first stage, the model is trained for 100 epochs with hyperparameters set to α = 0.05, β = 0.05, and γ = 0; in the second stage, the hyperparameters are adjusted to α = 0.01, β = 0.01, and γ = 0.03, and training is continued for 100 epochs. The model parameters are updated using the AdamW optimizer, and the initial learning rate is set to 6×10 -4 .

[0081] Step S3: input the low-resolution hyperspectral image to be fused and the corresponding high-resolution multispectral image into the fusion network model trained in step S2 to generate a high-resolution hyperspectral image.

[0082] In order to further illustrate the effect of the present invention, the effect of the present invention is verified by comparative experiments below:

[0083] The experiments used three public datasets covering remote sensing, medical imaging, and natural scenes. The Pavia Centre dataset contains aerial remote sensing hyperspectral images collected in the central area of ​​Pavia, Italy. The dataset contains 1096 × 715 effective pixels with 102 spectral bands ranging from 430 to 860 nanometers. In the experiments, the upper left corner area (512 × 128 pixels) was used for testing, and the rest was used for training. The MDC dataset contains hyperspectral and color images of bile duct cancer. Each image has a resolution of 320 × 256 pixels and 50 spectral bands. The dataset is divided into 280 training samples, 101 validation samples, and 106 test samples. The training and validation samples are cropped into non-overlapping patches of 64 × 64 pixels. The Harvard dataset contains 50 real-world images covering indoor and outdoor scenes. Each image has a spatial size of 1024 × 1392 and has 31 spectral bands ranging from 420 to 730 nanometers. 30 images were randomly selected for training, and the remaining images were used for testing. The training samples are cropped into non-overlapping patches of 128 × 128 pixels and random augmentation is applied.

[0084] Evaluation indicators adopt four widely used quality indicators to quantitatively evaluate the model performance: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), Spectral Angle Mapper (SAM) and Relative Global Adimensionless Synthesizer (ERGAS). The higher the PSNR and SSIM values, the better the quality of the fused HR-HSI; the lower the SAM and ERGAS values, the better the performance. The effectiveness of the fusion model proposed in this embodiment was evaluated and verified by comparison with a variety of advanced methods, including two traditional methods (HySure and NSSR) and six deep learning-based methods (3DFCNN, SSRNet, ResTFNet, Fusformer, HSRNet and GuidedNet).

[0085] The following table is Table 1: Comparison of quantitative indicators of the MDC dataset under different scaling factors

[0086]

[0087] The following table is Table 2: Comparison of quantitative indicators of Harvard dataset under different scaling factors

[0088]

[0089] The following table is Table 3: Comparison of quantitative indicators of the Pavia dataset under different scaling factors

[0090]

[0091] The experimental results of the method proposed in this embodiment on the three aforementioned datasets are shown in Tables 1, 2, and 3 above. This embodiment achieved state-of-the-art performance on all evaluation indicators. In terms of quantitative indicators, this embodiment generally achieved the highest scores on PSNR and SSIM, and the lowest scores on SAM and ERGAS, which shows that this embodiment can simultaneously achieve excellent spatial detail recovery and high spectral fidelity. In terms of visual quality, by comparing the reconstructed images and error maps generated by different methods, as shown in Figure 2, the reconstructed images and error maps generated by different methods are shown in Figure 2. Figure 4 、 Figure 5 and Figure 6 As shown, this embodiment can better restore fine textures and structures, reduce artifacts, and generate images that are closer to the reference image. For example, on the MDC dataset, this embodiment can clearly restore tissue textures. Figure 4; On the Pavia dataset, this embodiment can effectively process remote sensing scenes and maintain the structure and spectral characteristics of ground objects. Figure 5 ; On the Harvard dataset, this embodiment can retain complex details in natural scenes and reduce errors, refer to Figure 6 .

[0092] like Figure 7 As shown in Figure 2, box plots of PSNR and SSIM performance of different methods under different spectral groups on the MDC dataset are shown. Figure 7 (a) and 7(b) show the PSNR results under scaling factors of 4 and 8, Figure 7 Figures (c) and 7(d) show the SSIM results for the same scaling factor. The boxplots clearly demonstrate the performance distribution and stability of each method across different spectral groupings. This example demonstrates excellent performance across all groupings, with a more compact performance distribution and a higher median. This further demonstrates the effectiveness of the spectral grouping strategy and subnetwork processing employed in this invention, enabling more accurate fusion of specific spectral regions.

[0093] like Figure 8 As shown in Figure 2, the average PSNR and SSIM performance curves of different methods in each band are shown on the MDC dataset. Figure 8 (a) and 8(b) show the average PSNR curves under scaling factors of 4 and 8, Figure 8 Figures (c) and 8(d) show the average SSIM curves at the same scaling factor. As can be seen from the figures, this embodiment maintains the highest PSNR and SSIM values ​​across most bands, further verifying the robustness of this embodiment and its ability to maintain high reconstruction quality across different spectral regions.

[0094] like Figure 9 As shown, the spectral vector analysis results of two spatial position pixels (for example, (256,127) and (48,48)) selected in the Pavia dataset are shown. Figure 9 (a) shows the spectrum vector at a scaling factor of 4, Figure 9 (b) shows the spectral vector at a scaling factor of 8. The figure compares the spectral vector of the reference image (black curve) and the spectral vectors of the images reconstructed at that position using different methods. It can be seen that the spectral vector generated by this embodiment (green curve) is closest to the reference spectral vector, and the curve morphology is highly consistent, especially in complex spectral regions. This shows that this embodiment performs well in maintaining spectral fidelity. The HFCL module plays a significant role in aligning cross-modal features and preserving subtle spectral changes, effectively avoiding spectral distortion that may occur with other methods.

[0095] Through the above steps, the present invention achieves super-resolution reconstruction of hyperspectral images. Compared with existing technologies, this invention more effectively utilizes frequency domain information by introducing wavelet transforms and grouping processing. The CMWF and HFCL modules enhance the alignment and fusion quality of high-frequency features of heterogeneous modal data. A comprehensive loss function is used to balance pixel accuracy, spectral fidelity, and detail reconstruction. Experimental results demonstrate that this embodiment outperforms existing methods in both quantitative metrics and visual effects, generating hyperspectral images with richer detail and more consistent spectra, while also being less susceptible to information loss and improving processing efficiency.

[0096] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Any person skilled in the art can utilize the methods and technical contents disclosed above to make many possible variations and modifications to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, or modify them into equivalent embodiments with equivalent variations. Therefore, any equivalent variations made in accordance with the shape, structure, and principles of the present invention without departing from the content of the technical solutions of the present invention should be included in the scope of protection of the present invention.

Claims

1. A hyperspectral and multispectral image fusion method based on wavelet feature fusion and contrast learning, characterized in that: The following steps are involved: Step S1: constructing a fusion network model, wherein the fusion network model includes a wavelet transform module, a cross-modal wavelet feature fusion module, a high-frequency contrast learning module, and an image reconstruction module; Step S2: Obtain a training dataset, wherein the training dataset includes low-resolution hyperspectral images, high-resolution multispectral images, and their corresponding reference high-resolution hyperspectral images, construct training pairs by combining images of different modalities, and perform end-to-end supervised training using a composite loss function; Step S3: Input the low-resolution hyperspectral image to be fused and the corresponding high-resolution multispectral image into the trained fusion network model, and output the high-resolution hyperspectral image reconstruction result.

2. The method according to claim 1, wherein In step S1, the wavelet transform module uses stationary wavelet transform to perform multi-scale decomposition on the input image, extracting low-frequency sub-bands and high-frequency sub-bands in three directions respectively, and all sub-bands maintain the same spatial size as the original image.

3. The method according to claim 1, wherein In step S1, the cross-modal wavelet feature fusion module consists of multiple parallel fusion units. Each fusion unit receives the sub-band features of the hyperspectral image and the multispectral image in a high-frequency direction, and uses the attention mechanism to fuse the high-frequency features of the two modalities.

4. The method according to claim 1, wherein In step S1, the high-frequency contrast learning module includes a sample sampler, an embedding encoder and a loss calculation unit. The sample sampler is used to extract anchor samples, positive samples and negative samples from the fused high-frequency features. The embedding encoder is a multi-layer perceptron structure. The samples are projected into the latent space through the MLP layer. The MLP module includes two sequentially connected fully connected layers with a GELU activation function inserted in the middle. By optimizing the contrast loss, the model learns to distinguish textures in different directions and enhance the discriminative power and alignment effect of the features. The contrast loss is constructed based on the temperature-scaled dot product contrast loss function.

5. The method according to claim 1, wherein In step S1, the image reconstruction module includes an inverse stationary wavelet transform module, which takes as input the fused low-frequency sub-band and the directional high-frequency sub-band. A high-resolution image is generated by inverse wavelet reconstruction, and the reconstructed multiple sub-network output channels are spliced ​​to obtain the final image.

6. The method according to claim 1, wherein In step S2, the composite loss function is: Among them, L rec is the pixel-level L1 reconstruction loss; L cl is the high-frequency contrastive learning loss; L sam is the spectral angle error loss; L swt is the high-frequency reconstruction loss in the wavelet domain; I represents the number of samples in a batch; α, β, γ are weighting coefficients.

7. The method according to claim 6, wherein Pixel-level L1 reconstruction loss L rec for: in, is the reconstructed hyperspectral image of the i-th sample in a batch; Z i is the corresponding reference high-resolution hyperspectral image; g(·) represents the L1 loss calculation.

8. The method according to claim 6, wherein High-frequency contrastive learning loss L cl for: Among them, A i represents the i-th anchor point sample in a batch; P i is a positive sample; is a negative sample; sim(·) represents the dot product similarity after vector normalization; τ is the temperature scaling parameter; I represents the number of samples in a batch.

9. The method according to claim 6, wherein Spectral angle error loss L sam for: in, Represents the predicted spectrum vector of the i-th sample in a batch on the band group k; represents the corresponding reference spectrum vector; h(·) represents the cosine similarity calculation.

10. The method according to claim 6, wherein High-frequency reconstruction loss L in wavelet domain swt for: Among them, m is the direction of the wavelet high-frequency subband, and its values ​​include horizontal, vertical and diagonal directions; Represents the feature representation of the reconstructed image of the i-th sample in a batch on the m-th wavelet subband; represents the feature representation of the reference image on the same wavelet subband; g(·) represents the L1 loss calculation.

Citation Information

Cited By

  • Highway pavement disease intelligent detection method and system based on multi-source data fusion and YOLO optimization algorithm

    CN121190981A

  • Medical hyperspectral image enhancement method and system based on multi-domain fusion

    CN121582077A

  • Hyperspectral image fusion reconstruction method based on cross-modal interaction and difference perception

    CN122289033A

  • A hyperspectral image imaging method, device, system and storage medium

    CN122454346A

  • A hyperspectral image reconstruction method of a spatial spectral feature fusion double-branch network

    CN122617637A