Hyperspectral and multispectral data fusion method of Transform + KAN network

The hyperspectral and multispectral data fusion method using Transformer+KAN network solves the problems of spectral distortion and loss of spatial details in existing technologies, achieving high-fidelity data fusion that is suitable for complex natural scenes.

CN121746916APending Publication Date: 2026-03-27BEIJING RES INST OF SPATIAL MECHANICAL & ELECTRICAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies suffer from spectral distortion or loss of spatial details in hyperspectral and multispectral data fusion, and deep learning methods have failed to effectively mine multimodal data features, making it difficult to achieve high-fidelity fusion results.

Method used

A hyperspectral and multispectral data fusion method using Transformer+KAN network is proposed. By constructing a training dataset and a multispectral and hyperspectral fusion network, including shallow feature extraction, cross-attention fusion and deep fusion modules, and combining it with the Swin Transformer-KAN structure, cross-modal feature fusion and reconstruction are achieved.

Benefits of technology

It achieves high-fidelity hyperspectral and multispectral data fusion, fully explores long-distance dependencies between features, improves the flexibility and interpretability of fusion results, and is suitable for complex natural scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746916A_ABST
    Figure CN121746916A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral and multispectral data fusion method of a Transform + KAN network. The hyperspectral and multispectral data fusion method comprises the following steps: S1, constructing a training data set: acquiring high-resolution and hyperspectral reference data HR-HSI-Benchmark, low-resolution and hyperspectral input data LR-HSI-Input to be fused and high-resolution and multispectral input data HR-MS-Input to be fused; after radiation correction, geometric correction, spectrum matching and down-sampling processing are carried out on the high-resolution hyperspectral reference data, low-resolution hyperspectral data LR-HSI-Train for training and high-resolution multispectral data HR-MSI-Train for training are generated; the LR-HSI-Train and the HR-MSI-Train are combined to form a low-resolution hyperspectral-high-resolution multispectral-high-resolution hyperspectral reference data training sample set; s2, constructing a multispectral and hyperspectral fusion network; s3, performing data amplification on the training sample set, and inputting the training sample set into a multispectral hyperspectral fusion network to obtain a trained fusion model; s4, registering the hyperspectral and multispectral data to be fused; and S5, image fusion and output.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a hyperspectral multispectral data fusion method of a Transformer+KAN network and belongs to the technical field of remote sensing image processing and computer vision. BACKGROUND

[0002] Satellite remote sensing images are widely used in ecological environmental protection, resource exploration, emergency disaster reduction, city planning and other aspects. High-resolution multispectral satellite remote sensing images have rich spatial texture details and are easy to be recognized by humans and interpreted by machines, but they lack spectral information of ground objects and are difficult to effectively carry out fine classification of ground objects (such as tree species and crop identification) and quantitative inversion of key parameters (such as vegetation surface leaf area index and water suspended matter concentration). Hyperspectral satellite remote sensing images can obtain fine spectra of different ground objects, realize fine classification through reflection and absorption characteristics, and realize quantitative inversion of key parameters of ground objects by combining with a radiation transmission quantitative model, but due to limitations of optical systems, sensors and other processes, the current mainstream hyperspectral satellite images generally have low spatial resolution and cannot identify ground objects smaller than the spatial scale. The fusion of high-resolution multispectral satellite remote sensing images and hyperspectral satellite remote sensing images can obtain satellite images with high spatial resolution and high spectral resolution, effectively compensate for the defects of the two kinds of data, and greatly expand the application potential.

[0003] Traditional fusion methods mainly include a detail injection method and a model optimization method, and need to manually design prior constraints for solving, which has poor applicability to complex natural scenes and is prone to cause spectral distortion or loss of spatial details. With the improvement of hardware capability and the perfection of algorithm principle, deep learning methods have been widely used in the field of remote sensing image processing. Data-driven fusion based on deep learning methods can automatically and effectively mine the intrinsic characteristics of multi-modal data. However, current deep learning methods have problems such as insufficient mining of spectral characteristics of multispectral and hyperspectral data and poor fusion quality, and most of the methods are still in the data testing stage and cannot effectively serve the application needs of hyperspectral data. Therefore, it is necessary to propose a hyperspectral multispectral data fusion method of a Transformer+KAN network. SUMMARY

[0004] The technical problem to be solved by the application is to overcome the shortcomings of the prior art and provide a hyperspectral multispectral data fusion method of a Transformer+KAN network to obtain a high-fidelity fusion result.

[0005] The technical solution of the application is: The application discloses a hyperspectral multispectral data fusion method of a Transformer+KAN network, which comprises the following steps: S1, constructing a training data set: obtaining high-resolution hyperspectral benchmark data HR-HSI-Benchmark, to-be-fused low-resolution hyperspectral input data LR-HSI-Input, and to-be-fused high-resolution multispectral input data HR-MS-Input; after the high-resolution hyperspectral benchmark data is processed through radiation correction, geometric correction, spectral matching, and downsampling, low-resolution hyperspectral data LR-HSI-Train for training and high-resolution multispectral data HR-MSI-Train for training are generated; the low-resolution hyperspectral data LR-HSI-Train for training and the high-resolution multispectral data HR-MSI-Train for training are combined to form a low-resolution hyperspectral-high-resolution multispectral-high-resolution hyperspectral benchmark data training sample set; S2, constructing a multispectral hyperspectral fusion network: S3, inputting the training sample set after data augmentation into the multispectral hyperspectral fusion network to obtain a trained fusion model; S4, to-be-fused hyperspectral multispectral data registration: performing image space registration alignment on the LR-HSI-Input and the HR-MS-Input to obtain registered data; S5, image fusion and output: inputting the registered data into the trained fusion model to output a high-resolution hyperspectral fusion image.

[0006] Further, in the above method, the multispectral hyperspectral fusion network comprises a shallow feature extraction module, a cross-attention fusion module, a deep fusion module, and a fusion image reconstruction module; wherein the shallow feature extraction module is a double-branch structure, which extracts spectral features and spatial features from low-resolution hyperspectral data and high-resolution multispectral data, respectively; the cross-attention fusion module performs cross-attention fusion on the spectral features and the spatial features to obtain cross-modal fusion features; the deep fusion module adopts an improved structure of Swin Transformer-KAN, generates deep fusion features according to the cross-modal fusion features; and the fusion image reconstruction module reconstructs fused high-resolution hyperspectral fusion data according to the deep fusion features and outputs a high-resolution hyperspectral fusion image.

[0007] Further, in the above method, the selection of the data source in step S1 satisfies the following principles simultaneously: Principle 1: The imaging time difference of the to-be-fused low-resolution hyperspectral input data LR-HSI-Input and the to-be-fused high-resolution multispectral input data HR-MS-Input is less than 7 days, and the imaging area overlap degree is ≥95% or adjacent without spatial discontinuity; Principle 2: Diversity of ground objects: the imaging area of the high-resolution hyperspectral benchmark data HR-HSI-Benchmark, the low-resolution hyperspectral input data LR-HSI-Input to be fused and the high-resolution multispectral input data HR-MS-Input to be fused covers at least three different ground objects in vegetation, water body, building and bare soil; Principle 3: Data quality: the imaging area of the high-resolution hyperspectral benchmark data HR-HSI-Benchmark, the low-resolution hyperspectral input data LR-HSI-Input to be fused and the high-resolution multispectral input data HR-MS-Input to be fused is free of cloud, fog, shadow obstruction and atmospheric pollution, and the spectral and spatial information is pure; Principle 4: Spatial resolution requirement: the spatial resolution of the high-resolution hyperspectral benchmark data HR-HSI-Benchmark is at least 2 times or more than that of the low-resolution hyperspectral input data LR-HSI-Input to be fused, and is the same as the resolution of the high-resolution multispectral input data HR-MS-Input to be fused; Principle 5: Spectral channel requirement: the spectral channel coverage range of the high-resolution hyperspectral benchmark data HR-HSI-Benchmark is the same as or greater than the coverage range of the low-resolution hyperspectral input data LR-HSI-Input to be fused and the high-resolution multispectral input data HR-MS-Input to be fused.

[0008] Further, in the above method, the spectral matching and down-sampling processing in step S1 specifically includes: Spectrum response functions of the HR-HSI-Benchmark and the LR-HSI-Input are extracted, and effective spectral channels with a band overlap degree of ≥90% are screened; The HR-HSI-Benchmark after spectral matching is first denoised by Gaussian filtering with a standard deviation σ=1.2, and then down-sampled by bilinear interpolation to be consistent with the resolution of the LR-HSI-Input, to generate low-resolution hyperspectral data LR-HSI-Train for training; HR-HSI-Benchmark is down-sampled based on the spectrum response function of the HR-MS-Input, to generate intermediate data consistent with the number of bands of the HR-MS-Input; the intermediate data is then denoised by Gaussian filtering with a standard deviation σ=1.2, and then down-sampled by bilinear interpolation to be the same as the resolution of the HR-MS-Input, to generate high-resolution multispectral data HR-MSI-Train for training.

[0009] Further, in the above method, the construction of the training sample set in step S1 is as follows: the HR-HSI-Benchmark and the HR-MSI-Train are cropped by using 256x256 pixel blocks to obtain the cropped HR-HSI-Benchmark and the cropped HR-MSI-Train; the cropped LR-HSI-Train is obtained by cropping the LR-HSI-Train; wherein s represents the multiple of the difference between the spatial resolution of the HR-MS-Input and the spatial resolution of the LR-HSI-Input; the cropped HR-HSI-Benchmark, the cropped HR-MSI-Train, and the cropped LR-HSI-Train form a training sample; and all the training samples are divided into a training set and a validation set in a ratio of 8:2, wherein the cropped HR-HSI-Benchmark is used as the fusion ground truth. The cropped LR-HSI-Train is obtained by cropping the LR-HSI-Train; wherein s represents the multiple of the difference between the spatial resolution of the HR-MS-Input and the spatial resolution of the LR-HSI-Input; the cropped HR-HSI-Benchmark, the cropped HR-MSI-Train, and the cropped LR-HSI-Train form a training sample; and all the training samples are divided into a training set and a validation set in a ratio of 8:2, wherein the cropped HR-HSI-Benchmark is used as the fusion ground truth.

[0010] Further, in the above method, the specific structure of the shallow feature extraction module in step S2 is as follows: The spectral feature branch: the up-sampled low-resolution hyperspectral data is sequentially passed through two 3x3 convolution layers, a KAN spectral modeling network, and a KAN channel attention network to output a shallow spectral feature F_s; the KAN spectral modeling network includes two hidden layers, each layer has 128 nodes, and uses a 3rd-order piecewise polynomial activation function; The spatial feature branch: the high-resolution multispectral data is sequentially passed through a multi-scale convolution unit and a spatial attention network to output a shallow spatial feature F_sp; the multi-scale convolution unit includes parallel 1x1, 3x3, and 5x5 convolution layers.

[0011] Further, in the above method, the cross-attention fusion is specifically as follows: The query vector Q is obtained from the shallow spectral feature F_s, the key vector K and the value vector V are obtained from the shallow spatial feature F_sp, cross-attention calculation is performed, and a first calculation result is obtained; The query vector Q is obtained from the shallow spatial feature F_sp, the key vector K and the value vector V are obtained from the shallow spectral feature F_s, cross-attention calculation is performed, and a second calculation result is obtained; After the first calculation result and the second calculation result are spliced, they are input into a convolution layer, and a cross-attention fusion result F_cross is output.

[0012] Further, in the above method, the deep fusion module in step S2 includes three Swin Transformer modules and one 3x3 convolution layer; wherein, The 3-layer Swin Transformer module is cascaded with a 3*3 convolution layer, and the cross-modal fusion feature F_cross is subjected to cascade processing, and a deep fusion feature F_deep is output. The window size of the Swin Transformer module is 7*7, and the shift window step is 3. The Swin Transformer module is a KAN network replacing the output layer multi-layer perception (MLP) of the traditional Swin Transformer, wherein the KAN network comprises one hidden layer and 256 nodes, and the activation function is a 3-order segmented polynomial.

[0013] Further, in the above method, the fusion image reconstruction module in step S2 comprises a 3*3 convolution layer; wherein the input deep fusion feature is output as a high-resolution hyperspectral fusion image through the 3*3 convolution layer by using a Sigmoid activation function.

[0014] Further, in the above method, the data augmentation in step S3 is specifically: performing horizontal or vertical flipping, 90°, 180° or 270° rotation and adding Gaussian noise (variance, probability 0.2) on the training sample set; wherein the probability of horizontal or vertical flipping is 0.5; the probability of 90°, 180° or 270° rotation is 0.3; the probability of adding Gaussian noise is 0.2, and the variance is 0.01-0.05.

[0015] Further, in the above method, the combined loss function in step S3 is wherein L1 is the mean absolute error loss, SSIM is the structural similarity loss, SAM is the spectral angle matching loss, a, b and c are the weights of the three error terms; the training parameters are set as follows: Adam optimizer, weight decay coefficient less than 1e-5, initial learning rate less than 1e-4, decay to less than 0.5 of the current value every 50 rounds, batch size greater than or equal to 16, total iteration number greater than 200 rounds, and stop iteration when the loss function of the verification set does not decrease for 10-30 consecutive rounds.

[0016] The beneficial effects of the present application over the prior art are: (1) The present application uses a KAN spectral modeling network for spectral feature modeling in the shallow feature extraction module, which has better mathematical principle interpretability than convolution or fully connected networks, and fully realizes flexible, accurate and interpretable spectral feature extraction. (2) The application realizes feature depth fusion by cascading Swin Transformer-KAN modules, fully utilizes the ability of the Transformer to fully mine long-distance dependency relationships between features, effectively reduces the model parameter quantity by means of the KAN network in complex modeling, and is beneficial to engineering deployment and practice.

[0017] (3) The application fully mines long-distance dependency relationships between features by means of the Transformer network, and obtains a high-fidelity fusion result by combining the flexible, accurate and interpretable advantages of the KAN network (Kolmogorov-Arnold Network). BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is the fusion technology route provided by the application; Figure 2 is the KAN-Transformer fusion network provided by the application; Figure 3 is a schematic diagram of the shallow feature extraction module of the application; Figure 4 is a CAM_KAN schematic diagram of the application; Figure 5 is a schematic diagram of the cross-attention fusion module of the application; Figure 6 is a schematic diagram of the deep fusion module of the application; Figure 7 is a Swin-Transformer-KAN schematic diagram of the application; Figure 8 is a schematic diagram of the fusion image reconstruction module of the application. DETAILED DESCRIPTION

[0019] The application will be further described in detail below in combination with the drawings and specific embodiments.

[0020] The application obtains low-resolution hyperspectral and high-resolution multispectral image training data sets by simulating and processing on-orbit hyperspectral satellite image data, inputs the two to the KAN-Transformer fusion network designed by the application for model training, selects the optimal result as the fusion model after training is completed, inputs the low-resolution hyperspectral and high-resolution multispectral image data to be fused after image registration to the model, and finally obtains high-resolution hyperspectral fusion data.

[0021] As shown in Figure 1 The embodiment provides a hyperspectral multispectral data fusion method of a Transformer+KAN network, which comprises the following steps: Step 1: Constructing a training data set (1) Obtain high-resolution hyperspectral benchmark data (HR-HSI-Benchmark), low-resolution hyperspectral input data to be fused (LR-HSI-Input), and high-resolution multispectral input data to be fused (HR-MS-Input), wherein: High-resolution hyperspectral benchmark data (HR-HSI-Benchmark): high-resolution hyperspectral data obtained in orbit, with complete spectral dimension and high spatial resolution; Low-resolution hyperspectral input data to be fused (LR-HSI-Input) and high-resolution multispectral input data to be fused (HR-MS-Input): data to be fused, input to the final model to obtain high-resolution hyperspectral fusion data.

[0022] The screening of data sources meets the following principles: Space-time matching: the imaging time difference of LR-HSI-Input and HR-MS-Input is less than 7 days, and the imaging area overlap degree is ≥95% or adjacent without spatial discontinuity; Ground object diversity: the imaging area of HR-HSI-Benchmark, LR-HSI-Input and HR-MS-Input covers at least three different ground objects among vegetation, water, buildings and bare soil; Data quality: the imaging area of HR-HSI-Benchmark, LR-HSI-Input and HR-MS-Input is free of clouds, fog, shadow obstruction and atmospheric pollution, and the spectral and spatial information is pure; Spatial resolution requirement: the spatial resolution of HR-HSI-Benchmark is at least 2 times or more than that of LR-HSI-Input and the same as that of HR-MS-Input; Spectral channel requirement: the spectral channel coverage range of HR-HSI-Benchmark is the same as or greater than that of LR-HSI-Input and HR-MS-Input.

[0023] (2) Synchronously perform preprocessing on the three types of screened data sources to eliminate system errors and environmental interference, including radiation correction and orthographic correction.

[0024] (3) After preprocessing, the high-resolution hyperspectral benchmark data is processed by spectral matching and downsampling to generate low-resolution hyperspectral data for training (LR-HSI-Train) and high-resolution multispectral data for training (HR-MSI-Train), wherein: Spectral matching: extract the spectral response functions of HR-HSI-Benchmark and LR-HSI-Input, and screen effective spectral channels with a waveband coincidence degree ≥90%; Spatial down-sampling: After spectral matching, the HR-HSI-Benchmark is first denoised by Gaussian filter with standard deviation σ=1.2, and then down-sampled to the resolution consistent with LR-HSI-Input by bilinear interpolation to generate low-resolution hyperspectral data (LR-HSI-Train) for training; Spectral-spatial joint down-sampling: Based on the spectral response function of HR-MS-Input, the HR-HSI-Benchmark is spectrally down-sampled to generate intermediate data with the same number of bands, and then adjusted to the same resolution as HR-MS-Input by interpolation with the spatial down-sampling method to generate high-resolution multispectral data (HR-MSI-Train) for training.

[0025] (4) Constructing training sample data set. The construction method of the training sample set is as follows: 256x256 pixel blocks are used to crop HR-HSI-Benchmark and HR-MSI-Train, LR-HSI-Train is cropped, where s represents the difference multiple of the spatial resolution between HR-MS-Input and LR-HSI-Input, to ensure that the cropping areas of the three are completely consistent. Each set of cropped data constitutes a training sample, and all samples are divided into a training set and a validation set in a ratio of 8:2, wherein HR-HSI-Benchmark is used as the fusion truth value.

[0026] Step 2: Constructing multi / hyperspectral fusion network The fusion network constructed by the present application is shown in the accompanying Figure 2 The fusion network constructed by the present application is shown in the accompanying (1) Shallow feature extraction module As shown in the accompanying Figure 3 , the shallow feature extraction module schematic diagram, in which Conv represents convolution operation, 1x1, 3x3, 5x5, 7x7 represent convolution kernel size; KAN represents KAN network; CAM_KAN represents KAN channel attention module; SAM represents spatial attention module; Figure 4 is a schematic diagram of CAM_KAN, in which GAP represents global pooling operation, and ReLu represents ReLu activation.

[0027] Spectral feature branch: a. 2D convolution layer: After up-sampling, the input of LR-HSI is input to this branch, which first realizes preliminary extraction of spectral features through two 3x3 convolution layers; b. KAN spectral modeling: the 2D convolution layer output is input into the KAN network (2 layers of hidden layers, 128 nodes per layer, 3-order segmented polynomial activation) pixel by pixel to realize fine modeling of spectral features; c. KAN channel attention: the KAN spectral modeling output is reshaped into the original image size and input into the KAN channel attention network, and the channel attention weight is obtained through global pooling-KAN module transformation-Relu activation-KAN module transformation. After being multiplied by the input of the KAN channel attention network, the shallow spectral feature F_s is output.

[0028] Spatial feature branch: a. Multi-scale convolution unit: the HR-MSI is input into this branch, and is parallelly passed through 1x1, 3x3 and 5x5 convolution layers to extract multi-scale features, respectively, and finally is element-wise fused through 1x1 convolution; b. Spatial attention network: the multi-scale convolution unit output is input into the spatial attention network, and the shallow spatial feature F_sp is output.

[0029] (2) Cross-attention fusion module It contains two branches, wherein the calculation process of the first cross branch is that the shallow spectral feature F_s is linearly transformed to obtain the query vector Q, the shallow spatial feature F_sp is linearly transformed to obtain the key vector K and the value vector V, and the cross-attention calculation-layer normalization is performed, and then the residual connection is performed with the query vector Q, and then the multi-layer perceptron and the layer normalization are performed, and then the residual connection is performed with the input of the multi-layer perceptron, and finally the intermediate variable F_s_cross is output. The calculation process of the second cross branch is that the shallow spatial feature F_sp is linearly transformed to obtain the query vector Q, the shallow spectral feature F_s is linearly transformed to obtain the key vector K and the value vector V, and the cross-attention calculation-layer normalization is performed, and then the residual connection is performed with the query vector Q, and then the multi-layer perceptron and the layer normalization are performed, and then the residual connection is performed with the input of the multi-layer perceptron, and finally the intermediate variable F_sp_cross is output. The F_s_cross and the F_sp_cross are spliced and input into the 3x3 convolution layer, and the cross-attention fusion result F_cross is output.

[0030] (3) Deep fusion module F_cross is input into the deep fusion module, and a 3-layer Swin Transformer-KAN and a 3x3 convolution layer are cascaded to realize deep feature mining, wherein the output layer (multi-layer perception, MLP) of the traditional Swin Transformer is replaced by a KAN network (1 hidden layer, 256 nodes, 3-order segmented polynomial activation) in the structure of the Swin Transformer-KAN, so as to improve the feature mapping interpretability, the window size is 7x7, and the shift window step is 3, so as to realize local-global feature joint modeling through window division and shift; the deep fusion feature F_deep is output after cascading processing. Figure 6 Fig. 3 is a schematic diagram of the deep fusion module, Figure 7 Fig. 4 is a schematic diagram of Swin-Transformer-KAN, wherein W-MSA represents a window-based spatial attention module, and SW-MSA represents a shift window-based spatial attention module (4) Fusion image reconstruction module F_deep is input into the fusion image reconstruction module, the output channel number is adjusted through a 3x3 convolution layer to be consistent with the HR-HSI-Benchmark band number, and finally the pixel value is mapped to the [0, 1] interval through a Sigmoid function, and a high-resolution hyperspectral fusion result is output.

[0031] Step 3: training of the fusion network (1) Data augmentation: geometric transformation (horizontal / vertical flip, probability 0.5; 90° / 180° / 270° rotation, probability 0.3) and noise addition (Gaussian noise: variance 0.01-0.05, probability 0.2) are performed on the training set samples to improve the model generalization ability.

[0032] (2) Training parameter setting: Adam optimizer (weight decay coefficient 1e-5) is adopted, the initial learning rate is 1e-4, and it is attenuated to 0.5 of the current value every 50 rounds; batch size=16, the total iteration is generally 200 rounds, and the early stopping strategy (stop if the validation set loss does not decrease for 20 consecutive rounds) is adopted.

[0033] (3) Loss function design: a combined loss function is adopted:

[0034] wherein L1 is the mean absolute error loss, SSIM is the structural similarity loss, SAM is the spectral angle matching loss, and a, b and c represent the weights of the three error terms (4) Model determination: the network parameters at the time when the validation set loss is the smallest are selected as the final fusion model to ensure the optimal fusion performance.

[0035] Step 4: registration of the high / multi-spectral data to be fused Image registration is performed on the LR-HSI-Input and the HR-MS-Input.

[0036] Step 5: Image fusion and output The registered data is input into the trained fusion model, and the network is forward propagated to output the high-resolution hyperspectral fusion image.

[0037] Embodiment The embodiment provides a hyperspectral image and multispectral image fusion method combining a Kolmogorov-Arnold network (KAN) and a Transformer. The operation process of the method is as shown in the figure, and specifically includes the following steps: Figure 1 Step 1: Construct a training data set.

[0038] Step 2: Construct a multi / hyperspectral fusion network.

[0039] Step 3: Train the fusion network.

[0040] Step 4: Register the high / multispectral data to be fused.

[0041] Step 5: Image fusion and output.

[0042] The specific implementation case takes environmental disaster mitigation satellite hyperspectral data (with a spatial resolution of 100 m) as low-resolution hyperspectral data to be fused, takes Landsat-8 (with a spatial resolution of 30 m) as high-resolution multispectral data to be fused, and takes Gaofen-5 visible short-wave infrared hyperspectral camera data (with a spatial resolution of 30 m) as a fusion true value. The Gaofen-5 data meets the requirements that the spatial resolution is at least 2 times or more than the hyperspectral data to be fused, the resolution is the same as or slightly higher than the multispectral data to be fused, the spectral channel coverage range is consistent with or slightly larger than the data to be fused.

[0043] Specifically, the implementation of step 1 is as follows: (1) Obtain high-resolution hyperspectral benchmark data (HR-HSI-Benchmark), low-resolution hyperspectral input data to be fused (LR-HSI-Input), and high-resolution multispectral input data to be fused (HR-MS-Input), wherein: The high-resolution hyperspectral benchmark data (HR-HSI-Benchmark) is high-resolution hyperspectral data obtained in orbit, which has complete spectral dimension and high spatial resolution. The low-resolution hyperspectral input data to be fused (LR-HSI-Input) and the high-resolution multispectral input data to be fused (HR-MS-Input) are data to be fused, which are input into the final model to obtain high-resolution hyperspectral fusion data.​

[0044] The screening of the data sources satisfies the following principles: Temporal-spatial matching: the imaging time difference between LR-HSI-Input and HR-MS-Input is less than 7 days, and the imaging area overlap degree is ≥95% or adjacent without spatial fault; Object diversity: the imaging area of HR-HSI-Benchmark, LR-HSI-Input and HR-MS-Input covers at least 3 different objects among vegetation, water, building and bare soil; Data quality: the imaging area of HR-HSI-Benchmark, LR-HSI-Input and HR-MS-Input is free of cloud, fog, shadow obstruction and atmospheric pollution, and the spectral and spatial information is pure; Spatial resolution requirement: the spatial resolution of HR-HSI-Benchmark is at least 2 times or more than that of LR-HSI-Input and is the same as that of HR-MS-Input; Spectral channel requirement: the spectral channel coverage range of HR-HSI-Benchmark is the same as or greater than that of LR-HSI-Input and HR-MS-Input.

[0045] (2) The three types of data sources after screening are simultaneously preprocessed to eliminate systematic errors and environmental interference, including radiation correction and orthographic correction.

[0046] (3) The high-resolution hyperspectral benchmark data after preprocessing is processed by spectral matching and down-sampling to generate training low-resolution hyperspectral data (LR-HSI-Train) and training high-resolution multispectral data (HR-MSI-Train), wherein: Spectral matching: the spectral response functions of HR-HSI-Benchmark and LR-HSI-Input are extracted, and effective spectral channels with a waveband coincidence degree ≥90% are screened; Spatial down-sampling: the HR-HSI-Benchmark after spectral matching is first denoised by Gaussian filtering with a standard deviation σ=1.2, and then down-sampled by bilinear interpolation to the resolution consistent with that of LR-HSI-Input to generate training low-resolution hyperspectral data (LR-HSI-Train); Spectral-spatial joint down-sampling: the HR-HSI-Benchmark is spectrally down-sampled based on the spectral response function of HR-MS-Input to generate intermediate data with the same number of wavebands, and then interpolated and adjusted to the same resolution as HR-MS-Input by the spatial down-sampling method described in ② to generate training high-resolution multispectral data (HR-MSI-Train).

[0047] (4) Constructing the training sample dataset. The training sample set is constructed in the following manner: 256x256 pixel blocks are used to crop HR-HSI-Benchmark and HR-MSI-Train, and s indicates that the spatial resolution of HR-MS-Input and LR-HSI-Input differs by a factor of s, and the cropping areas of the three are completely consistent. Each set of cropped data constitutes a training sample, and all samples are divided into a training set and a validation set in a ratio of 8:2, wherein HR-HSI-Benchmark is used as the fusion truth value.

[0048] Specifically, the implementation of step 2 is as follows: A hyperspectral-multispectral dual-branch network is used to realize efficient fusion of spectral and spatial information through four stages of "feature extraction-cross-modal fusion-depth mining-resolution reconstruction", and the network structure is as follows: (1) Shallow feature extraction module, as shown in Figure 3 .

[0049] Spectral feature branch: a. 2D convolution layer: LR-HSI is input to this branch after upsampling, and spectral features are first extracted through two 3x3 convolution layers; b. KAN spectral modeling: the output of the 2D convolution layer is input into the KAN network (2 layers of hidden layers, 128 nodes per layer, 3-order segmented polynomial activation) pixel by pixel to realize fine modeling of spectral features; c. KAN channel attention: the KAN spectral modeling output is reshaped to the original image size and input into the KAN channel attention network, and the channel attention weight is obtained through global pooling-KAN module transformation-Relu activation-KAN module transformation, and then multiplied by the input of the KAN channel attention network to output the shallow spectral feature F_s.

[0050] Spatial feature branch: a. Multi-scale convolution unit: HR-MSI is input to this branch and is parallelly passed through 1x1, 3x3 and 5x5 convolution layers to extract multi-scale features, and finally fused element by element through a 1x1 convolution; b. Spatial attention network: the multi-scale convolution unit output is input into the spatial attention network to output the shallow spatial feature F_sp.

[0051] (2) Cross-attention fusion module as shown in Figure 5 , which is a schematic diagram of the cross-attention fusion module of the present application, wherein Q, K and V represent linear transformations of the query vector Q, the key vector K and the value vector V, LN represents layer normalization, MCA represents cross-attention calculation, and MLP represents a multi-layer perceptron.

[0052] The first cross branch includes the following calculation process: the shallow spectral feature F_s is linearly transformed to obtain a query vector Q, the shallow spatial feature F_sp is linearly transformed to obtain a key vector K and a value vector V, the cross attention calculation layer is normalized, and then the normalized result is connected in residual with the query vector Q, and then the multi-layer perceptron is connected in residual with the input of the multi-layer perceptron, and finally the intermediate variable F_s_cross is output. The second cross branch includes the following calculation process: the shallow spatial feature F_sp is linearly transformed to obtain a query vector Q, the shallow spectral feature F_s is linearly transformed to obtain a key vector K and a value vector V, the cross attention calculation layer is normalized, and then the normalized result is connected in residual with the query vector Q, and then the multi-layer perceptron is connected in residual with the input of the multi-layer perceptron, and finally the intermediate variable F_sp_cross is output. The F_s_cross and the F_sp_cross are spliced and input to a 3x3 convolution layer, and the cross attention fusion result F_cross is output.

[0053] (3) Deep fusion module The F_cross is input to the deep fusion module, and a 3-layer Swin Transformer-KAN and a 3x3 convolution layer are cascaded to realize deep feature mining, wherein the output layer (multi-layer perceptron MLP) of the traditional Swin Transformer is replaced by a KAN network (1 layer of hidden layer, 256 nodes, 3-order segmented polynomial activation) in the structure of the Swin Transformer-KAN, the feature mapping interpretability is improved, the window size is 7x7, the shift window step is 3, and the local-global feature joint modeling is realized through window division and shift; the deep fusion feature F_deep is output after cascading processing.

[0054] (4) Fusion image reconstruction module The F_deep is input to the fusion image reconstruction module, the output channel number is adjusted through a 3x3 convolution layer to be consistent with the number of HR-HSI-Benchmark bands, and finally the pixel value is mapped to the interval [0, 1] through a Sigmoid function, and a high-resolution hyperspectral fusion result is output. As shown in Figure 8 .

[0055] Specifically, the implementation of step 3 is as follows: (1) Data augmentation: geometric transformation (horizontal / vertical flip, probability 0.5; 90° / 180° / 270° rotation, probability 0.3) and noise addition (Gaussian noise: variance 0.01-0.05, probability 0.2) are performed on the training set samples to improve the model generalization ability.

[0056] (2) Training parameter setting: Adam optimizer (weight decay coefficient 1e-5) is adopted, the initial learning rate is 1e-4, and is attenuated to 0.5 of the current value every 50 rounds; batch size = 16, the total iteration is generally 200 rounds, and an early stopping strategy (stopping when the validation set loss does not decrease for 20 consecutive rounds) is adopted.

[0057] (3) Loss function design: a combination loss function is adopted:

[0058] Wherein L1 is the mean absolute error loss, SSIM is the structural similarity loss, SAM is the spectral angle matching loss, and a, b and c represent the weights of the three error terms (4) Model determination: the network parameters at the time when the validation set loss is the smallest are selected as the final fusion model to ensure the optimal fusion performance.

[0059] Specifically, the implementation of step 4 is as follows: Image registration is performed on the LR-HSI-Input and the HR-MS-Input.

[0060] Specifically, the implementation of step 5 is as follows: The registered data is input into the trained fusion model, the network is forward propagated, and a high-resolution hyperspectral fusion image is output.

[0061] Although the content of the present application has been described in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as limiting the present application. After reading the above content, various modifications and alternatives of the present application will be obvious to those skilled in the art. Therefore, the protection scope of the present application should be defined by the appended claims.

[0062] The content not described in detail in the specification of the present application belongs to the known technology of those skilled in the art.

Claims

1. A method for fusion of hyperspectral and multispectral data using a Transformer+KAN network, characterized in that, include: S1. Constructing the training dataset: Obtain the high-resolution hyperspectral benchmark data HR-HSI-Benchmark, the low-resolution hyperspectral input data to be fused LR-HSI-Input, and the high-resolution multispectral input data to be fused HR-MS-Input; After radiometric correction, geometric correction, spectral matching, and downsampling processing of the high-resolution hyperspectral benchmark data, generate the low-resolution hyperspectral data LR-HSI-Train and the high-resolution multispectral data HR-MSI-Train for training; Combine the low-resolution hyperspectral data LR-HSI-Train and the high-resolution multispectral data HR-MSI-Train to form a training sample set of low-resolution hyperspectral-high-resolution multispectral-high-resolution hyperspectral benchmark data; S2. Construct a multispectral hyperspectral fusion network: S3. After data augmentation of the training sample set, input it into the multispectral and hyperspectral fusion network to obtain the trained fusion model; S4. Registration of hyperspectral and multispectral data to be fused: Perform image spatial registration and alignment on LR-HSI-Input and HR-MS-Input to obtain the registered data; S5. Image Fusion and Output: Input the registered data into the trained fusion model and output a high-resolution hyperspectral fused image.

2. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The multispectral hyperspectral fusion network includes a shallow feature extraction module, a cross-attention fusion module, a deep fusion module, and a fused image reconstruction module. The shallow feature extraction module has a dual-branch structure, extracting spectral and spatial features from low-resolution hyperspectral data and high-resolution multispectral data, respectively. The cross-attention fusion module performs cross-attention fusion on the spectral and spatial features to obtain cross-modal fusion features. The deep fusion module uses an improved SwinTransformer-KAN structure to generate deep fusion features based on the cross-modal fusion features. The fused image reconstruction module reconstructs the fused high-resolution hyperspectral data based on the deep fusion features and outputs a high-resolution hyperspectral fusion image.

3. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The data source filtering in step S1 simultaneously satisfies the following principles: Principle 1: The imaging time difference between the low-resolution hyperspectral input data LR-HSI-Input to be fused and the high-resolution multispectral input data HR-MS-Input to be fused is less than 7 days, and the overlap of the imaging areas is ≥95% or there are no spatial breaks between adjacent areas. Principle 2: Land cover diversity: The imaging area of ​​the high-resolution hyperspectral benchmark data HR-HSI-Benchmark, the low-resolution hyperspectral input data to be fused LR-HSI-Input, and the high-resolution multispectral input data to be fused HR-MS-Input covers at least three different types of land cover, including vegetation, water bodies, buildings, and bare soil. Principle 3: Data Quality: The high-resolution hyperspectral benchmark data HR-HSI-Benchmark, the low-resolution hyperspectral input data to be fused LR-HSI-Input, and the high-resolution multispectral input data to be fused HR-MS-Input imaging areas are free from clouds, fog, shadows, and atmospheric pollution, and the spectral and spatial information is pure. Principle 4: Spatial resolution requirement: The spatial resolution of the high-resolution hyperspectral reference data HR-HSI-Benchmark must be at least twice that of the low-resolution hyperspectral input data LR-HSI-Input to be fused, and the same as that of the high-resolution multispectral input data HR-MS-Input to be fused. Principle 5: Spectral channel requirements: The spectral channel coverage of the high-resolution hyperspectral benchmark data HR-HSI-Benchmark should be the same as or greater than the coverage of the low-resolution hyperspectral input data LR-HSI-Input and the high-resolution multispectral input data HR-MS-Input to be fused.

4. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The spectral matching and downsampling process in step S1 specifically includes: Extract the spectral response functions of HR-HSI-Benchmark and LR-HSI-Input, and screen the effective spectral channels with band overlap ≥90%; After spectral matching, the HR-HSI-Benchmark is first denoised by Gaussian filtering with a standard deviation of σ=1.2, and then downsampled to the same resolution as the LR-HSI-Input by bilinear interpolation to generate the low-resolution hyperspectral data LR-HSI-Train for training. The spectral response function of HR-MS-Input is used to downsample the HR-HSI-Benchmark to generate intermediate data with the same number of bands as HR-MS-Input. The intermediate data is then denoised by Gaussian filtering with a standard deviation of σ=1.2, and then downsampled to the same resolution as HR-MS-Input by bilinear interpolation to generate high-resolution multispectral data HR-MSI-Train for training.

5. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The training sample set in step S1 is constructed as follows: HR-HSI-Benchmark and HR-MSI-Train are cropped using 256×256 pixel blocks to obtain the cropped HR-HSI-Benchmark and HR-MSI-Train; The LR-HSI-Train is cropped by pixel blocks to obtain the cropped LR-HSI-Train; where s represents the spatial resolution difference between HR-MS-Input and LR-HSI-Input; the cropped HR-HSI-Benchmark, HR-MSI-Train and the cropped LR-HSI-Train are used to form a training sample; all training samples are divided into training set and validation set in an 8:2 ratio, where the cropped HR-HSI-Benchmark is used as the fusion ground truth.

6. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The specific structure of the shallow feature extraction module in step S2 is as follows: Spectral feature branch: After upsampling of low-resolution hyperspectral data, it is passed sequentially through two 3×3 convolutional layers, a KAN spectral modeling network, and a KAN channel attention network to output shallow spectral features F_s; the KAN spectral modeling network contains two hidden layers with 128 nodes per layer and uses a third-order piecewise polynomial activation function. Spatial feature branch: High-resolution multispectral data are sequentially passed through multi-scale convolutional units and spatial attention networks to output shallow spatial features F_sp; the multi-scale convolutional units include parallel 1×1, 3×3, and 5×5 convolutional layers.

7. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 5, characterized in that, The specific method for cross-attention fusion is as follows: The query vector Q is obtained using shallow spectral features F_s, and the key vector K and value vector V are obtained using shallow spatial features F_sp. Cross-attention calculation is then performed to obtain the first calculation result. The query vector Q is obtained using shallow spatial features F_sp, and the key vector K and value vector V are obtained using shallow spectral features F_s. Cross-attention calculation is then performed to obtain the second calculation result. The first and second calculation results are concatenated and then input into the convolutional layer to output the cross-attention fusion result F_cross.

8. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The deep fusion module in step S2 includes a 3-layer Swing Transformer module and a 3×3 convolutional layer; wherein, After cascading 3 layers of Swin Transformer modules, they are cascaded with a 3×3 convolutional layer to process the cross-modal fusion feature F_cross and output the deep fusion feature F_deep. The Swin Transformer module has a window size of 7×7 and a shift window step size of 3. The Swin Transformer module replaces the output layer of the traditional Swin Transformer, the Multilayer Perceptron (MLP), with a KAN network. The KAN network has one hidden layer, 256 nodes, and the activation function is a third-order piecewise polynomial.

9. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The fused image reconstruction module in step S2 includes a 3×3 convolutional layer; wherein, the Sigmoid activation function is used to pass the input deep fusion features through the 3×3 convolutional layer to output a high-resolution hyperspectral fused image.

10. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The data augmentation in step S3 specifically involves performing horizontal or vertical flipping, 90°, 180°, or 270° rotation, and adding Gaussian noise (variance, probability 0.2) on the training sample set. The probability of horizontal or vertical flipping is 0.5; the probability of 90°, 180°, or 270° rotation is 0.3; and the probability of adding Gaussian noise is 0.2, with a variance of 0.01-0.

05.

11. The hyperspectral and multispectral data fusion method using a Transformer+KAN network according to claim 1, characterized in that, The combined loss function in step S3 is Where L1 is the mean absolute error loss, SSIM is the structural similarity loss, SAM is the spectral angle matching loss, and a, b, and c are the weights of the three error terms; the training parameters are set as follows: Adam optimizer, weight decay coefficient less than 1e-5, initial learning rate less than 1e-4, decay to less than 0.5 of the current value every 50 rounds, batch size greater than or equal to 16, total number of iterations greater than 200 rounds, and iteration stops if the validation set loss function does not decrease for 10 to 30 consecutive rounds.