RGB reconstruction hyperspectral network and method based on multi-scale heterogeneous codebook auto-encoder

Through the multi-scale heterogeneous codebook autoencoder, the generalization ability and spectral fidelity of the RGB to HSI reconstruction model are improved, and the feature limitations caused by training of a single data set is solved, achieving better reconstruction results.

CN120355568APending Publication Date: 2025-07-22SHENZHEN TECH UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510442895.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing RGB to HSI reconstruction methods are limited by single dataset training and cannot effectively fuse complementary information from heterogeneous datasets, resulting in weak generalization capabilities of the model, and the reconstruction results have limitations in detail retention and spectral fidelity.

Method used

A multi-scale heterogeneous codebook autoencoder is used to build a hybrid codebook by fusing complementary band information of heterogeneous data sets such as HySpecNet-11k, ARAD-1k and HyperGlobal-450K, and combined with multi-task joint optimization strategies to improve the model's generalization ability of complex spectral distribution.

Benefits of technology

It significantly improves the model's generalization ability of complex spectral distributions, enhances spectral fidelity and detail retention capabilities, and solves the feature limitations caused by training in a single data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355568A_ABST
    Figure CN120355568A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of hyperspectral images, and particularly relates to an RGB reconstruction hyperspectral network and method based on a multi-scale heterogeneous codebook autoencoder, and the RGB reconstruction hyperspectral network comprises a multi-scale heterogeneous codebook construction module, a multi-scale feature extraction module, a multi-scale quantization mapping module, a multi-scale reconstruction module and a loss function calculation module. According to the method, complementary wave band information of heterogeneous data sets such as HySpecNet-11k (224 wave band), ARAD-1k (31 wave band) and HyperGlobal-450K (191 wave band) is effectively fused, and the problem of feature limitation caused by single data set training is solved; the hybrid codebook dynamic adaptation mechanism can flexibly call codewords of different scales according to input image features, and the generalization ability of the model to complex spectral distribution is significantly improved; in combination with a multi-task joint optimization strategy (HSI-HSI representation learning and RGB-HSI reconstruction task collaborative optimization), the spectrum fidelity is enhanced while details are reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral image technology, and particularly relates to an RGB reconstruction hyperspectral network and method based on a multi-scale heterogeneous codebook autoencoder. Background Art

[0002] With the wide application of hyperspectral imaging technology in fields such as agricultural monitoring, environmental assessment, and medical diagnosis, how to efficiently reconstruct hyperspectral images from low-cost RGB images has become a research hotspot. Traditional hyperspectral imaging devices are limited by optical systems and mechanical structures, and have problems such as slow imaging speed and high cost, making it difficult to meet the requirements of real-time and large-scale monitoring. Therefore, researchers have proposed various RGB-to-HSI reconstruction methods based on deep learning, but their performance is limited by the following technical bottlenecks: Existing RGB-to-HSI reconstruction methods are usually trained based on a single dataset and cannot effectively fuse the complementary information of heterogeneous datasets. Specifically manifested as: Data island problem: Hyperspectral datasets collected by different sensors (such as satellite remote sensing data, laboratory spectrometer data) are difficult to directly share feature representations due to differences in band ranges, resolutions, and noise characteristics; Low utilization rate of prior knowledge: Existing methods do not fully utilize the potential complementary band information in heterogeneous datasets (such as the correlation between visible light bands and near-infrared bands), resulting in limitations in detail retention and spectral fidelity of the reconstruction results; Weak model generalization ability: Models trained relying on a single dataset have a significant drop in reconstruction accuracy when facing new scenarios (such as new sensor data or complex lighting conditions).

[0003] How to fuse the complementary information of heterogeneous hyperspectral datasets to construct an RGB-to-HSI reconstruction model with strong generalization ability to break through the performance bottleneck of single-dataset training is an urgent problem to be solved currently. Summary of the Invention

[0004] The purpose of the present invention is to provide an RGB reconstruction hyperspectral network and method based on a multi-scale heterogeneous codebook autoencoder, which significantly improves the generalization ability of the model for complex spectral distributions to solve the problems proposed in the above background art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An RGB reconstruction hyperspectral network based on a multi-scale heterogeneous codebook autoencoder, comprising: The multi-scale heterogeneous codebook construction module is used to generate a hybrid codebook; the multi-scale feature extraction module includes convolutional units with gradually decreasing dimensions and is used to perform multi-scale feature decomposition on the input RGB image; the multi-scale quantization mapping module includes quantization units corresponding to the feature decomposition levels, and each quantization unit calls the corresponding scale codebook in the hybrid codebook to perform vector quantization on the features; the multi-scale reconstruction module uses convolutional units with gradually increasing dimensions to perform channel-level splicing on the quantized features and the original features and reconstructs the hyperspectral image; the loss function calculation module optimizes the network parameters by combining the embedding loss, the commitment loss, and the reconstruction loss.

[0006] Preferably, the generation of the hybrid codebook includes: Training multiple independent multi-scale VQ-VAE models using at least two heterogeneous hyperspectral datasets, and the multiple independent models correspond to the band features of different datasets; Extracting the quantization codebooks corresponding to the preset scales in each independent model to form a heterogeneous codebook set; Splicing the codebooks in the heterogeneous codebook set according to a preset rule to construct a hybrid codebook covering a multi-band range.

[0007] Preferably, the formula for calculating the embedding loss is: , where represents the stop gradient operation, is the quantized feature, is the original feature; The formula for calculating the commitment loss is: ; The formula for calculating the reconstruction loss is: , where is the real hyperspectral image, is the reconstruction result, is the setting of the smoothing term constant, which is set to .

[0008] Preferably, in the multi-scale heterogeneous codebook construction module: The heterogeneous dataset contains hyperspectral data collected by different sensors, and the number of bands in each dataset is at least two of 224, 31, and 191; When splicing the codebooks, keep the spectral dimension of each codebook as 512, and the spatial dimension is scaled according to the ratio of the number of bands in the original dataset; The total spectral dimension of the hybrid codebook is the sum of the spectral dimensions of the heterogeneous codebooks.

[0009] Preferably, in the multi-scale quantization mapping module: Set S quantization scales, and the feature reduction factor corresponding to each scale i is ; The quantization process uses the nearest neighbor search method to select the codeword with the smallest Euclidean distance in the codebook for replacement; The quantization losses at each scale are accumulated according to the formula: , where is the feature after dimensionality reduction at the current scale, is the quantization result, and are the embedding loss and the commitment loss respectively.

[0010] Preferably, in the multi-scale reconstruction module: The reverse operation corresponding to the upsampling factor and the quantization scale is adopted in the reconstruction process; The channel dimension consistency is maintained during feature concatenation, and the final output spectral dimension matches the number of bands of the input RGB image; The residual connection mechanism is used to maintain the gradient propagation path.

[0011] On the other hand, the present invention proposes an RGB reconstruction hyperspectral reconstruction method for an RGB reconstruction hyperspectral network, including the following steps: Generate a hybrid codebook by using the multi-scale heterogeneous codebook construction module of the network; input the RGB image to be reconstructed into the multi-scale feature extraction module of the network to obtain multi-scale decomposition features; perform vector quantization on each scale feature by calling the hybrid codebook through the multi-scale quantization mapping module; adopt the multi-scale reconstruction module to concatenate the quantization features and the original features step by step and reconstruct the hyperspectral image; optimize the network parameters by using the loss function calculation module until the reconstruction error converges.

[0012] Preferably, the construction of the hybrid codebook includes: Train the first VQ-VAE model with the HySpecNet-11k dataset at the S = 2 scale to generate a -dimensional codebook; Train the second VQ-VAE model with the ARAD-1k dataset to generate a -dimensional codebook; Train the third VQ-VAE model with the HyperGlobal-450K dataset to generate a -dimensional codebook; Concatenate each codebook in ascending order of the spatial dimension to form a hybrid codebook.

[0013] Preferably, the vector quantization includes: Perform S-dimensionality reduction operations on the RGB image features, and the dimensionality reduction factor is 2 each time; Insert a quantization unit into the channel dimension after each dimensionality reduction, and call the corresponding scale codebook for quantization; After the quantization result is dimensionally upscaled, it is concatenated with the original features to form a multi-scale feature sequence.

[0014] Preferably, the optimization of the loss function includes: Total loss function formula: , where and are the quantization losses of the HSI-HSI representation learning and RGB-HSI mapping processes, respectively.

[0015] Technical effects and advantages of the present invention: The RGB reconstruction hyperspectral network and method based on the multi-scale heterogeneous codebook autoencoder proposed by the present invention have the following advantages compared with the prior art: The present invention effectively integrates complementary band information of heterogeneous datasets such as HySpecNet-11k (224 bands), ARAD-1k (31 bands), and HyperGlobal-450K (191 bands), solving the problem of feature limitations caused by training with a single dataset; its hybrid codebook dynamic adaptation mechanism can flexibly call codewords of different scales according to the input image features, significantly improving the generalization ability of the model to complex spectral distributions; combined with the multi-task joint optimization strategy (cooperative optimization of HSI-HSI representation learning and RGB-HSI reconstruction tasks), it enhances spectral fidelity while retaining details. Brief Description of the Drawings

[0016] Figure 1 is a module diagram of the RGB reconstruction hyperspectral network of the present invention; Figure 2 is a flowchart of the RGB reconstruction hyperspectral reconstruction method of the present invention; Figure 3 is an output flowchart of the reconstructed HSI of the present invention. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment 1 The present invention provides a RGB reconstruction hyperspectral network based on a multi-scale heterogeneous codebook autoencoder as shown in Figure 1 , including a multi-scale heterogeneous codebook construction module, a multi-scale feature extraction module, a multi-scale quantization mapping module, a multi-scale reconstruction module, and a loss function calculation module.

[0019] In this embodiment, the multi-scale heterogeneous codebook construction module is used to generate a hybrid codebook through the following steps: S1: Train multiple independent multi-scale VQ-VAE models using at least two heterogeneous hyperspectral datasets. The multiple independent models correspond to the band characteristics of different datasets. S2: Extract the quantization codebooks corresponding to the preset scales in each independent model to form a heterogeneous codebook set. S3: Stitch the codebooks in the heterogeneous codebook set according to preset rules to construct a hybrid codebook covering a multi-band range.

[0020] Furthermore, the heterogeneous datasets include hyperspectral data collected by different sensors, and the number of bands in each dataset is at least two of 224, 31, and 191. When stitching the codebooks, the spectral dimension of each codebook is kept as 512, and the spatial dimension is scaled according to the ratio of the number of bands in the original dataset. The total spectral dimension of the hybrid codebook is the sum of the spectral dimensions of the heterogeneous codebooks.

[0021] In this embodiment, the multi-scale feature extraction module includes convolutional units for gradually reducing the dimension, which are used to perform multi-scale feature decomposition on the input RGB image. In this embodiment, the multi-scale quantization mapping module includes quantization units corresponding to the feature decomposition levels. Each quantization unit calls the corresponding scale codebook in the hybrid codebook to perform vector quantization on the features. Specifically, set S quantization scales, and the feature reduction factor for each scale i is ; The quantization process uses the nearest neighbor search method to select the codeword with the smallest Euclidean distance in the codebook for replacement. The quantization losses at each scale are accumulated according to the formula: , where is the feature reduced in dimension at the current scale, is the quantization result, and are the embedding loss and the commitment loss respectively.

[0022] In this embodiment, the multi-scale reconstruction module stitches the quantization features and the original features at the channel level through convolutional units for gradually increasing the dimension and reconstructs the hyperspectral image. Specifically, the reconstruction process uses the reverse operation with the upsampling factor corresponding to the quantization scale. When stitching the features, the channel dimension is kept consistent, and the final output spectral dimension matches the number of bands of the input RGB image. The residual connection mechanism is used to maintain the gradient propagation path.

[0023] In this embodiment, the loss function calculation module optimizes the network parameters by combining the embedding loss, the commitment loss, and the reconstruction loss.

[0024] Among them, the formula for calculating the embedding loss is: , where represents the stop gradient operation, is the quantization feature, is the original feature; The formula for the commitment loss is as follows: ; The formula for the reconstruction loss is as follows: , where is the true hyperspectral image, is the reconstruction result, is the setting of the smoothing term constant, which is set to .

[0025] Example 2 In this example, an RGB reconstruction hyperspectral reconstruction method is proposed. As Figure 2 shown, it includes the following steps: Using the multi-scale heterogeneous codebook construction module of the network to generate a hybrid codebook. Specifically, the construction of the hybrid codebook includes: Training the first VQ-VAE model with the HySpecNet-11k dataset at the S = 2 scale to generate a -dimensional codebook; Training the second VQ-VAE model with the ARAD-1k dataset to generate a -dimensional codebook; Training the third VQ-VAE model with the HyperGlobal-450K dataset to generate a -dimensional codebook; Concatenating each codebook in ascending order of the spatial dimension to form a hybrid codebook.

[0026] Inputting the RGB image to be reconstructed into the multi-scale feature extraction module of the network to obtain multi-scale decomposition features; Through the multi-scale quantization mapping module, calling the hybrid codebook to perform vector quantization on each scale feature. Specifically, it includes: performing S times of dimensionality reduction operations on the RGB image features, with a dimensionality reduction factor of 2 each time; inserting quantization units in the channel dimension after each dimensionality reduction, and calling the corresponding scale codebook for quantization; after the quantization result is upsampled, it is concatenated with the original feature to form a multi-scale feature sequence.

[0027] Adopting the multi-scale reconstruction module to concatenate the quantized features and the original features step by step and reconstruct the hyperspectral image; Using the loss function calculation module to optimize the network parameters until the reconstruction error converges. The optimization of the loss function includes: the formula for the total loss function: , where and are the quantization losses of the HSI-HSI representation learning and the RGB-HSI mapping processes, respectively.

[0028] Example 3 In this embodiment, a multi-scale VQ-VAE framework for extracting the latent representation of heterogeneous hyperspectral datasets is proposed based on the VQ-VAE (Vector Quantized-Variational Autoencoder) technology. This framework sequentially includes a multi-scale quantization and a multi-scale reconstruction process. The overall process includes two processes: the characterization learning of HSI-HSI and the mapping of RGB-HSI.

[0029] Overall process of HSI-HSI characterization learning: Input the hyperspectral HSI. Since the number of hyperspectral bands may be very large, MaskedEncoder (any convolutional structure) performs dimensionality reduction (masking) in the channel dimension. Further perform the multi-scale quantization process to obtain the intermediate product, the multi-scale features . For Further perform the multi-scale reconstruction process to obtain .

[0030] Finally, for Perform any convolutional structure mapping to restore its channels to the original stage, completing the latent characterization learning of HSI. The product of this stage: the trained (codebook) and the multi-scale reconstruction network.

[0031] The multi-scale quantization process starts from the input hyperspectral image , where , and represent the number of channels (number of bands), length, and width of the hyperspectral image respectively. During model initialization, is randomly initialized from a normal distribution as a model parameter; at the start of training, a channel mask is applied to . Here, is used for iterative dimensionality reduction to obtain .

[0032] In this process, is affected by , and their respective codebooks are , represents the convolutional mapping.

[0033] The pseudo-code of multi-scale quantization is as follows: 1:Input:Hyperspectral image ; 2:Output:Featureset H=[], ; 3: ; 4:for to do; 5: ; 6: ; 7: ; 8: ; 9: ; 10: ; 11: end for; 12: return .

[0034] Quantization: Convert to , indicating there are spectral vectors with dimension . Formulas (1) and (2) respectively represent calculating the distances between each spectral vector and all vectors in , and the index of the nearest vector in . Use this index to replace each vector in with the vector in , thus obtaining : , (2).

[0035] Assume is performed, that is, quantization at the 2-scale. For the input (1:), first perform random channel masking or one-dimensional convolution to reduce the dimension to (3:). Loop from 1 to 2 (inclusive), perform one-dimensional convolution on to obtain (5:), and add it to the queue. Then use to quantize to obtain the latent variable (7:); calculate the loss once here and accumulate it into (8:). Then perform one-dimensional upsampling convolution on to restore the spatial resolution to obtain (9:). Finally, perform one-dimensional convolution mapping on and assign it to , while performs the next loop (10:).

[0036] In the second loop, perform one-dimensional convolution on to obtain (5:), and add it to In the queue. Then use to perform quantization to obtain latent variables (7:); Here, calculate the loss once and accumulate it to in (8:). Then perform an upsampling convolution on to restore the spatial resolution to obtain perform a convolution mapping on and assign it to The loop stops. After the multi-scale quantization loop ends, the multi-scale feature queue and the accumulated loss

[0037] The multi-scale reconstruction pseudocode describes the multi-scale reconstruction process, which uses the extracted feature set to restore the hyperspectral image. This process iterates from the smallest scale to the largest scale, extracting the dimensionality-reduced features from and quantizing them to obtain . Concatenate the quantized features with , and then perform spatial upsampling. This ensures the minimum loss of the original information. This process iteratively concatenates and in the channel dimension. After the iteration is completed, apply a convolution to to obtain .

[0038] The multi-scale reconstruction pseudocode is as follows: 1: Input: Featureset 2: Output: Featureset , 3: for to do 4: 5: 6: 7: 8: 9: end for 10: 10: return .

[0039] Assume that , i.e., quantization at 2 scales. The input is the output result of multi-scale quantization, and the multi-scale features . At this time, perform a reverse loop from 2 to 1 (inclusive), and pop the feature from (4:). Then use to for quantization to obtain the latent variable (5:); calculate the loss once here and accumulate it into (6:). Then, concatenate and in the channel dimension and perform a spatial upsampling once to obtain (7:). Perform a convolutional mapping on and concatenate it with in the channel dimension and assign the result to (8:).

[0040] In the second loop, pop the feature from (4:). Then use to for quantization to obtain the latent variable (5:); calculate the loss once here and accumulate it into (6:). Then, concatenate and in the channel dimension and perform a spatial upsampling once to obtain (7:). Perform a convolutional mapping on and concatenate it with in the channel dimension and assign the result to (8:).

[0041] The loop stops. Perform a convolutional mapping on to restore the number of channels to the same as the input, .

[0042] Update the total loss function of the model as shown in the formula: ; The reconstruction loss is used to update the parameters of the convolutional operator, which is set to : .

[0043] Use the embedding loss and the commitment loss, which is consistent with the typical VQ-VAE, to update the parameters of the codebook. The embedding loss is used to update the codebook to make it closer to the distribution of . The commitment loss For regularizing parameter updates to stabilize during training and prevent its unrestricted growth while ensuring the stability of the codebook distribution.

[0044] ; ; where represents the stop gradient operation.

[0045] After independently training datasets and extracting their respective codebooks, they are concatenated to obtain a multi-scale heterogeneous codebook and a trained multi-scale VQ-VAE.

[0046] Finally, the RGB-to-HSI reconstruction is completed. In any RGB-HSI reconstruction network, the multi-scale heterogeneous codebook is regularly used to quantize the layer features and convert them into the latent representation of the real HSI as much as possible.

[0047] As Figure 3 shown, at the end of any network, the multi-scale reconstruction network of the multi-scale VQ-VAE is connected, and finally the reconstructed HSI is output. This process requires freezing the parameters of the multi-scale heterogeneous codebook and the multi-scale reconstruction network.

[0048] The mapping process of RGB-HSI, taking ResNet as an example, is assumed to be composed of 2 residual blocks (the quantity here is consistent with the previous multi-scale S): 1. Input the RGB image and enter the first residual block of the ResNet network. Use H to retain the features of this scale, and at the same time use to quantize this scale, and the quantization result is passed to the next residual block.

[0049] 2. Use H to retain the output of the second residual block.

[0050] 3. Send H into the multi-scale reconstruction network to obtain the HSI output and complete the RGB-HSI mapping.

[0051] 4. Taking Unet as an example (U-shaped network), it is assumed to be composed of 2 downsampling blocks (Encoder) and 2 upsampling blocks (Decoder).

[0052] 5. Input the RGB image. After passing through the 1st and 2nd downsampling blocks, it enters the upsampling block stage.

[0053] 6. After the first upsampling block is calculated, use H to retain the features of this scale. At the same time use to quantize this scale, and the quantization result is passed to the next upsampling block.

[0054] 7. Use H to retain the output of the second upsampling block.

[0055] 8. Feed H into the multi-scale reconstruction network to obtain the HSI output, completing the RGB-HSI mapping.

[0056] 9. Taking Transformer as an example, assuming 3 Transformer blocks are stacked, the corresponding multi-scale VQ-VAE in the previous text also needs to be trained according to the scale of S = 3.

[0057] 10. Input the RGB image. After being calculated by the first Transformer block, use H to retain this feature, and at the same time use , perform quantization on this scale, and transfer the quantization result to the next Transformer block.

[0058] 11. After being calculated by the second Transformer block, use H to retain this feature, and at the same time use , perform quantization on this scale, and transfer the quantization result to the next Transformer block.

[0059] 12. After being calculated by the third Transformer block, use H to retain this output.

[0060] 13. Feed H into the multi-scale reconstruction network to obtain the HSI output, completing the RGB-HSI mapping.

[0061] Use H to retain the multi-scale features, and the shape of this feature needs to meet the shape requirements of the multi-scale reconstruction network. It is required to regularly use to perform quantization on the layer features and transfer them into the latent space of HSI. Finally, feed H into the trained multi-scale reconstruction network to obtain HSI.

[0062] The RGB-reconstructed hyperspectral network based on the multi-scale heterogeneous codebook autoencoder proposed in the present invention can make full use of the existing heterogeneous hyperspectral datasets, because the heterogeneous hyperspectral datasets have similar or complementary information in the latent space, and this rich information is preserved in the multi-scale heterogeneous codebook: Taking three currently public datasets as examples: HySpecNet-11k, which contains 11k RGB-HSI pairs, The resolution is 224 bands in the spectral range of 420 - 2450 nm, collected from a satellite. The HySpecNet authors divided it into hard mode and easy mode according to whether the training set, test set, and validation set contain patches from the same image. All hard mode data was used in this experiment, which means there is no data leakage problem. The HySpec-train-hard dataset contains approximately 8k samples, the HySpec-val-hard dataset contains 2k samples, and the HySpec-test hard dataset contains 1k samples.

[0063] ARAD-1k, containing 1k RGB-HSI pairs, The resolution is 31 bands in the range of 400 - 700 nm, with an interval of 10 nm. When training, it is cropped into resolution. ARAD-train contains 900 independent data, ARAD-valid contains 50 independent data, and the training set will be divided into patch.

[0064] HyperGlobal-450K, containing 1,701 RGB-HSI pairs, with a resolution of 64x64 and 191 bands. This dataset is used to participate in the construction of the Mixture of Codebooks. The three heterogeneous datasets are reflected in the use of different sensors to record hyperspectral information with different numbers of bands, and there is a certain complementary relationship. The multi-scale VQVAE model proposed in this patent is used to train on these three datasets respectively to obtain three heterogeneous codebook representations and three trained multi-scale reconstruction networks. We splice the three codebooks together to obtain a hybrid codebook representation.

[0065] Next, assume that the current task is to reconstruct hyperspectral bands consistent with the ARAD-1K dataset. Then, using the hybrid codebook representation and the multi-scale reconstruction network trained on the ARAD-1K dataset combined with any RGB-HSI reconstruction network can better complete the task of reconstructing for a certain band range. The hybrid codebook introduces richer and more extensive prior knowledge, making the RGB-HSI algorithm more generalizable.

[0066] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An RGB-reconstructed hyperspectral network based on a multi-scale heterogeneous codebook autoencoder, characterized in that Including: A multi-scale heterogeneous codebook construction module for generating a hybrid codebook; A multi-scale feature extraction module, including convolutional units for progressive dimensionality reduction, for multi-scale feature decomposition of the input RGB image; A multi-scale quantization mapping module, including quantization units corresponding to the feature decomposition levels, and each quantization unit calls the corresponding scale codebook in the hybrid codebook to perform vector quantization on the features; A multi-scale reconstruction module, which splices the quantized features and the original features at the channel level through convolutional units for progressive dimensionality increase and reconstructs the hyperspectral image; A loss function calculation module, which optimizes the network parameters by combining the embedding loss, the commitment loss, and the reconstruction loss.

2. The RGB reconstruction hyperspectral network based on the multi-scale heterogeneous codebook autoencoder according to claim 1, wherein The generation of the hybrid codebook includes: Training multiple independent multi-scale VQ-VAE models using at least two heterogeneous hyperspectral datasets, and the multiple independent models correspond to the band features of different datasets; Extracting the quantization codebooks corresponding to the preset scales in each independent model to form a heterogeneous codebook set; Splicing the codebooks in the heterogeneous codebook set according to a preset rule to construct a hybrid codebook covering a multi-band range.

3. The RGB reconstruction hyperspectral network based on the multi-scale heterogeneous codebook autoencoder according to claim 2, characterized in that, The embedding loss calculation formula is as follows: , where represents the stop gradient operation, is the quantized feature, is the original feature; The committed loss calculation formula is: ; The reconstruction loss calculation formula is as follows: , where is the real hyperspectral image, is the reconstruction result, is the setting of the smoothing term constant, which is set to .

4. The RGB reconstruction hyperspectral network based on the multi-scale heterogeneous codebook autoencoder according to claim 3, wherein In the multi-scale heterogeneous codebook construction module: The heterogeneous datasets contain hyperspectral data collected by different sensors, and the number of bands in each dataset is at least two of 224, 31, and 191; When splicing the codebooks, the spectral dimension of each codebook is kept as 512, and the spatial dimension is scaled according to the ratio of the number of bands in the original dataset; The total spectral dimension of the hybrid codebook is the sum of the spectral dimensions of the heterogeneous codebooks.

5. The RGB reconstruction hyperspectral network based on the multi-scale heterogeneous codebook autoencoder according to claim 4, characterized in that In the multi-scale quantization mapping module: Set S quantization scales, and for each scale i, the feature dimensionality reduction factor is ; The quantization process uses the nearest neighbor search method to select the codeword with the smallest Euclidean distance in the codebook for replacement; The quantization losses at each scale are accumulated according to the formula: , Among them, is the current scale dimension-reduced feature, is the quantization result, and are the embedding loss and the commitment loss respectively.

6. The RGB reconstruction hyperspectral network based on the multi-scale heterogeneous codebook autoencoder according to claim 5, wherein In the multi-scale reconstruction module: The reconstruction process uses the reverse operation with the upsampling factor corresponding to the quantization scale; When splicing the features, the channel dimension is kept consistent, and the final output spectral dimension matches the number of bands of the input RGB image; The residual connection mechanism is used to maintain the gradient propagation path.

7. An RGB hyperspectral reconstruction method for an RGB reconstruction hyperspectral network according to any one of claims 1-6, characterized in that, Including the following steps: Using the multi-scale heterogeneous codebook construction module of the network to generate a hybrid codebook; Inputting the RGB image to be reconstructed into the multi-scale feature extraction module of the network to obtain multi-scale decomposition features; Through the multi-scale quantization mapping module, calling the hybrid codebook to perform vector quantization on each scale feature; Using the multi-scale reconstruction module to splice the quantized features and the original features step by step and reconstruct the hyperspectral image; Using the loss function calculation module to optimize the network parameters until the reconstruction error converges.

8. The RGB reconstruction hyperspectral reconstruction method according to claim 7, characterized in that The construction of the hybrid codebook includes: Train the first VQ-VAE model with S = 2 scale using the HySpecNet-11k dataset to generate dimensional codebooks; Train the second VQ-VAE model using the ARAD-1k dataset to generate dimensional codebook; Train the third VQ-VAE model using the HyperGlobal-450K dataset to generate a dimensional codebook; Splicing each codebook into a hybrid codebook in ascending order of the spatial dimension.

9. The RGB reconstruction hyperspectral reconstruction method according to claim 7, wherein The vector quantization includes: Performing S times of dimensionality reduction operations on the RGB image features, and the dimensionality reduction factor for each time is 2; Inserting a quantization unit into the channel dimension after each dimensionality reduction and calling the corresponding scale codebook for quantization; After the quantization result is dimensionally increased, it is spliced with the original feature to form a multi-scale feature sequence.

10. The RGB reconstruction hyperspectral reconstruction method according to claim 7, characterized in that, The optimization of the loss function includes: Total loss function formula: , where and are the quantization losses of the HSI-HSI representation learning and RGB-HSI mapping processes, respectively.

Citation Information

Cited By

  • Unmanned aerial vehicle target detection method and system based on infrared modal privilege information

    CN121353652A

  • Unmanned aerial vehicle target detection method and system based on infrared modal privileged information

    CN121353652B

  • Three-dimensional physical field data compression method based on multi-scale vector quantization

    CN122247430A