Hyperspectral calculation imaging method and system based on multi-domain fusion of spatial, spectral and frequency domains, and medium

By converting the RGB image to the frequency domain, extracting and fusing the frequency information, the problems of poor detailed information and low reconstruction accuracy of high-spectral images in the prior art are solved, and efficient spectral reconstruction effect is achieved.

WO2025112554A1PCT designated stage expired Publication Date: 2025-06-05HUNAN UNIV

Patent Information

Application Number
PCT/CN2024/105425
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-07-15
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing hyperspectral calculation and imaging methods fail to effectively utilize frequency domain information, resulting in poor detailed information of hyperspectral images and low reconstruction accuracy.

Method used

A hyperspectral image is generated by converting the RGB image to the frequency domain, extracting the frequency information and fusing it into the null spectral domain features. The specific steps include converting the RGB image to the frequency domain, extracting the frequency information, transforming it into the spatial domain, and fusing it into the null spectral domain features.

Benefits of technology

It significantly improves the spectral reconstruction accuracy and visual effects, and can more effectively reconstruct the fine details of hyperspectral images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024105425_05062025_PF_FP_ABST
    Figure CN2024105425_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a hyperspectral calculation imaging method and system based on multi-domain fusion of spatial, spectral and frequency domains, and a medium. The method of the present invention comprises: using a two-dimensional offline discrete cosine transform (DCT) to convert an RGB image Y into a frequency domain to obtain a frequency domain feature map Yfreq; extracting a frequency information image (I) from the frequency domain feature map Yfreq; using a two-dimensional offline inverse discrete cosine transform (IDCT) to transform the frequency information image (I) into a spatial domain to obtain a frequency information image Xfreq of the spatial domain; and fusing the frequency information image Xfreq of the spatial domain into the spatial-spectral domain features of the RGB image Y to generate a hyperspectral image (II). The present invention aims to solve the problems of poor detail information and low reconstruction accuracy of hyperspectral images in existing hyperspectral calculation imaging, and realizes high-fidelity reconstruction of a target spectrum.
Need to check novelty before this filing date? Find Prior Art

Description

A hyperspectral computational imaging method, system and medium for space-spectrum-frequency multi-domain fusion

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on and claims priority to a Chinese patent application with an application date of "November 30, 2023", application number "202311622750.8", and invention name "A hyperspectral computational imaging method, system and medium with spatial-spectral-frequency multi-domain fusion". The full text of the Chinese patent application is hereby cited in this application as a part of this application.

Technical field

[0003] The present invention relates to the technical field of hyperspectral imaging, and in particular to a hyperspectral computational imaging method, system and medium for space-spectrum-frequency multi-domain fusion. [Background Technology]

[0004] Hyperspectral images typically contain dozens to hundreds of continuous spectral bands. This high number of spectral bands records rich spectral information about objects. Because different objects have varying reflectivity and absorptivity of electromagnetic waves at different wavelengths, accurate object identification is possible. Currently, a variety of spectral imagers are commercially available, but their imaging speeds are slow and cannot meet real-time requirements. Furthermore, hyperspectral imaging equipment is bulky, difficult to carry around, and relatively expensive, limiting its application and development. Acquiring high-spatial-resolution hyperspectral images at a low cost has become a research hotspot in the field of hyperspectral imaging.

[0005] Hyperspectral computational imaging technology aims to design an algorithm to reconstruct high-spatial-resolution hyperspectral images from three-band RGB images. It has the advantages of low cost, fast imaging speed, and high spatial resolution. Compared with hyperspectral images, RGB images lose a lot of spectral information. Therefore, recovering spectral information from RGB images to reconstruct hyperspectral images is an inverse problem and is severely ill-posed. Existing research methods can be roughly divided into three categories: hardware-based methods, prior-based methods, and deep learning-based methods. Hardware-based spectral imaging methods mainly design a specific system to increase the spectral dimension information obtained from RGB cameras, mainly including modifying the RGB camera system, increasing the number of cameras, or controlling the imaging environment. Prior-based methods use mathematical relationships to model the mapping relationship between RGB and hyperspectral images, and learn specific priors based on the inherent properties and statistical information of the hyperspectral image to reconstruct the spectral information. The spectral imaging accuracy of the above-mentioned methods is significantly affected by manual priors. In recent years, due to the outstanding performance of convolutional neural networks in the field of computer vision, spectral imaging technology based on deep learning methods has gradually become a research hotspot. This method generally uses a large number of available RGB and hyperspectral image pairs to characterize the hidden mapping relationship between the two, achieving accurate reconstruction of high-spatial-resolution hyperspectral images, and is therefore also known as a data-driven method. In addition, some scholars have studied spectral super-resolution technology that combines physical models with deep learning. This transforms the super-resolution problem into a target optimization problem, applies optimization theory to solve the problem, and makes the model physically interpretable.

[0006] In summary, existing hyperspectral computational imaging methods focus solely on information in the spatial or spectral domain, failing to utilize information in the frequency domain. Furthermore, the complex spatial-spectral structure of hyperspectral images makes it difficult to accurately represent single-dimensional information. Therefore, leveraging information in the frequency domain to achieve hyperspectral computational imaging has become a critical technical challenge that needs to be addressed.

[0007] [Summary of the invention]

[0008] The technical problem to be solved by the present invention is as follows: In view of the above-mentioned problems of the prior art, a hyperspectral computational imaging method, system and medium with spatial-spectral-frequency multi-domain fusion are provided. The present invention aims to solve the problems of poor hyperspectral image detail information and low reconstruction accuracy in existing hyperspectral computational imaging, and to achieve high-fidelity reconstruction of the target spectrum.

[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0010] A hyperspectral computational imaging method for space-spectrum-frequency multi-domain fusion, comprising:

[0011] S1, convert the RGB image Y to the frequency domain to obtain the frequency domain feature map Y freq ;

[0012] S2, from the frequency domain feature map Y freq Extract frequency information from the graph

[0013] S3, the frequency information graph Transform to the spatial domain to obtain the frequency information graph Y in the spatial domain freq ;

[0014] S4, the frequency information map of the spatial domain X freq Fusion into the spatial spectrum domain features of RGB image Y to generate a hyperspectral image

[0015] Optionally, in step S1, the RGB image Y is converted to the frequency domain to obtain the frequency domain feature map Y freq include:

[0016] S1.1, convert the input RGB image Y into YCbCr space image and split it into image blocks according to color channels;

[0017] S1.2, use two-dimensional offline discrete cosine transform (DCT) to project each image block into the frequency domain to obtain frequency coefficients;

[0018] S1.3, the frequency coefficients of each image block are stacked together in the order of Y, Cb, and Cr channels, and vectorized to obtain frequency bands; the frequency bands are then reorganized in a Z-shaped manner to obtain the frequency domain feature map Y freq , and the frequency domain feature map Y freq Each channel corresponds to a frequency band.

[0019] Optionally, in step S2, from the frequency domain feature map Y freq Extract frequency information from the graph include:

[0020] S2.1, the frequency domain feature map Y freq According to the Y, Cb, and Cr channels of the YCbCr space, the channel dimension is divided into three equal parts. First, the features of each equal part are spatially downsampled and divided into low-frequency information and high-frequency information. Then, the low-frequency information and high-frequency information in each equal part are channel-recombined to obtain the low-frequency features. With high frequency features And the size is Where C is the frequency domain feature map Y freq The number of channels, k represents the size after spatial downsampling;

[0021] S2.2, respectively for low-frequency features With high frequency features Use m dense residual blocks RDB for mapping learning, and use GeLu function as the activation function in the dense residual block RDB;

[0022] S2.3, the low-frequency features after mapping learning and the high-frequency features are stacked in the channel dimension and then mapped using n dense residual blocks RDB to obtain a frequency reorganization map, and then the frequency reorganization map is upsampled to obtain a frequency information map.

[0023] Optionally, in step S3, the frequency information map Transform to the spatial domain to obtain the frequency information graph X in the spatial domain freq include:

[0024] S3.1, dividing the frequency information map into image blocks;

[0025] S3.2, project each image block into the spatial domain using a two-dimensional offline inverse discrete cosine transform (IDCT);

[0026] S3.3, combine the image blocks projected into the spatial domain to obtain the frequency information map Y in the spatial domain freq .

[0027] Optionally, the spatial spectrum domain feature of the RGB image Y in step S4 includes extracting the spatial spectrum shallow feature Y of the RGB image Y using a convolution kernel size of 1×1. sfe , and the shallow feature Y sfe The convolution kernel size is 3×3 and then fed into the symmetric convolutional neural network to further extract the spatial spectrum deep features X of the RGB image Y. sase ; And the frequency information diagram of the spatial domain is X freq Fusion into the spatial spectrum domain features of RGB image Y to generate a hyperspectral image The function expression is:

[0028] In the above formula, F conv3×3 Indicates the convolution operation with a convolution kernel size of 3×3, F conv1×1 Indicates a convolution operation with a convolution kernel size of 1×1, X freq It is the frequency information diagram in the spatial domain.

[0029] Optionally, the symmetric convolutional neural network includes a feature extraction upper branch, a spatial attention module SA, a feature extraction lower branch, a channel stacking module CAT and a convolution module, the feature extraction upper branch includes N local feature extraction modules LFEM connected in sequence, the feature map of the input symmetric convolutional neural network enters from the first local feature extraction module LFEM of the feature extraction upper branch, and the local feature extraction modules LFEM in the feature extraction upper branch each include three outputs, the first output to the channel stacking module CAT, the second output to a corresponding spatial attention module SA, the third output of the first N-1 local feature extraction modules LFEM to the next local feature extraction module LFEM in the feature extraction upper branch, the feature extraction lower branch includes N The output of the first spatial attention module SA is fused with the output of the Nth local feature extraction module LFEM in the feature extraction upper branch and serves as the input of the first local feature extraction module LFEM in the feature extraction lower branch. The outputs of the remaining spatial attention modules SA are fused with the output of the previous local feature extraction module LFEM in the feature extraction lower branch and serve as the input of the corresponding local feature extraction module LFEM in the feature extraction lower branch. Each local feature extraction module LFEM in the feature extraction lower branch includes an output connected to the channel stacking module CAT. The output of the channel stacking module CAT is refined by a convolution module with a convolution kernel size of 1×1 to obtain the spatial spectrum deep feature X of the RGB image Y. sase .

[0030] Optionally, the local feature extraction module LFEM includes three multi-feature fusion dual attention modules MFFDAB, two group convolutions with a convolution kernel size of 1×1, a channel stacking module and a convolution with a convolution kernel size of 1×1. The function expression of the local feature extraction module LFEM for extracting features from the input feature map is:

[0031] In the above formula, They are the output features of three multi-feature fusion dual attention modules MFFDAB, Represent three multi-feature fusion dual attention modules MFFDAB, and Respectively represent the input and output features of the local feature extraction module LFEM, F CAT Indicates the channel stacking module, F conv1×1 represents a convolution with a kernel size of 1×1, and Respectively represent two group convolutions with a convolution kernel size of 1×1; the multi-feature fusion dual attention module MFFDAB is used to extract deep features of the image in the spatial spectral domain, and the function expression for extracting deep features of the image in the spatial spectral domain is: Y CA =F CA (F CAT (Y1,Y2,Y3)), T SA =F SA (F CAT (Y1,Y2,Y3)),

[0032] In the above formula, Y1~Y3 are three groups of intermediate features, F conv1×1 Indicates a convolution with a kernel size of 1×1, F conv3×3 represents a convolution with a kernel size of 3×3, F conv5×5 represents a convolution with a kernel size of 5×5, F conv7×7 represents a convolution with a kernel size of 7×7, F CA represents the channel attention module, F SA represents the spatial attention module, Y CA is the output feature of the channel attention module, Y SA is the output feature of the spatial attention module, and They represent the input and output features of the j-th multi-feature fusion dual attention module MFFDAB respectively.

[0033] Optionally, steps S1 to S2 are implemented based on a hyperspectral computing imaging model, and the hyperspectral computing imaging model includes:

[0034] A frequency domain transformation module, configured to execute step S1;

[0035] A frequency domain learning module FDLM, configured to execute step S2;

[0036] A frequency domain inverse transform module, configured to execute step S3;

[0037] Feature fusion module FFM, used to execute step S4;

[0038] The spatial spectrum domain feature extraction module is used to extract the spatial spectrum domain features of the RGB image Y;

[0039] The loss function L used in the training of the hyperspectral computational imaging model is all The function expression of L is: all =L sase +γL freq ,

[0040] In the above formula, L sase is the loss function in the empty spectrum domain, γ is the balance parameter, L freq It is a frequency domain loss function. The spatial spectrum domain loss function is the L1 norm between the hyperspectral image generated by the hyperspectral computing imaging model and the real hyperspectral image as a label. The frequency domain loss function is the L1 norm between the frequency information map of the spatial domain generated by the hyperspectral computing imaging model and the frequency information map of the spatial domain in the real hyperspectral image.

[0041] The present invention also provides a hyperspectral computational imaging system utilizing spatial, spectral, and frequency multi-domain fusion, comprising an interconnected microprocessor and memory, wherein the microprocessor is programmed or configured to execute the hyperspectral computational imaging method utilizing spatial, spectral, and frequency multi-domain fusion. The present invention also provides a computer-readable storage medium storing a computer program configured to be programmed or configured by the microprocessor to execute the hyperspectral computational imaging method utilizing spatial, spectral, and frequency multi-domain fusion.

[0042] Compared with the prior art, the present invention has the following advantages: the present invention includes converting the RGB image Y into the frequency domain to obtain the frequency domain feature map Y freq , from the frequency domain feature map Y freq Extract frequency information from the graph Frequency Infographic Transform to the spatial domain to obtain the frequency information graph X in the spatial domain freq , the frequency information graph X in the spatial domain freq Fusion into the spatial spectrum domain features of RGB image Y to generate a hyperspectral image By introducing frequency information into spectral super-resolution, a spectral super-resolution method that fuses spatial, spectral and frequency multi-domain features is constructed, which effectively reconstructs the fine details of hyperspectral images. Compared with existing methods, it can significantly improve the spectral reconstruction accuracy and visual effects.

Brief Description of the Drawings

[0043] FIG1 is a schematic diagram of a basic flow chart of a method according to an embodiment of the present invention.

[0044] FIG2 is a schematic diagram of a network structure of a method according to an embodiment of the present invention.

[0045] FIG3 is a schematic diagram showing the principle of step S1 in an embodiment of the present invention.

[0046] FIG4 is a schematic diagram showing the principle of step S2 in an embodiment of the present invention.

[0047] FIG5 is a schematic diagram of the network structure of the local feature extraction module LFEM according to an embodiment of the present invention.

[0048] FIG6 is a schematic diagram of the network structure of the multi-feature fusion dual attention module MFFDAB in an embodiment of the present invention.

[0049] FIG7 is a schematic diagram of the network structure of the spatial attention module in an embodiment of the present invention.

[0050] FIG8 is a schematic diagram of the network structure of the channel attention module in an embodiment of the present invention.

[0051] FIG9 is a comparison experiment result of the imaging results calculated by the method according to the embodiment of the present invention and the existing method. [Specific implementation method]

[0052] As shown in FIG1 and FIG2 , the hyperspectral computational imaging method of the present embodiment using spatial-spectral-frequency multi-domain fusion includes:

[0053] S1, convert the RGB image Y to the frequency domain to obtain the frequency domain feature map Y freq ;

[0054] S2, from the frequency domain feature map Y freq Extract frequency information from the graph

[0055] S3, the frequency information graph Transform to the spatial domain to obtain the frequency information graph X in the spatial domain freq ;

[0056] S4, the frequency information map of the spatial domain X freq Fusion into the spatial spectrum domain features of RGB image Y to generate a hyperspectral image

[0057] In step S1, the RGB image Y is converted to the frequency domain to obtain the frequency domain feature map Y freq It can be expressed as: X freq =F freq (Y),

[0058] In the above formula, F freq (.) is the frequency domain information extraction operation, X freq ∈R W×H×3 is the final output feature of this branch, W×H are the width and height of RGB image Y respectively.

[0059] As shown in FIG3 , in step S1 of this embodiment, the RGB image Y is converted to the frequency domain to obtain the frequency domain feature map Y freq include:

[0060] S1.1, converting the input RGB image Y into a YCbCr space image and segmenting the image into image blocks according to the color channel; the number of rows and columns for segmenting the image blocks in step S1.1 can be selected as needed, and generally the number of rows and columns is the same. For example, as an optional implementation, in this embodiment, the image blocks are segmented into 4×4=16 blocks;

[0061] S1.2, using two-dimensional offline discrete cosine transform (DCT) to project each image block into the frequency domain to obtain frequency coefficients, each of which corresponds to the intensity of a specific frequency band;

[0062] S1.3, the frequency coefficients of each image block are stacked together in the order of Y, Cb, and Cr channels, and vectorized to obtain frequency bands; the frequency bands are then reorganized in a Z-shaped manner to obtain the frequency domain feature map Y freq , and the frequency domain feature map Y freq Each channel in the frequency band corresponds to a frequency band. The frequency bands are reorganized in a zigzag pattern, that is, first reorganizing the first row from the beginning to the end, then returning to the beginning of the second row, and reorganizing the second row from the beginning to the end, and so on, finally returning to the beginning of the last row, and reorganizing the last row from the beginning to the end.

[0063] Since the frequency domain information Y obtained in the previous step freq It cannot well represent all the features of the image, and it is necessary to further learn the frequency domain information with the help of the nonlinear representation ability of deep learning, so as to better supplement the spatial spectrum domain information. As shown in Figure 4, in step S2 of this embodiment, the frequency domain feature map Y freq Extract frequency information from the graph include:

[0064] S2.1, the frequency domain feature map Y freq According to the Y, Cb, and Cr channels of the YCbCr space, the channel dimension is divided into three equal parts. First, the features of each equal part are spatially downsampled and divided into low-frequency information and high-frequency information. Then, the low-frequency information and high-frequency information in each equal part are channel-recombined to obtain the low-frequency features. With high frequency features And the size is Where C is the frequency domain feature map Y freq The number of channels, k represents the size after spatial downsampling; in this embodiment, C=48, k is W / 8, and W is the width of the RGB image Y.

[0065] S2.2, respectively for low-frequency features With high frequency features m dense residual blocks (RDBs) are used for mapping learning, and the GeLu function is used as the activation function in the dense residual block RDB. Based on the principle of not increasing the complexity of the model, multiple dense residual blocks (RDBs) are selected to respectively map high-frequency information and low-frequency information, ensuring that the richness of information is captured without adding too many parameters. Because the frequency domain signal contains the possibility of negative numbers, the GeLu function is used as the activation function in the dense residual block.

[0066] S2.3, the low-frequency features after mapping learning and the high-frequency features are stacked in the channel dimension and then mapped using n dense residual blocks RDB to obtain a frequency reorganization map, and then the frequency reorganization map is upsampled to obtain a frequency information map. Based on the principle of not increasing the complexity of the model, multiple dense residual blocks (RDBs) are selected to map high-frequency information to low-frequency information respectively, ensuring that the richness of information is captured without adding too many parameters.

[0067] Since the learned frequency domain information is mainly used for feature supplementation, it is necessary to convert the information into the spatial domain and use 2DIDCT to convert it into the spatial domain. Similar to the offline DCT transformation, it is also necessary to cut the feature map one by one and then inversely transform it. Finally, the final output feature map X with the same size as the input RGB is obtained based on the shape rearrangement. freq In this embodiment, in step S3, the frequency information map Transform to the spatial domain to obtain the frequency information graph X in the spatial domain freq include:

[0068] S3.1, dividing the frequency information map into image blocks;

[0069] S3.2, project each image block into the spatial domain using a two-dimensional offline inverse discrete cosine transform (IDCT);

[0070] S3.3, combine the image blocks projected into the spatial domain to obtain the frequency information map X in the spatial domain freq .

[0071] It should be noted that the spatial spectral domain features of the RGB image Y are realized by the spatial spectral domain information extraction branch, which can be specifically realized by using a known spatial spectral domain feature extraction method as needed. As an optional embodiment, as shown in FIG2 , the spatial spectral domain features of the RGB image Y in step S4 of this embodiment include extracting the spatial spectral shallow feature Y of the RGB image Y using a convolution kernel size of 1×1. freq (This feature has the same number of channels as the hyperspectral image) and can be expressed as: Y sfe =F sfe (Y),

[0072] In the above formula, F sfe(·) represents a convolution operation with a kernel size of 1×1.

[0073] Then, the spatial spectrum shallow feature Y sfe The convolution kernel size is 3×3 and then fed into the symmetric convolutional neural network to further extract the spatial spectrum deep features X of the RGB image Y. sase , which can be expressed as: X sase =FCNN(F↑(Y sfe )),

[0074] In the above formula, F ↑ (·) is the feature dimension-raising operation, F CNN (.) is a symmetric convolutional neural network.

[0075] Finally, the frequency information graph X in the spatial domain freq Fusion into the spatial spectrum domain features of RGB image Y to generate a hyperspectral image The function expression is:

[0076] In the above formula, F conv3×3 Indicates the convolution operation with a convolution kernel size of 3×3, F conv1×1 Indicates a convolution operation with a convolution kernel size of 1×1, X freq It is the frequency information diagram in the spatial domain.

[0077] As shown in Figure 2, the symmetric convolutional neural network (symmetric CNN) described in this embodiment includes a feature extraction upper branch, a spatial attention module SA, a feature extraction lower branch, a channel stacking module CAT and a convolution module. The feature extraction upper branch includes N local feature extraction modules LFEM connected in sequence. The feature map of the input symmetric convolutional neural network enters from the first local feature extraction module LFEM of the feature extraction upper branch, and the local feature extraction modules LFEM in the feature extraction upper branch all include three outputs, the first output is to the channel stacking module CAT, the second output is to a corresponding spatial attention module SA, and the third output of the first N-1 local feature extraction modules LFEM is to the next local feature extraction module LFEM in the feature extraction upper branch. The lower branch includes N local feature extraction modules LFEM. The output of the first spatial attention module SA is fused with the output of the Nth local feature extraction module LFEM in the feature extraction upper branch and serves as the input of the first local feature extraction module LFEM in the feature extraction lower branch. The outputs of the remaining spatial attention modules SA are fused with the output of the previous local feature extraction module LFEM in the feature extraction lower branch and serve as the input of the corresponding local feature extraction module LFEM in the feature extraction lower branch. Each local feature extraction module LFEM in the feature extraction lower branch includes an output connected to the channel stacking module CAT. The output of the channel stacking module CAT is refined by a convolution module with a convolution kernel size of 1×1 to obtain the spatial spectrum deep feature X of the RGB image Y. sase The output features of each local feature extraction module LFEM retain the useful information at the current stage. Therefore, the method proposed in this embodiment concatenates the output features of all local feature extraction modules LFEM of the upper and lower branches in the channel dimension and uses a 1×1 convolution to refine these features to produce a higher-level feature representation. Finally, through these operations, the final output feature X of the spatial spectral domain information extraction branch is obtained. sase By combining the output features of the two branches, the network can better capture the texture and structural information of the input image. At the same time, 1×1 convolution is used to refine the features, which also helps to reduce the dimension of the features and speed up the calculation, achieving the goal of lightweight model and high performance. The deep spatial spectral feature X of the RGB image Y is obtained. sase The process is expressed by the formula:

[0078] Among them, F conv1×1 represents 1×1 convolution, F CAT Indicates a channel stacking module.

[0079] In this embodiment, a symmetric CNN spatial spectrum feature extraction structure is designed. The symmetric CNN consists of some paired local feature extraction modules LFEM and spatial attention modules SA. Symmetric CNN is a neural network with a dual-branch structure. Here, the shallow feature Y is first sfe The image is then fed into the upper branch for feature extraction. In the upper branch, the output of the previous LFEM becomes the input to the next LFEM. This output also becomes part of the input to the corresponding LFEM in the lower branch through the spatial attention module (SA). More importantly, the network employs a parameter sharing mechanism for the corresponding LFEMs between the two branches, enabling the network to learn image features more efficiently while significantly reducing the number of parameters.

[0080] For the feature extraction branch, the specific process can be expressed as:

[0081] in, Represents the output features of the i-th LFEM in the upper branch, there are 3 in total Represents the i-th LFEM. For the feature extraction branch, the specific process can be expressed as:

[0082] in, It can be seen from the above formula that for the first LFEM in the lower branch, since there is no previous output feature, the output of the last LFEM in the upper branch is selected as the output feature of the 0th LFEM in the lower branch. represents the i-th LFEM, and its parameters are consistent with those in the upper branch. represents the i-th spatial attention module.

[0083] Based on the excellent performance of the dense residual block, this embodiment proposes a local feature extraction module LFEM as the core module of the symmetric CNN in the spatial spectrum domain extraction branch. As shown in Figure 5, the local feature extraction module LFEM of this embodiment includes three multi-feature fusion dual attention modules MFFDAB, two group convolutions with a convolution kernel size of 1×1, a channel stacking module and a convolution with a convolution kernel size of 1×1. In this example, the multi-feature fusion dual attention module MFFDAB replaces the "convolution layer-activation function" combination in the traditional dense residual block, and adds a 1×1 group convolution for feature dimensionality reduction. The function expression of the local feature extraction module LFEM for feature extraction of the input feature map is:

[0084] In the above formula, They are the output features of three multi-feature fusion dual attention modules MFFDAB, Represent three multi-feature fusion dual attention modules MFFDAB, and Respectively represent the input and output features of the local feature extraction module LFEM, F CAT Indicates the channel stacking module, F conv1×1 represents a convolution with a kernel size of 1×1, and They represent two group convolutions with kernel size of 1×1.

[0085] The multi-feature fusion dual attention module MFFDAB in this embodiment is used to extract deep features in the image spatial spectral domain and improve the model's expressiveness. As shown in Figure 6, the multi-feature fusion dual attention module MFFDAB in this embodiment is used to extract deep features in the image spatial spectral domain. The function expression for extracting deep features in the image spatial spectral domain is: Y CA =F CA (F CAT (Y1,Y2,Y3)), Y SA =F SA (F CAT (Y1,Y2,Y3)),

[0086] In the above formula, Y1~Y3 are three groups of intermediate features, F conv1×1 Indicates a convolution with a kernel size of 1×1, F conv3×3 represents a convolution with a kernel size of 3×3, F conv5×5 represents a convolution with a kernel size of 5×5, F conv7×7 represents a convolution with a kernel size of 7×7, F CA represents the channel attention module, F SA represents the spatial attention module, Y CA is the output feature of the channel attention module, Y SA is the output feature of the spatial attention module, and They represent the input and output features of the jth multi-feature fusion dual-attention module MFFDAB. As shown in Figure 6, the input is first passed through convolutional layers with different convolution kernel sizes, with convolution sizes of 3×3, 5×5, and 7×7, respectively. This aims to obtain features with different receptive fields, capture more feature information, and improve the model's accuracy and generalization. The feature channels of the two different receptive fields are then stacked and fused, and then an overall fusion is performed to obtain the multi-feature fusion features. The channel attention module CA then uses the weight of each channel to weight the feature channels. The spatial attention module SA models the spatial contextual relationship and reweights each pixel. Finally, the outputs of the two attention modules are superimposed and fused to obtain the final feature representation. As shown in Figure 7, the spatial attention module in this embodiment processes the input features by performing average pooling and maximum pooling on the input features, stacking the processed results, and then performing a convolution operation with a convolution kernel size of 7×7 and a Sigmoid activation function before concatenating them with the input features to obtain the output features. As shown in Figure 8, the channel attention module in this embodiment processes the input features by performing adaptive pooling, convolution with a convolution kernel size of 1×1, activation of the Relu activation function, convolution with a convolution kernel size of 1×1, activation of the Sigmoid activation function, and then connecting with the input features to obtain output features. In this embodiment, a feature fusion module FFM is designed to fuse the spatial spectral domain features, frequency domain information and initial features to reconstruct a high-fidelity hyperspectral image. Since the spatial spectral domain features, frequency domain information and initial feature dimensions are inconsistent and cannot be simply superimposed and fused, it is necessary to design a feature fusion module (FFM) to fuse these three features and output the final super-resolution hyperspectral image. Since the spatial spectral domain features and frequency domain information have been fully extracted.

[0087] As shown in FIG2 , steps S1 to S2 of this embodiment are implemented based on a hyperspectral computing imaging model, which includes:

[0088] A frequency domain transformation module, configured to execute step S1;

[0089] A frequency domain learning module FDLM, configured to execute step S2;

[0090] A frequency domain inverse transform module, configured to execute step S3;

[0091] Feature fusion module FFM, used to execute step S4;

[0092] The spatial spectrum domain feature extraction module is used to extract the spatial spectrum domain features of the RGB image Y;

[0093] The loss function L used in the training of the hyperspectral computational imaging model is all The function expression of L is:all =L sase +γL freq ,

[0094] In the above formula, L sase is the loss function in the empty spectrum domain, γ is the balance parameter, L freq is a frequency domain loss function, the spatial spectral domain loss function is the L1 norm between the hyperspectral image generated by the hyperspectral computing imaging model and the real hyperspectral image as a label, and the frequency domain loss function is the L1 norm between the frequency information map of the spatial domain generated by the hyperspectral computing imaging model and the frequency information map of the spatial domain in the real hyperspectral image. In terms of the loss function, this embodiment also takes the difference between the hyperspectral computing imaging and the real hyperspectral image in the frequency domain as one of the constraints of the network, and designs the above-mentioned joint loss function of the spatial spectral domain and the frequency domain. Define X as the real hyperspectral image, represents the reconstructed hyperspectral image, X freq and are the real hyperspectral frequency domain information and the reconstructed hyperspectral frequency domain information respectively, then L sase and L freq Can be expressed as:

[0095] In the above formula, ||.||1 represents the L1 norm.

[0096] In order to verify the hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion in this embodiment, a verification experiment was conducted using the CAVE dataset and Harvard data in this embodiment. The CAVE dataset contains a total of 32 indoor images, each with 31 bands, a spectral resolution of 400-700nm, and a spatial resolution of 512*512; the Harvard dataset contains 50 indoor and outdoor scene images taken under natural daylight, with a wavelength range of 420-720nm and a spatial resolution of 1040*1392. In generating experimental data, based on the mapping relationship between RGB and HSI, the hyperspectral image is spectrally downsampled to obtain the corresponding RGB dataset, which constitutes the RGB-HSI dataset required for the experiment. In the experiment, the spectral response function of the Nikon D700 camera is selected as the spectral downsampling matrix. In terms of the division of training data and test data, 20 pairs of data (about 60%) are randomly selected from the CAVE dataset as the training set, and the remaining 12 pairs of data are used as the test set. In the Harvard dataset, 35 pairs of data (70%) were randomly selected as training sets, and the remaining 15 pairs were used for testing. In order to accurately evaluate the performance of the proposed method and more intuitively compare the pros and cons of different methods, four widely used objective evaluation indicators were selected, namely, spectral angular distance SAM, root mean square error RMSE, global image quality index UIQI, and structural similarity SSIM. This embodiment is implemented using a deep learning framework, and the kaiming initialization method is used to initialize the network parameters. At the same time, the network parameters are optimized based on the Adam optimizer. The optimizer parameters are set to ∈ = 10 -8 , β1 = 0.9, β2 = 0.999. The initial learning rate was set to 0.0001, and a cosine annealing strategy was used to decay the learning rate. The training set image block size was 64*64. To increase the sample size, 32 was set as the overlap size during cropping. During the entire network training process, a total of 100 iterations were set with a batch size of 32. In the experiments, the network was compared with six existing spectral computational imaging methods, including:

[0097] HSCNN-R (Shi Z, Chen C, Xiong Z, et al. HSCNN+: Advanced CNN-based hyperspectral recovery from RGB images. In: Proc of IEEE Conference on Computer Vision and Pattern Recognition Workshops. 2018, 939-947);

[0098] DFMN(Zhang L,Lang Z,Wang P,et al.Pixel-aware deep function-mixture network for spectral superresolution.In:Proc ofAAAI Conference on Artificial Intelligence,volume 34.2020,12821-12828);

[0099] HRNet(Zhao Y,Po L M,Yan Q,et al.Hierarchical regression network for spectral reconstruction from RGB images.In:Proc of IEEE / CVF Conference on Computer Vision and Pattern RecognitionWorkshops.2020,422-423);

[0100] AWAN(Li J,Wu C,Song R,et al.Adaptive weighted attention network with camera spectral sensitivity prior for spectral reconstruction from RGB images.In:Proc of IEEE / CVF Conference on ComputerVision andPatternRecognitionWorkshops.2020,462-463);

[0101] HSRNet(He J,Li J,Yuan Q,et al.Spectral response function-guided deep optimization-driven network for spectral super-resolution.IEEE Transactions on Neural Networks and Learning Systems,2021,33(9):4213-4227);

[0102] Prinet+(Hang R, Liu Q, Li Z. Spectral super-resolution network guided by intrinsic properties of hyperspectral imagery. IEEE Transactions on Image Processing, 2021, 30: 7256-7265).

[0103] Table 1 shows the various objective evaluation indicators of different methods in the CAVE dataset and Harvard dataset.

[0104] Table 1: Objective performance indicators of the method in this embodiment and six existing methods on the CAVE and Harvard datasets.

[0105] As can be observed in Table 1, the method of this embodiment achieved the best evaluation indicators in both datasets, indicating that the proposed method can more effectively improve the spatial and spectral quality of hyperspectral images. In addition to objective evaluation indicators, this embodiment also performed a visual evaluation to more intuitively experience the effects of different spectral super-resolution imaging methods. In this visual evaluation, a band image of the super-resolved hyperspectral image is displayed to evaluate the spatial quality, and a specific area is magnified. The spectral quality of the image is evaluated by analyzing the spectral error map. The spectral error map is obtained by calculating the spectral error between the original hyperspectral image and the reconstructed image, and is related to the evaluation indicator SAM.

[0106] Figure 9 is a visual evaluation diagram of different methods on the CAVE dataset, where (a) is the reconstructed band diagram of the HSCNN-R method, (b) is the reconstructed band diagram of the DFMN method, (c) is the reconstructed band diagram of the HRNet method, (d) is the reconstructed band diagram of the AWAN method, (e) is the spectral error diagram of the HSCNN-R method, (f) is the spectral error diagram of the DFMN method, (g) is the spectral error diagram of the HRNet method, (h) is the spectral error diagram of the AWAN method, (i) is the reconstructed band diagram of the HSRNet method, (j) is the reconstructed band diagram of the Prinet+ method, (k) is the reconstructed band diagram of the method of this embodiment, (l) is the reference reconstructed band diagram, (m) is the spectral error diagram of the HSRNet method, (n) is the spectral error diagram of the Prinet+ method, and (o) is the spectral error diagram of the method of this embodiment. 9 , the hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion in this embodiment obtains the smallest spectral error map, which indicates that the hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion in this embodiment can well reconstruct spatial detail information.

[0107] In summary, the hyperspectral computational imaging method for space-spectrum-frequency multi-domain fusion in this embodiment includes a frequency domain information learning module, a space-spectrum-frequency multi-domain information extraction module, and a feature fusion module. In the frequency domain information learning module, the RGB image is converted to the frequency domain using an offline discrete cosine transform (DCT). Frequency domain features are extracted using the frequency domain learning module (FDLM). Finally, the frequency domain information is transformed to the spatial domain using an inverse discrete cosine transform (IDCT). In the space-spectrum-frequency information extraction branch, a symmetrical convolutional neural network (CNN) structure is constructed based on the proposed local feature extraction module (LFEM) and the spatial attention module to extract the image's space-spectrum-domain features. The feature fusion module (FFM) fuses the features extracted from these two branches with the initial features to generate a high-resolution hyperspectral image. This method addresses the issues of poor detail information and low reconstruction accuracy in existing spectral imaging technologies. Unlike traditional network models, the hyperspectral computational imaging method for space-spectrum-frequency multi-domain fusion in this embodiment introduces frequency domain information into hyperspectral imaging, constructs a space-spectrum-frequency multi-domain feature fusion framework, and designs a symmetrical convolutional neural network space-spectrum feature extraction structure to achieve high-fidelity reconstruction of the target spectrum.

[0108] In addition, this embodiment also provides a hyperspectral computational imaging system for spatial-spectral-frequency multi-domain fusion, comprising an interconnected microprocessor and memory, wherein the microprocessor is programmed or configured to execute the hyperspectral computational imaging method for spatial-spectral-frequency multi-domain fusion. This embodiment also provides a computer-readable storage medium storing a computer program for being programmed or configured by the microprocessor to execute the hyperspectral computational imaging method for spatial-spectral-frequency multi-domain fusion.

[0109] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow chart and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine so that the instruction executed by the processor of the computer or other programmable data processing device produces a device for realizing the function specified in one flow chart flow or multiple flows and / or one block or multiple blocks of the block diagram. These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.

[0110] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A hyperspectral computational imaging method with spatial-spectral-frequency multi-domain fusion, characterized in that: include: S1, convert the RGB image Y to the frequency domain to obtain the frequency domain feature map Y freq ; S2, from the frequency domain feature map Y freq Extract frequency information from Including: S2.1, the frequency domain feature map Y freq According to the three channels of Y, Cb, and Cr in the YCbCr space, the channel dimension is divided into three equal parts. First, the features of each equal part are spatially downsampled and divided into two parts: low-frequency information and high-frequency information. Then, the low-frequency information and high-frequency information in each equal part are channel-recombined to obtain the low-frequency features. With high frequency characteristics And the size is Where C is the frequency domain feature map Y freq The number of channels, k represents the size after spatial downsampling; S2.2, respectively, for low-frequency features With high frequency characteristics m dense residual blocks RDB are used for mapping learning, and the GeLu function is used as the activation function in the dense residual block RDB; S2.3, the low-frequency features after mapping learning and the high-frequency features are stacked in the channel dimension, and then n dense residual blocks RDB are used for mapping learning to obtain a frequency reorganization map, and then the frequency reorganization map is upsampled to obtain a frequency information map S3, frequency information map Transform to the spatial domain to obtain the frequency information graph X in the spatial domain freq ; S4, the frequency information map of the spatial domain X freq Fusion into the spatial spectral domain features of the RGB image Y generates a hyperspectral image The acquisition of the spatial spectrum domain features of the RGB image Y includes: extracting the spatial spectrum shallow feature Y of the RGB image Y using a convolution with a convolution kernel size of 1×1 sfe , and the shallow feature Y sfe The convolution kernel size is 3×3, and then the convolution kernel is fed into a symmetric convolutional neural network to further extract the spatial spectrum deep features X of the RGB image Y. sase As the spatial spectrum domain feature of RGB image Y; the frequency information map x in the spatial domain freq Fusion into the spatial spectral domain features of the RGB image Y generates a hyperspectral image The function expression is: In the above formula, F conv3×3 represents a convolution operation with a convolution kernel size of 3×3, F conv1×1 represents a convolution operation with a convolution kernel size of 1×1, X freq It is the frequency information diagram in the spatial domain.

2. The hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion according to claim 1 is characterized in that: In step S1, the RGB image Y is converted to the frequency domain to obtain the frequency domain feature map Y freq include: S1.1, convert the input RGB image Y into a YCbCr space image and divide it into image blocks according to color channels; S1.2, using two-dimensional offline discrete cosine transform DCT to project each image block into the frequency domain to obtain frequency coefficients; S1.3, the frequency coefficients of each image block are stacked together in the order of the three channels Y, Cb, and Cr, and vectorized to obtain the frequency band; then the frequency band is reorganized in a Z-shape to obtain the frequency domain feature map Y freq , and the frequency domain feature map Y freq Each channel corresponds to a frequency band.

3. The hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion according to claim 1 is characterized in that: In step S3, the frequency information map Transform to the spatial domain to obtain the frequency information graph X in the spatial domain freq include: S3.1, dividing the frequency information map into image blocks; S3.2, project each image block into the spatial domain using a two-dimensional offline inverse discrete cosine transform IDCT; S3.3, combine the image blocks projected into the spatial domain to obtain the frequency information map x in the spatial domain freq .

4. The hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion according to claim 1 is characterized in that: The symmetric convolutional neural network includes a feature extraction upper branch, a spatial attention module SA, a feature extraction lower branch, a channel stacking module CAT and a convolution module. The feature extraction upper branch includes N local feature extraction modules LFEM connected in sequence. The feature map of the input symmetric convolutional neural network enters from the first local feature extraction module LFEM of the feature extraction upper branch, and the local feature extraction modules LFEM in the feature extraction upper branch all include three outputs, the first output is to the channel stacking module CAT, the second output is to a corresponding spatial attention module SA, the third output of the first N-1 local feature extraction modules LFEM is to the next local feature extraction module LFEM in the feature extraction upper branch, and the feature extraction lower branch includes N local The output of the first spatial attention module SA is fused with the output of the Nth local feature extraction module LFEM in the feature extraction upper branch as the input of the first local feature extraction module LFEM in the feature extraction lower branch, and the outputs of the remaining spatial attention modules SA are fused with the output of the last local feature extraction module LFEM in the feature extraction lower branch as the input of the corresponding local feature extraction module LFEM in the feature extraction lower branch, and each local feature extraction module LFEM in the feature extraction lower branch includes an output connected to a channel stacking module CAT, and the output of the channel stacking module CAT is refined by a convolution module with a convolution kernel size of 1×1 to obtain the spatial spectrum deep feature X of the RGB image Y sase .

5. The hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion according to claim 4 is characterized in that: The local feature extraction module LFEM includes three multi-feature fusion dual attention modules MFFDAB, two group convolutions with a convolution kernel size of 1×1, a channel stacking module and a convolution with a convolution kernel size of 1×1. The function expression of the local feature extraction module LFEM for extracting features from the input feature map is: In the above formula, They are the output features of three multi-feature fusion dual attention modules MFFDAB, They represent three multi-feature fusion dual attention modules MFFDAB, and They represent the input and output features of the local feature extraction module LFEM, F CAT Indicates the channel stacking module, F conv1×1 represents a convolution with a kernel size of 1×1, and Respectively represent two group convolutions with a convolution kernel size of 1×1; the multi-feature fusion dual attention module MFFDAB is used to extract deep features of the image in the spatial spectral domain, and the function expression for extracting deep features of the image in the spatial spectral domain is: <h2 style=";text-align:left;direction:ltr">Y<h2 style=";text-align:left;direction:ltr"> CA <h2 style=";text-align:left;direction:ltr"> =F<h2 style=";text-align:left;direction:ltr"> CA <h2 style=";text-align:left;direction:ltr"> (F<h2 style=";text-align:left;direction:ltr"> CAT <h2 style=";text-align:left;direction:ltr"> (Y1,Y2,Y3)), <h2 style=";text-align:left;direction:ltr">Y<h2 style=";text-align:left;direction:ltr"> SA <h2 style=";text-align:left;direction:ltr"> =F<h2 style=";text-align:left;direction:ltr"> SA <h2 style=";text-align:left;direction:ltr"> (F<h2 style=";text-align:left;direction:ltr"> CAT <h2 style=";text-align:left;direction:ltr"> (Y1,Y2,Y3)),<h2 style=";text-align:left;direction:ltr"> In the above formula, Y1~Y3 are three groups of intermediate features, F conv1×1 represents a convolution with a kernel size of 1×1, F conv3×3 represents a convolution with a kernel size of 3×3, F conv5×5 represents a convolution with a kernel size of 5×5, F conv7×7 represents a convolution with a kernel size of 7×7, F CA represents the channel attention module, F SA represents the spatial attention module, Y CA is the output feature of the channel attention module, Y SA is the output feature of the spatial attention module, and They respectively represent the input and output features of the j-th multi-feature fusion dual attention module MFFDAB.

6. The hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion according to claim 5 is characterized in that: Steps S1 to S4 are implemented based on a hyperspectral computing imaging model, and the hyperspectral computing imaging model includes: A frequency domain transformation module, used to execute step S1; A frequency domain learning module FDLM, used to execute step S2; A frequency domain inverse transform module, used to execute step S3; A feature fusion module FFM, used to execute step S4; A spatial spectrum domain feature extraction module is used to extract the spatial spectrum domain features of the RGB image Y; The loss function L used in the training of the hyperspectral imaging model is all The function expression is: THE all =L sase +γL freq , In the above formula, L sase is the loss function in the empty spectrum domain, γ is the balance parameter, L freq It is a frequency domain loss function. The spatial spectral domain loss function is the L1 norm between the hyperspectral image generated by the hyperspectral computing imaging model and the real hyperspectral image as a label. The frequency domain loss function is the L1 norm between the frequency information map in the spatial domain generated by the hyperspectral computing imaging model and the frequency information map in the spatial domain in the real hyperspectral image.

7. A hyperspectral computational imaging system with spatial-spectral-frequency multi-domain fusion, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion as described in any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored therein, characterized in that: The computer program is used to be programmed or configured by a microprocessor to execute the hyperspectral computational imaging method of space-spectrum-frequency multi-domain fusion as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Remote sensing image missing data reconstruction method based on deep convolutional neural network

    CN108876754A

  • A hyperspectral and multispectral image fusion method based on a two-way dense residual network

    CN109636769A

  • Multi-scale and global feature hyperspectral and multispectral remote sensing fusion method

    CN115861083A

  • Garbage metal rescreening method for reconstructing hyperspectral image based on RGB image

    CN116665051A

  • Hyperspectral calculation imaging method and system based on spatial-spectral frequency multi-domain fusion and medium

    CN117314757A

Cited By

  • Image processing method and system

    CN120339146A

  • Image processing method and system

    CN120339146B

  • OCT image classification method

    CN120807476A

  • Physique and visceral organ identification method based on multi-domain fusion

    CN120852864A

  • Ground penetrating radar data multi-frequency fusion method, device, equipment, medium and product

    CN120951277A