Low resolution hyperspectral image processing method, apparatus, computer program product
By using a self-attention mechanism model to group and perform sub-pixel convolution processing on low-resolution hyperspectral images, the problem of low spatial resolution in hyperspectral images is solved, and the accuracy of target recognition is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2026-04-14
AI Technical Summary
Because hyperspectral sensors need to receive optical information from more spectral bands, the spatial resolution of hyperspectral images is low, which reduces the accuracy of target recognition.
A self-attention mechanism model is used to group low-resolution hyperspectral images. By connecting multiple self-attention mechanism models and performing sub-pixel convolution processing, the spatial resolution of the images is improved, thereby increasing the accuracy of target recognition.
By enhancing the learning ability of global spatial information and long-range features through a self-attention mechanism model, the spatial resolution of low-resolution hyperspectral images is improved, thereby increasing the accuracy of target recognition.
Smart Images

Figure CN115439325B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a low-resolution hyperspectral image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the development of image processing technology, hyperspectral imaging is a sophisticated technique that can capture and analyze the spectrum at individual points within a spatial region. Based on hyperspectral sensors, optical information in spectral bands can be received, and based on this optical information, hyperspectral images can be obtained. Furthermore, target localization can be achieved based on hyperspectral images.
[0003] However, because hyperspectral sensors need to ensure the reception of optical information across more spectral bands, the spatial resolution of the hyperspectral images obtained is typically low, which reduces the accuracy of target recognition. Summary of the Invention
[0004] Therefore, it is necessary to provide a low-resolution hyperspectral image processing method that can improve the accuracy of target recognition in order to address the above-mentioned technical problems.
[0005] In a first aspect, this application provides a low-resolution hyperspectral image processing method, comprising: acquiring shallow features of each grouped image of a low-resolution hyperspectral image; each grouped image being obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each grouped image being obtained by performing a first convolution process on each grouped image respectively; processing the shallow features of each grouped image based on a processing network, and performing pixel-wise addition of the shallow features of each grouped image with the processed shallow features of each grouped image to obtain the global depth of each grouped image. Layer features; the processing network includes multiple self-attention mechanism models, which have the same structure, and the output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to it; based on the spectral features of the low-resolution hyperspectral image, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image is obtained; the spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each grouped image, and then performing the first convolution processing and cascade processing on the global deep features of each grouped image after the sub-pixel convolution processing. Secondly, this application provides a low-resolution hyperspectral image processing apparatus, the apparatus comprising: an acquisition module for acquiring shallow features of each grouped image of a low-resolution hyperspectral image; each grouped image is obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each grouped image are obtained by performing a first convolution process on each grouped image respectively; and a feature determination module for processing the shallow features of each grouped image based on a processing network, and performing pixel-wise addition of the shallow features of each grouped image with the processed shallow features of each grouped image to obtain each grouped image. The processing network includes multiple self-attention mechanism models, each with the same structure, and the output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to it. A processing module is used to obtain a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image based on its spectral features. The spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing the first convolution processing and cascading processing on the global deep features of each group of images after the sub-pixel convolution processing.
[0006] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, performs the following steps: acquiring shallow features of each grouped image of a low-resolution hyperspectral image; each grouped image is obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each grouped image are obtained by performing a first convolution process on each grouped image respectively; processing the shallow features of each grouped image based on a processing network, and performing pixel-wise addition of the shallow features of each grouped image with the processed shallow features of each grouped image to obtain... The global deep features of each group of images; the processing network includes multiple self-attention mechanism models, the multiple self-attention mechanism models have the same structure, and the output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to each self-attention mechanism model; based on the spectral features of the low-resolution hyperspectral image, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image is obtained; the spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing the first convolution processing and cascade processing on the global deep features of each group of images after sub-pixel convolution processing.
[0007] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps: acquiring shallow features of each grouped image of a low-resolution hyperspectral image; each grouped image is obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each grouped image are obtained by performing a first convolution process on each grouped image respectively; processing the shallow features of each grouped image based on a processing network, and performing pixel-wise addition of the shallow features of each grouped image with the processed shallow features of each grouped image to obtain the shallow features of each grouped image. The global deep features of the group of images; the processing network includes multiple self-attention mechanism models, the multiple self-attention mechanism models have the same structure, and the output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to each self-attention mechanism model; based on the spectral features of the low-resolution hyperspectral image, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image is obtained; the spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing the first convolution processing and cascade processing on the global deep features of each group of images after sub-pixel convolution processing.
[0008] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps: acquiring shallow features of each grouped image of a low-resolution hyperspectral image; each grouped image is obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each grouped image are obtained by performing a first convolution process on each grouped image respectively; processing the shallow features of each grouped image based on a processing network, and performing pixel-wise addition of the shallow features of each grouped image with the processed shallow features of each grouped image to obtain each grouped image. The processing network includes multiple self-attention mechanism models, each with the same structure. The output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to it. Based on the spectral features of the low-resolution hyperspectral image, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image is obtained. The spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing the first convolution processing and cascading processing on the global deep features of each group of images after the sub-pixel convolution processing.
[0009] The aforementioned low-resolution hyperspectral image processing methods, apparatus, computer equipment, storage media, and computer program products, By grouping low-resolution hyperspectral images based on their spectral data, we can obtain grouped images. Each group is then subjected to a first convolution to obtain shallow features. Further processing of these shallow features using a processing network, along with pixel-wise summation of the shallow features and processed features, yields the global deep features of each group. This processing network consists of multiple self-attention mechanisms with identical structures. The output of each self-attention mechanism serves as the input to the next self-attention mechanism model it connects to. Sub-pixel convolution is then applied to the global deep features of each group, followed by a first convolution and cascaded processing. Finally, based on these spectral features, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image can be obtained. In this way, by using a self-attention mechanism model, the learning ability of global spatial information and long-range features can be enhanced, enabling the acquisition of more spectral band optical information of low-resolution hyperspectral images, thereby improving the spatial resolution of low-resolution hyperspectral images and improving the accuracy of target recognition. Attached Figure Description
[0010] Figure 1 This is an application environment diagram of a low-resolution hyperspectral image processing method in one embodiment;
[0011] Figure 2 This is a flowchart illustrating a low-resolution hyperspectral image processing method in one embodiment;
[0012] Figure 3 This is a flowchart illustrating how a processing network processes the shallow features of each group of images, and then adds the shallow features of each group of images pixel by pixel to obtain the global deep features of each group of images.
[0013] Figure 4 This is a schematic diagram illustrating the process of obtaining intermediate features based on shallow features of grouped images in one embodiment.
[0014] Figure 5 This is a flowchart illustrating the process of obtaining the input features of the second self-attention mechanism model in the corresponding processing network based on the corresponding intermediate features in one embodiment.
[0015] Figure 6 This is a schematic diagram illustrating how, in one embodiment, the input features of the second self-attention mechanism model in the corresponding processing network are obtained based on the corresponding intermediate features.
[0016] Figure 7 This is a flowchart illustrating the process of processing the input features of a 3D convolutional module to obtain the output features of the corresponding convolutional module in one embodiment.
[0017] Figure 8 This is a schematic diagram illustrating how a 3D convolutional module processes the input features of a corresponding convolutional module to obtain the output features of that convolutional module in one embodiment.
[0018] Figure 9 This is a schematic diagram of the process of obtaining a high-resolution hyperspectral image corresponding to a low-resolution hyperspectral image based on the spectral features of a low-resolution hyperspectral image in one embodiment.
[0019] Figure 10 This is a schematic diagram of the structure of a low-resolution hyperspectral image processing method in one embodiment;
[0020] Figure 11 and Figure 12 The images show spatial image details and error plots in the 550nm and 600nm spectral bands of the CAVE dataset test images at magnifications of 4x and 8x, respectively.
[0021] Figure 13 and Figure 14 The images show spatial image details and error plots at the 60th and 80th spectral bands in the test results of different methods using the Chikusei dataset at magnifications of 4x and 8x, respectively.
[0022] Figure 15 This is a structural block diagram of a low-resolution hyperspectral image processing device in one embodiment;
[0023] Figure 16 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0025] With the development of image processing technology, hyperspectral imaging technology is a sophisticated technique that can capture and analyze the spectrum at individual points within a spatial region. Because it can detect the unique spectral features of a single object at different spatial locations, it can detect substances that are visually indistinguishable.
[0026] Hyperspectral sensors are sensors based on hyperspectral imaging technology. Typically, hyperspectral sensors can receive optical information in spectral bands, and based on this optical information, hyperspectral images can be obtained. Furthermore, target localization can be achieved based on hyperspectral images. Compared with the RGB three bands of natural images, the bands in hyperspectral images are narrower spectral bands, for example, 10nm-20nm. Hyperspectral images contain tens to hundreds of spectral bands.
[0027] However, because hyperspectral sensors need to ensure the reception of optical information in more spectral bands, the spatial resolution of hyperspectral images obtained is usually low. Low spatial resolution results in unclear pixel outlines, making it impossible to accurately identify whether a target is being measured in target detection and image recognition tasks, thus reducing the accuracy of target recognition and image recognition.
[0028] In the field of super-resolution, upstream vision tasks involving super-resolution can be processed using specific pre-trained Transformer structures. With only fine-tuning on specific task datasets, better recovery results than convolutional neural network methods can be obtained. Although IPT and SwinIR models have been explored for super-resolution of natural images and have achieved better reconstruction results than pure convolutional neural networks, in the field of hyperspectral super-resolution of single images, the mainstream network structure still uses 2D or 3D convolutional neural networks for feature extraction. The Transformer structure has not been used to design the network. That is, the Transformer structure has not been used to improve the global receptive field and spatial clarity in the hyperspectral super-resolution reconstruction of single images.
[0029] In view of this, this application provides a low-resolution hyperspectral image processing method, which can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. The data storage system can store the images that server 104 needs to process. The data storage system can be integrated on server 104 or placed on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc.
[0030] Specifically, after acquiring a low-resolution hyperspectral image, terminal 102 can transmit the low-resolution hyperspectral image to server 104. Server 104, based on the spectral count of the low-resolution hyperspectral image, performs grouping processing on the low-resolution hyperspectral image to obtain individual grouped images. Each grouped image undergoes a first convolution process to obtain shallow features. Server 104 then processes these shallow features using a processing network and adds them pixel-wise to obtain the individual grouped images. The global deep features are obtained by the processing network, which consists of multiple self-attention mechanism models with the same structure. The output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to it. Then, the server 104 performs sub-pixel convolution processing on the global deep features of each group of images, and performs first convolution processing and cascade processing on the global deep features of each group of images after sub-pixel convolution processing to obtain the spectral features of the low-resolution hyperspectral image. Based on the spectral features of the low-resolution hyperspectral image, the server 104 can obtain the high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image.
[0031] In one embodiment, such as Figure 2 As shown, a low-resolution hyperspectral image processing method is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:
[0032] S202, Obtain the shallow features of each group of images from the low-resolution hyperspectral image.
[0033] In this embodiment, the grouped images of the low-resolution hyperspectral image are obtained by grouping the low-resolution hyperspectral image based on its spectral channels. For example, if the low-resolution hyperspectral image is represented by ILR, then I... LR Divided into N groups, then I LR It can be represented as: Where N is less than the number of spectra in the low-resolution hyperspectral image, and the overlap coefficient α is 1.
[0034] In this embodiment, the shallow features of each group of images are obtained by performing a first convolution process on each group of images. The first convolution process is a preprocessing step, and it can be a 3×3 convolution process. By performing a 3×3 convolution process on each group of images, not only can the line-level features of each group of images be extracted, but the feature channel dimension can also be increased to retain more information, thereby improving the accuracy of target recognition.
[0035] For example, if This represents the nth group image of a low-resolution hyperspectral image. This represents the 3×3 convolution process performed on the nth group of images. Let represent the shallow features obtained after performing a 3×3 convolution on the nth group image. It can be represented as:
[0036] S204. Based on the processing network, the shallow features of each group of images are processed separately, and the shallow features of each group of images and the processed features of each group of images are added pixel by pixel to obtain the global deep features of each group of images.
[0037] In this embodiment, the processing network includes multiple self-attention mechanism models, which are Transformer models. These Transformer models have the same structure, meaning they share parameters. Specifically, the processing network may include B Transformer models. By using these B Transformer models to process the shallow features of each group of images, the global deep features of each group of images can be obtained. For example, if... The shallow features representing the nth group of images. Let represent the global deep features of the nth group image, then It can be represented as: This represents the b-th Transformer model for the n-th group of images. The outputs of the B Transformer modules for the nth group image are the global deep features of the nth group image.
[0038] It should be noted that the B Transformer modules are connected by local skip connections. That is, the output of each Transformer model is the input of the next Transformer model connected by the local skip connections of each Transformer model. By processing the shallow features of each group of images through the B Transformer modules, the vanishing or exploding gradients can be reduced.
[0039] S206. Based on the spectral features of the low-resolution hyperspectral image, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image is obtained. The spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing a first convolution processing and a concatenation processing on the global deep features of each group of images after the sub-pixel convolution processing.
[0040] Specifically, by performing sub-pixel convolution processing on the global deep features of each group of images, the feature maps of each group of images can be magnified to half of the expected magnification factor. For example, if Represents the global spatial features of the nth group of images. This represents the sub-pixel convolution processing performed on the global deep features of the nth group image. The output after performing sub-pixel convolution on the global deep features of the nth group image is then... It can be represented as:
[0041] Specifically, by performing a first convolution on the global deep features of each group of images after subpixel convolution processing, the dimensionality of each group of images after subpixel convolution processing can be reduced to the input dimension. The first convolution is a 3×3 convolution. For example, if... This represents the output obtained after sub-pixel convolution processing of the global deep features of the nth group image. This represents the 3×3 convolution process performed on the output of the global deep features of the nth group image after sub-pixel convolution. Let represent the output obtained after performing a 3×3 convolution on the output of the global deep features of the nth group image after subpixel convolution. It can be represented as:
[0042] In this process, cascading operations can be used to combine the outputs of sub-pixel convolutions of global deep features from multiple grouped images into a single feature map with the same spectral dimension after a 3×3 convolution. For example, if concat(·) represents cascading processing in the spectral dimension, F branch F represents the spectral features of a low-resolution hyperspectral image. branch It can be represented as:
[0043] In summary, Figure 2In the illustrated embodiment, by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, we can obtain each group of low-resolution hyperspectral images. Each group of images undergoes a first convolution process to obtain shallow features. The shallow features of each group of images are then processed using a processing network, and the shallow features of each group of images are pixel-wise summed with the processed features of each group of images to obtain the global deep features of each group of images. The processing network consists of multiple self-attention mechanism models with identical structures. The output of each self-attention mechanism model serves as the input to the next self-attention mechanism model connected to it. Furthermore, by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing a first convolution and cascade processing on the global deep features of the group of images after sub-pixel convolution processing, we can obtain the spectral features of the low-resolution hyperspectral image. Based on the spectral features of the low-resolution hyperspectral image, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image can be obtained. In this way, by using a self-attention mechanism model, the learning ability of global spatial information and long-range features can be enhanced, enabling the acquisition of more spectral band optical information of low-resolution hyperspectral images, thereby improving the spatial resolution of low-resolution hyperspectral images and improving the accuracy of target recognition.
[0044] exist Figure 2 Based on the illustrated embodiments, in one embodiment, a flowchart is provided that processes the shallow features of each group of images using a processing network, and then adds the shallow features of each group of images pixel by pixel to obtain the global deep features of each group of images, as shown below. Figure 3 As shown, this method is applied to Figure 1 Taking server 104 as an example, the following steps are included:
[0045] S302, the shallow features of each group of images are determined as the input of the first self-attention mechanism model in the processing network, and the input is subjected to layer normalization processing to obtain the first feature after the corresponding layer normalization processing.
[0046] S304, the first feature after the corresponding layer normalization is subjected to a second convolution to obtain the first feature after the second convolution.
[0047] In this embodiment, by performing layer normalization on the shallow features of each group of images, the feature weight range can be changed. The second convolution process includes sequential 1×1 convolution and 3×3 depth-separable convolution.
[0048] S306, perform channel segmentation on the first feature after the corresponding second convolution to obtain the corresponding query feature, key feature and value feature after channel segmentation.
[0049] Among them, based on the content described in S302 to S306, the self-attention mechanism model is the Transformer model, if LN(·) represents the input to the first Transformer model, and W represents the layer normalization process. 1×1 (·) indicates a 1×1 convolution process, W DW3×3 (·) represents a 3×3 depthwise separable convolution, split(·) represents channel splitting, Q represents the query feature, K represents the key feature, and V represents the value feature. Therefore,
[0050] S308, the query features, key features and value features after the corresponding channel segmentation are subjected to tensor reshaping to obtain the corresponding reshaped query features, key features and value features.
[0051] S310, the corresponding reshaped key features are transposed, and the transposed key features are multiplied with the corresponding reshaped query features. The features after matrix multiplication are then processed by the softmax activation function to obtain the corresponding softmax activation function processed features.
[0052] S312, perform matrix multiplication on the features processed by the corresponding softmax activation function and the corresponding reshaped value features, and perform tensor reshaping on the features after matrix multiplication to obtain the corresponding reshaped first features.
[0053] In conjunction with the descriptions in S308 to S312, if Q, K, and V represent query features, key features, and value features respectively, then as well as These can be used to represent the reshaped query features, key features, and value features, respectively. The key features after transposition are represented by... This indicates that the features processed by the corresponding softmax activation function are... It can be represented as: The first feature after reshaping can be represented as Where Q,K,V∈R h×w×c ,but
[0054] S314, the corresponding reshaped first feature is subjected to a third convolution, and the first feature after the third convolution is added pixel by pixel to the input of the first self-attention mechanism model in the processing network to obtain the intermediate feature of the first self-attention mechanism model in the processing network.
[0055] The third convolution process is a 1×1 convolution process, if This represents the input to the first Transformer model. W represents the first feature after reshaping. 1×1 (·) indicates a 1×1 convolution process, F SA Let F represent the intermediate features of the first Transformer model. SA It can be represented as:
[0056] Combining the content shown in S302 to S314, such as Figure 4 The diagram illustrates a method for obtaining intermediate features based on shallow features of grouped images, where each group can be based on... Figure 4 The steps shown yield the corresponding intermediate features.
[0057] S316. Based on the corresponding intermediate features, obtain the input features of the second Transformer model in the corresponding processing network.
[0058] S318, the output features of the last self-attention mechanism model in the corresponding processing network are determined as the shallow features of each grouped image after processing.
[0059] The self-attention mechanism model is a Transformer model. The implementation method of obtaining the output features of the last Transformer model based on the input features of the last Transformer model can be referred to the content adaptation description in S302 to S316, which will not be repeated here.
[0060] S320: The shallow features of each group of images are added pixel by pixel to the shallow features of the processed group of images to obtain the global deep features of each group of images.
[0061] exist Figure 3 Based on the content shown, in one embodiment, a flowchart is provided to obtain the input features of the second self-attention mechanism model in the corresponding processing network according to the corresponding intermediate features, as shown below. Figure 5 As shown, this method is applied to Figure 1 Taking server 104 as an example, the following steps are included:
[0062] S502, perform layer normalization on the corresponding intermediate features to obtain the corresponding second feature after layer normalization.
[0063] S504, the second feature after the corresponding layer normalization is subjected to a second convolution to obtain the second feature after the second convolution.
[0064] S506, the second feature after the corresponding second convolution is processed into layers to obtain the corresponding layered convolution module input features and gated branch input features.
[0065] In this embodiment, the second convolution processing includes sequential 1×1 convolution processing and 3×3 depthwise separable convolution processing. Specifically, referring to the content described in S502 to S506, if F SA LN(·) represents intermediate features, and LN(·) represents layer normalization. 1×1 (·) indicates a 1×1 convolution process, W DW3×3 (·) indicates a 3×3 depth separable convolution, and split(·) indicates layered processing. This represents the input features of the convolutional module. If the gating branch input characteristics are represented, then...
[0066] S508 processes the input features of the corresponding convolutional module based on the 3D convolutional module to obtain the output features of the corresponding convolutional module.
[0067] Among them, if This represents the input features of the convolutional module. f represents the output feature of the convolutional module. 3D (·) indicates processing by the 3D convolution module. It can be represented as:
[0068] S510 processes the corresponding gated branch input features using the GELU activation function to obtain the corresponding gated branch output features.
[0069] Among them, if The input features of the gated branch are represented by GELU(·), and GELU(·) represents the GELU activation function processing. To represent the output characteristics of the gated branch, then It can be represented as:
[0070] S512 performs pixel-wise multiplication on the output features of the corresponding convolutional module and the output features of the corresponding gated branch to obtain the corresponding pixel-wise multiplied features.
[0071] S514, the features after multiplying the corresponding pixels are processed by a third convolution, and the features after the third convolution are added to the corresponding intermediate features by pixels to obtain the input features of the second self-attention mechanism model in the processing network.
[0072] The third convolution process is a 1×1 convolution, and the self-attention mechanism model is a Transformer model. Combining the descriptions in S512 and S514, if "⊙" represents pixel multiplication, W... 1×1 (·) indicates a 1×1 convolution. Indicates the output characteristics of the gated branch. F represents the output features of the convolutional module. SA Indicates intermediate features. Let the input feature representation of the second Transformer model be... It can be represented as:
[0073] Combination Figure 5 The content shown, in one embodiment, is as follows: Figure 6 As shown, this diagram illustrates how the input features of the second self-attention mechanism model in the corresponding processing network are obtained based on the corresponding intermediate features. The self-attention mechanism model is the Transformer model, and the structure shown in the diagram can be called the feedforward network structure of the Transformer model. In the feedforward propagation part, the Transformer model uses GELU as a gating mechanism to constrain the feedforward information and improve the information flow in the network. At the same time, a 3D convolutional module branch is added to the feedforward propagation part to strengthen the correlation between spatial and spectral information. By extracting spatial and spectral feature information, the reconstruction effect is further improved, and the acquisition of high-frequency spatial information is increased.
[0074] exist Figure 5 and Figure 6 Based on the content shown, in one embodiment, such as Figure 7 The diagram illustrates a process for processing the input features of a 3D convolutional module to obtain the output features of that module. This method is then applied to... Figure 1 Taking server 104 as an example, the explanation may include the following steps:
[0075] S702, the corresponding convolutional module input features are subjected to tensor reshaping processing, and the corresponding tensor reshaping features are subjected to dimensional expansion processing to obtain the corresponding dimensional expansion features.
[0076] In this embodiment, dimensionality augmentation refers to 3D convolution with a kernel size of 1×1×1. This represents the input features of the convolutional module. Indicates to Features after tensor reshaping, W 1×1×1 (·) indicates a 1×1×1 convolution process, F unsqueese Let F represent the features after dimensionality expansion. unsqueese It can be represented as: Among them, for Performing dimensional expansion processing is equivalent to... Dimensional upscaling is performed to accommodate 3D convolution operations, and then the new dimension is expanded to R dimensions through 3D convolution.
[0077] S704 uses two parallel asymmetric 3D convolutions to perform convolution processing on the corresponding dimension-expanded features, resulting in the corresponding convolution-processed third and fourth features.
[0078] S706 performs pixel-by-pixel summation on the third and fourth features after convolution processing to obtain the corresponding pixel-summed features.
[0079] In this embodiment, spatial and spectral correlations can be found by using two parallel asymmetric 3D convolutions. Therefore, convolution processing using asymmetric 3D convolutions processes features through spatial and spectral dimensions respectively. Moreover, using two asymmetric 3D convolutions simultaneously can also reduce the computational cost and number of parameters of the network.
[0080] For example, if F unsqueese W represents the features after dimensionality expansion. 1×k×k (·) and W k×1×1 (·) represent asymmetric 3D convolutions with kernel sizes of 1×k×k and k×1×1, respectively. find If F represents the feature after pixel summation, then... find It can be represented as: F find =W 1×k×k (F unsqueese )+W k×1×1 (F unsqueese ); where the specific value of k can be set according to the actual application scenario, and this embodiment does not impose a specific limitation.
[0081] S708 performs dimensionality reduction on the features obtained by adding the corresponding pixels, and then performs tensor reshaping on the dimensionality-reduced features to obtain the features after tensor reshaping.
[0082] S710 performs pixel-wise addition of the tensor-reshaped features and the corresponding convolutional module input features to obtain the corresponding convolutional module output features.
[0083] In this embodiment, the dimensionality reduction of the features obtained by summing the corresponding pixels can be performed using a 3D convolution with a kernel size of 1×1×1. Then, local residual connections are used to obtain the output features of the convolution module. Where W... 1×1×1 (·) indicates a 3D convolution with a kernel size of 1×1×1, F find F represents the feature after pixel summation. find ' represents the feature after adding pixels together, then F find '=W 1×1×1 (F find Furthermore, if F find "Indicates the characteristics after tensor reshaping." This represents the input features of the convolutional module. If the convolutional module outputs features, then... It can be represented as:
[0084] Combination Figure 7 The content shown, in one embodiment, is as follows: Figure 8 The diagram illustrates a method for processing input features of a 3D convolutional module to obtain the corresponding output features. Figure 8 The content shown can be used as a reference. Figure 7 The content shown is for illustrative purposes only and will not be repeated here.
[0085] exist Figure 2 Based on the content shown, in one embodiment, such as Figure 9 The diagram illustrates a process for obtaining a high-resolution hyperspectral image corresponding to a low-resolution hyperspectral image based on the spectral features of the low-resolution hyperspectral image. This method is then applied to… Figure 1 Taking server 104 as an example, the explanation may include the following steps:
[0086] S902, perform the first convolution process on the spectral features of the low-resolution hyperspectral image to obtain the spectral features of the processed low-resolution hyperspectral image.
[0087] The first convolution process is a 3×3 convolution process, if F branch W represents the spectral features of a low-resolution hyperspectral image. 3×3 (·) indicates a 3×3 convolution process, which can extract global shallow features F through a preprocessing step. pre Preprocessing is essentially 3×3 convolution, then F pre It can be represented as Fpre =W 3×3 (F branch ).
[0088] S904 processes the global shallow features of the low-resolution hyperspectral image based on the processing network, and adds the global shallow features with the processed global shallow features pixel by pixel to obtain the global deep features of the low-resolution hyperspectral image.
[0089] Where, if F pre The global shallow features of a low-resolution hyperspectral image can be represented by B Transformer models, which can be used to extract the global deep features F. TM Then F TM It can be represented as:
[0090]
[0091] in, F represents the output feature of the b-th Transformer model. TM The output of the entire Transformer model is then used. Finally, sub-pixel convolution is applied to upsample the output features of the entire Transformer model, mapping the features to the magnification factor of the target. This yields the global deep features F of the low-resolution hyperspectral image. primary Then F primary It can be represented as: F primary =f UP (F TM ), f UP (·) is used for amplification processing of subpixel convolutions in the network.
[0092] S906 performs pixel-wise addition of the spectral features of the upsampled preprocessed image and the global deep features to obtain the pixel features of the low-resolution hyperspectral image.
[0093] S908 performs a first convolution process on the pixel features of the low-resolution hyperspectral image to obtain a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image.
[0094] In the descriptions of S906 and S908, the upsampling preprocessed image is obtained by upsampling a low-resolution hyperspectral image and then performing a third convolution on the upsampled low-resolution hyperspectral image. The third convolution is a 1×1 convolution, which is also an interpolation amplification process. The first convolution is a 3×3 convolution operation.
[0095] For example, if f↑(·) is an interpolation amplification process, W 1×1 (·) represents a 1×1 convolution, W3×3 (·) represents a 3×3 convolution; low-resolution hyperspectral images use I0. LR It means that F primary The global deep features of a low-resolution hyperspectral image are represented by I. The corresponding high-resolution hyperspectral image is represented by I. SR Indicate, then I SR It can be represented as: I SR =W 3×3 (F primary + W 1×1 (f↑(I LR ))).
[0096] It should be noted that the S906 uses progressive upsampling processing, which can reduce the huge amount of computation brought about by preprocessing upsampling and solve the problem of blurred reconstructed images caused by insufficient extraction of high-frequency information in postprocessing upsampling.
[0097] Combination Figures 2 to 9 The content shown, in one embodiment, is as follows: Figure 10 As shown, a schematic diagram of a low-resolution hyperspectral image processing method is provided, wherein, Figure 10 The content shown can be adapted to the above description and will not be repeated here.
[0098] It should be noted that, Figure 10 The structure shown is the overall network structure F used to obtain a high-resolution hyperspectral image corresponding to a low-resolution hyperspectral image. Net The overall network structure also includes Figure 4 , Figure 6 as well as Figure 8 The structure shown, where if I LR I represents the input low spatial resolution hyperspectral image. SR Indicates with I LR The corresponding high spatial resolution hyperspectral image, then I SR It can be represented as: I SR =F Net (I LR ).
[0099] Based on the above, it should be noted that the loss functions used to construct the overall network structure include the L1 loss function, the Spectral Angle Mapping (SAM) loss function, and the Spatial-Spectral Total Variation (SSTV) loss function. Specifically, the processing network is constructed, and the network is built... Figure 4 , Figure 6 as well as Figure 8The loss functions used in the structure shown are the SAM loss function and the spatial spectral total variational SSTV loss function.
[0100] Understandably, most super-resolution models use L1 loss function for spatial information constraints. Compared to the Mean Squared Error (MSE) loss function, the L1 loss function converges faster during training and is more sensitive to brightness and color changes in textureless regions of the image. Therefore, in this application, L1 loss is selected as the spatial information constraint for reconstructing hyperspectral images. In addition, SSTV loss function and SAM loss function are used simultaneously as spectral distortion loss functions to reduce spectral distortion of the reconstructed image.
[0101] The L1 loss function can be expressed as: The SSTV loss function can be expressed as: The SAM loss function can be expressed as: Furthermore, the overall loss function can be expressed as: Where N is the number of hyperspectral images used during model training. This represents the Nth high-resolution hyperspectral image. This represents the Nth super-resolution hyperspectral image generated by the model. They represent calculations respectively. It is a function of the horizontal, vertical, and spectral gradients. α and β represent adjustable hyperparameters.
[0102] Based on the above, it can be understood that in this application, the Transformer model, which extracts global spatial features and long-range information, is applied to hyperspectral image super-resolution as the overall network structure. The purpose of this overall network structure is to predict the corresponding high spatial resolution hyperspectral image from a low spatial resolution hyperspectral image using the end-to-end network proposed in this application. Specifically, the structure of the Space-Spectral Prior Network (SSPSR) replaces the feature extraction part with a Transformer model, enhancing the learning ability of global features in the spatial dimension and improving the network's ability to preserve long-range information. To enhance spatial detail representation, and to explore the correlation between spectral and spatial dimensions in hyperspectral images, this application proposes a 3D convolution module. This module exists in the feedforward network portion of the Transformer model, and can extract latent features between spectral and spatial dimensions, improving reconstruction performance. Finally, since spectral distortion can lead to errors in accuracy and precision for advanced computer vision tasks, this application employs L1 loss, SSTV loss, and SAM loss to constrain spatial details and spectral bias, reducing spectral distortion without affecting spatial recovery performance.
[0103] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0104] To better verify the performance of the overall network proposed in this application, the following sections will introduce the performance of the overall network from four aspects: dataset, implementation details, comparison of results on the CAVE dataset, and comparison of results on the Chikusei dataset.
[0105] In this application, the performance of the overall network is validated using the CAVE dataset and the Chikusei dataset, respectively, from daily and remote sensing hyperspectral images. The CAVE dataset was generated by Yasuma et al. at Columbia University using a cooled charge-coupled device (CCD) camera, for example, an Apogee Alta U260 camera, containing 32 daily hyperspectral images with a spatial size of 512×512, spanning 31 spectral bands from 400nm to 700nm. Specifically, 20 images are randomly selected as the training set, and 10% of the data is used as the validation set. The images were used as the test set; the Chikusei dataset consists of hyperspectral images taken by Yokoya et al. of the University of Tokyo using the Headwall Hyperspec-VNIR-C sensor, containing 128 spectral bands from 363nm to 1018nm, with a spatial resolution of 2517×2335; due to the lack of edge information, blurred edges were first removed, retaining a spatial resolution of 2304×2048, and then four non-overlapping images with a spatial resolution of 512×512 were selected as the test set, with 90% of the remaining images used as the training set and 10% as the validation set.
[0106] In this application, the implementation details are as follows: Except for the reconstruction part, all 2D convolutional networks have 256 output channels, and the 3D convolutional networks have 16 output channels. The number of Transformer models N in the processing network is 3. Specifically, the SSPSR grouping method is used to process low-resolution and high-resolution images in groups. For the CAVE dataset, 32 spectral bands are divided into 5 branches, each with 8 spectral channels, and the number of overlapping spectra between groups is 2. For the Chikusei dataset, 128 spectral bands are divided into 21 groups, each with 8 spectral channels, and the number of overlapping spectra is 2. ADAM is used as the optimizer, with an initial learning rate of 1e-4 and 40 epochs of training. The learning rate is halved at the 30th epoch. The experimental environment uses a graphics processing unit (GPU) version of PyTorch and is trained using an RTX 2070 Super GPU.
[0107] In validating the overall network performance, six evaluation metrics were established, including: Root Mean Square Error (RMSE), Peak Signal to Noise Ratio (PSNR), Structural Similarity (SSIM), Cross-Correlation (CC) Spectral Angle Mapping (SAM), and Erreur Relative Global Error (ERGAS). Among these, RMSE, PSNR, and SSIM are commonly used image restoration quality metrics, while CC, SAM, and ERGAS are widely used evaluation metrics in hyperspectral fusion tasks. Lower values for SAM, ERGAS, and RMSE are better, with optimal values of 0, 0, and 0, respectively. For CC, PSNR, and SSIM, values closer to 1, positive infinity, and 1, respectively, are better.
[0108] As shown in Table 1, the average quantitative comparison of five different methods on six evaluation metrics for test images of the CAVE dataset is provided. Table 1 shows the comparison of the bicubic interpolation method (bicubic method) with three deep learning-based single-image hyperspectral algorithms (3DCNN method, GDRRN method, and SSPSR method) and the method proposed in this application (3D-THSR method) on six evaluation metrics on 12 CAVE test images of size 512×512×31. In the table, bold indicates the best performance and underline indicates the second best performance.
[0109] As can be seen from Table 1, regardless of whether the magnification is 4x or 8x, the overall network proposed in this application achieves the best results in most evaluation metrics. In particular, at a magnification of 4x, the PSNR value of the overall network proposed in this application is 0.63dB higher than that of the SSPSR method, and the SAM value of the overall network is also 0.06 lower than that of the SSPSR method. Therefore, it can be shown that the Transformer model in the overall network proposed in this application extracts more high-frequency information when extracting global features, which plays a certain role in the spatial reconstruction of hyperspectral images. The addition of SAM loss can also effectively reduce the spectral angle matching value.
[0110] Table 1
[0111]
[0112] Among them, the bicubic method is a bicubic linear interpolation method; the 3D method refers to a method based on a three-dimensional fully convolutional neural network, which comes from the foreign literature titled "Hyperspectral Image Spatial Super-Resolution via 3D Full Convolutional Neural Network"; the GDRRN method refers to a method based on a grouped deep recursive residual network, which comes from the foreign literature titled "Single Hyperspectral Image Super-Resolution with Grouped Deep Recursive Residual Network"; the SSPSR method refers to a method based on a spatial spectral prior super-resolution network, which comes from the foreign literature titled "Learning Spatial-Spectral Prior for Super-Resolution of Hyperspectral Imagery"; and the 3D-THSR method refers to a method based on a super-resolution network with 3D convolution and Transformer structures, which is also the method proposed in this application.
[0113] like Figure 11 and Figure 12 As shown, the spatial image details and error maps of different methods in the CAVE dataset test images at magnifications of 4x and 8x are presented. Visually, it can be seen that the overall network proposed in this application... Figure 11 Compared with other methods, the proposed overall network can recover more high-frequency spatial details and can correctly recover the text below the color card. In contrast, the SSPSR method lost details in the recovery of the letters "m" and "e", and other methods failed to recover the text information correctly. Figure 12 The overall network proposed in this application can also more clearly restore the eye outline of the plush toy. Therefore, it can be shown that the overall network proposed in this application uses Transformer to extract global features, which improves the restoration of high-frequency spatial details. In addition, the overall network proposed in this application can also achieve lower spectral error. In comparison, the bicubic method and 3D method produce higher errors in high-frequency information, i.e., the outline.
[0114] As shown in Table 2, the average quantitative comparison of five different methods on six evaluation metrics for test images of the Chikusei dataset is provided. Table 2 shows the comparison of the bicubic interpolation method (bicubic method) with three deep learning-based single-image hyperspectral algorithms (3DCNN method, GDRRN method, SSPSR method) and the method proposed in this application (3D-THSR method) on six evaluation metrics on 12 Chikusei test images of size 512×512×128. Bold text indicates the best performance, and underlined text indicates the second best performance.
[0115] As shown in Table 2, the overall network proposed in this application achieves the best results at a magnification of 4x. However, the overall network proposed in this application does not achieve the best results among all evaluation metrics at a magnification of 8x. This may be because the Chikusei dataset is a remote sensing image, which is easily affected by airborne water vapor or particulate matter, resulting in noise. This makes it impossible for the 3D convolutional network to correctly extract spectral spatial correlation features, leading to a poor SAM index. Furthermore, at a magnification of 8x, the input low-resolution image loses too many high-frequency details, failing to leverage the role of the Transformer module in the overall network proposed in this application in extracting global information, resulting in a slightly lower SSIM index than the SSPSR method.
[0116] Table 2
[0117]
[0118]
[0119] like Figure 13 and Figure 14 As shown, the spatial image details and error maps at the 60th and 80th spectral bands in the test results using the Chikusei dataset for different methods at magnifications of 4x and 8x are displayed. The spatial detail maps show that... Figure 13 In this application, the overall network display of buildings presented by the present application exhibits higher sharpness and more detail compared to other methods; Figure 14 In the overall network proposed in this application, the outline of the farmland is clearer than that of other methods. In the error map, although the overall network proposed in this application is not much different from other methods, it has lower spectral error on some edges and outlines.
[0120] In the results analysis, this application will use ablation experiments of the Transformer model and 3D convolutional modules, ablation experiments of the number of 3D convolutional channels, and ablation experiments of the loss function to discuss the effectiveness of the overall network proposed in this application.
[0121] In this application, the impact of the Transformer module and the 3D convolution module in the overall network on the hyperspectral super-resolution reconstruction effect is discussed, ensuring the consistency of the training environment and loss function. The consistency of the loss function refers to the use of L1 loss function, SSTV loss function and SVM loss function.
[0122] Table 3 shows a comparison of the performance of the Transformer model and the 3D convolution module in the ablation study on the CAVE dataset at a magnification of 4x. The performance of w / o Transformer & 3Dconv is significantly better than that of the SSPSR method, which proves that the added Transformer model can effectively extract global feature information through the self-attention module and improve the reconstruction effect. The performance of w / o 3Dconv is further improved on the w / o Transformer & 3Dconv model, which proves that 3D convolution can further extract the potential information between the spectrum and space on the basis of the original network structure, and further improve the super-resolution reconstruction effect.
[0123] Table 3
[0124]
[0125] This application discusses the number of channels in the augmented dimension R of the proposed 3D convolution module, as shown in Table 4. It provides a comparison of the ablation performance of the 3D convolution layer number in the CAVE dataset at a magnification of 4x, evaluating the effectiveness on evaluation metrics. The experiments included 1, 8, 16, and 24 channels. The experiments with 1, 8, and 16 channels show that as the number of channels in the 3D convolution augmented dimension increases, more spectral and spatial depth information is extracted, resulting in better reconstruction. However, the performance decreases with 24 channels. This suggests that increasing the number of channels in the augmented dimension leads to increased redundant features, reducing reconstruction effectiveness. Furthermore, increasing the number of channels in the augmented dimension also increases the number of parameters and computational cost, reducing the algorithm's efficiency. Therefore, in this application, the number of channels in the 3D convolution augmented dimension R can be set to 16.
[0126] Table 4
[0127]
[0128] This application also discusses the effectiveness of L1 loss function, SSTV loss function, and SAM loss function in the training results of the proposed model. Using L1 loss function as a benchmark, both the combination of L1 loss and SAM loss, and the combination of L1 loss and SSTV loss, improve the reconstruction effect. Adding SSTV loss to L1 loss improves the reconstruction effect more than adding SAM loss. Using L1 loss, SSTV loss, and SAM loss together, except for a slight decrease in the SAM loss metric, further improves other metrics. It can be inferred that the use of the two loss functions targeting spectral distortion further constrains spatial information, but at the cost of a slight increase in spectral distortion. Based on the ablation experiments of the above loss functions, it is recommended to choose to use L1 loss, SSTV loss, and SAM loss together, and set the adjustable hyperparameters α and β to 10. -3 .
[0129] In summary, the image processing method proposed in this application can be applied to processing low spatial resolution hyperspectral images, reducing the overhead of replacing them with high spatial resolution hyperspectral optical sensors. Furthermore, the proposed method is a single-image hyperspectral super-resolution reconstruction method based on the Transformer model and 3D convolution, belonging to low-level computer vision tasks. It enhances high-level computer vision tasks such as hyperspectral target detection and target recognition. Specifically, in this application, the 3D-THSR single-image hyperspectral reconstruction algorithm based on the Transformer model improves spatial high-frequency detail learning by extracting global spatial features. By adding a 3D convolution module to the feedforward propagation part of the Transformer model, it can be used to explore potential features between spatial and spectral information, improving reconstruction results. Simultaneously, this application uses three loss functions to constrain spatial and spectral information respectively. Compared with existing single-image hyperspectral super-resolution algorithms, the proposed method not only improves spatial high-frequency detail information but also reduces spectral errors.
[0130] Based on the same inventive concept, this application also provides a low-resolution hyperspectral image processing apparatus for implementing the low-resolution hyperspectral image processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more image processing apparatus embodiments provided below can be found in the limitations of the low-resolution hyperspectral image processing method described above, and will not be repeated here.
[0131] In one embodiment, such as Figure 15As shown, a low-resolution hyperspectral image processing apparatus is provided, including: an acquisition module 1502, a feature determination module 1504, and a processing module 1506, wherein: the acquisition module 1502 is used to acquire shallow features of each group of images of the low-resolution hyperspectral image; each group of images is obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each group of images are obtained by performing a first convolution process on each group of images respectively; the feature determination module 1504 is used to process the shallow features of each group of images respectively based on the processing network, and perform pixel-wise addition of the shallow features of each group of images with the processed shallow features of each group of images. The global deep features of each group of images are obtained; the processing network includes multiple self-attention mechanism models with the same structure, and the output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to it; the processing module 1506 is used to obtain a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image based on the spectral features of the low-resolution hyperspectral image; the spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing a first convolution processing and cascading processing on the global deep features of each group of images after the sub-pixel convolution processing.
[0132] Each module in the aforementioned low-resolution hyperspectral image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0133] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 16 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores low-resolution hyperspectral images. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a low-resolution hyperspectral image processing method.
[0134] Those skilled in the art will understand that Figure 16The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0135] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, implements the steps of the methods described in the above embodiments. In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when executed by a processor, the computer program implements the steps of the methods described in the above embodiments. In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps of the methods described in the above embodiments.
[0136] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited thereto.
[0137] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A low-resolution hyperspectral image processing method, characterized in that, include: Obtain shallow features from grouped images of low-resolution hyperspectral images; Each of the grouped images is obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each of the grouped images are obtained by performing a first convolution process on each of the grouped images respectively; The shallow features of each grouped image are determined as the input to the first self-attention mechanism model in the processing network, and the input is subjected to layer normalization to obtain the corresponding first feature after layer normalization. The first feature after the corresponding layer normalization is subjected to a second convolution process to obtain the first feature after the second convolution process. The first feature after the corresponding second convolution is subjected to channel segmentation to obtain the corresponding query feature, key feature and value feature after channel segmentation. The query features, key features, and value features after the corresponding channel segmentation are subjected to tensor reshaping to obtain the corresponding reshaped query features, key features, and value features; The corresponding reshaped key features are matrix transposed, and the transposed key features are multiplied by the corresponding reshaped query features. The features after matrix multiplication are then processed by the softmax activation function to obtain the features processed by the softmax activation function. The feature processed by the corresponding softmax activation function is multiplied by the corresponding reshaped value feature, and the feature after matrix multiplication is reshaped by tensor to obtain the corresponding reshaped first feature. The first feature after reshaping is subjected to a third convolution, and the first feature after the third convolution is added pixel by pixel to the input of the first self-attention mechanism model in the processing network to obtain the intermediate feature of the first self-attention mechanism model in the processing network. Based on the corresponding intermediate features, the input features of the second self-attention mechanism model in the corresponding processing network are obtained; The output features of the last self-attention mechanism model in the corresponding processing network are determined as the shallow features of each grouped image after processing. The shallow features of each grouped image are pixel-wise added to the shallow features of the processed grouped images to obtain the global deep features of each grouped image; the processing network includes multiple self-attention mechanism models, the multiple self-attention mechanism models have the same structure, and the output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to each self-attention mechanism model; Based on the spectral features of the low-resolution hyperspectral image, a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image is obtained; the spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing the first convolution processing and cascading processing on the global deep features of each group of images after the sub-pixel convolution processing.
2. The method according to claim 1, characterized in that, The step of obtaining the input features of the second self-attention mechanism model in the processing network based on the corresponding intermediate features includes: The corresponding intermediate features are subjected to layer normalization to obtain the corresponding layer-normalized second features; The second feature after the corresponding layer normalization is subjected to the second convolution process to obtain the second feature after the second convolution process. The second feature after the second convolution is processed into layers to obtain the corresponding layered convolution module input features and gated branch input features. Based on the 3D convolutional module, the corresponding input features of the convolutional module are processed to obtain the corresponding output features of the convolutional module; The corresponding gated branch input features are processed by the GELU activation function to obtain the corresponding gated branch output features; The corresponding output features of the convolution module and the corresponding output features of the gated branch are multiplied by pixels to obtain the corresponding multiplied features. The features obtained by multiplying the corresponding pixels are then subjected to the third convolution process, and the features obtained by the third convolution process are then added to the corresponding intermediate features to obtain the input features of the second self-attention mechanism model in the processing network.
3. The method according to claim 2, characterized in that, The process of processing the input features of the 3D convolutional module to obtain the corresponding output features includes: The corresponding convolutional module input features are subjected to tensor reshaping processing, and the corresponding tensor reshaping features are subjected to dimension expansion processing to obtain the corresponding dimension-expanded features. Based on two parallel asymmetric 3D convolutions, the features after corresponding dimension expansion are convolved to obtain the corresponding convolutional third and fourth features. The third and fourth features after convolution are pixel-wise summed to obtain the corresponding pixel-summed features. The feature obtained by adding the corresponding pixels is subjected to dimensionality reduction, and the feature obtained by dimensionality reduction is subjected to tensor reshaping to obtain the feature after tensor reshaping. The features after tensor reshaping are pixel-wise added to the corresponding input features of the convolutional module to obtain the corresponding output features of the convolutional module.
4. The method according to claim 1, characterized in that, The step of obtaining a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image based on the spectral features of the low-resolution hyperspectral image includes: The spectral features of the low-resolution hyperspectral image are subjected to the first convolution process to obtain the global shallow features of the low-resolution hyperspectral image. The global shallow features of the low-resolution hyperspectral image are processed based on the processing network, and the global shallow features are added pixel by pixel to the processed global shallow features to obtain the global deep features of the low-resolution hyperspectral image. The spectral features of the upsampled preprocessed image are pixel-wise added to the global deep features to obtain the pixel features of the low-resolution hyperspectral image; the upsampled preprocessed image is obtained by upsampling the low-resolution hyperspectral image and then performing a third convolution on the upsampled low-resolution hyperspectral image. The pixel features of the low-resolution hyperspectral image are subjected to the first convolution process to obtain a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image.
5. The method according to any one of claims 1 to 4, characterized in that, The loss functions used to construct the processing network include the L1 loss function, the spectral angle matching SAM loss function, and the spatial spectral total variation SSTV loss function.
6. A low-resolution hyperspectral image processing device, characterized in that, The device includes: The acquisition module is used to acquire shallow features of grouped images of a low-resolution hyperspectral image; each grouped image is obtained by grouping the low-resolution hyperspectral image based on the spectral number of the low-resolution hyperspectral image, and the shallow features of each grouped image are obtained by performing a first convolution process on each grouped image respectively. The feature determination module is used to determine the shallow features of each grouped image as the input of the first self-attention mechanism model in the processing network, and perform layer normalization processing on the input to obtain the corresponding layer-normalized first feature; perform a second convolution processing on the corresponding layer-normalized first feature to obtain the corresponding second convolution processing first feature; perform channel segmentation processing on the corresponding second convolution processing first feature to obtain the corresponding channel-segmented query feature, key feature, and value feature; perform tensor reshaping processing on the corresponding channel-segmented query feature, key feature, and value feature to obtain the corresponding reshaped query feature, key feature, and value feature; perform matrix transpose processing on the corresponding reshaped key feature, and perform matrix multiplication processing on the transposed key feature and the corresponding reshaped query feature, and process the matrix multiplied feature with a softmax activation function to obtain the corresponding softmax activation function processed feature; and perform tensor reshaping processing on the corresponding reshaped query feature and the corresponding key feature and value feature. The value features are subjected to matrix multiplication, and the features after matrix multiplication are subjected to tensor reshaping to obtain the corresponding reshaped first features; the corresponding reshaped first features are subjected to third convolution, and the first features after third convolution are added pixel by pixel to the input of the first self-attention mechanism model in the processing network to obtain the corresponding intermediate features of the first self-attention mechanism model in the processing network; based on the corresponding intermediate features, the input features of the corresponding second self-attention mechanism model in the processing network are obtained; the output features of the corresponding last self-attention mechanism model in the processing network are determined as the shallow features of each grouped image after processing; the shallow features of each grouped image are added pixel by pixel to the shallow features of each grouped image after processing to obtain the global deep features of each grouped image; the processing network includes multiple self-attention mechanism models, the multiple self-attention mechanism models have the same structure, and the output of each self-attention mechanism model is the input of the next self-attention mechanism model connected to each self-attention mechanism model; The processing module is used to obtain a high-resolution hyperspectral image corresponding to the low-resolution hyperspectral image based on the spectral features of the low-resolution hyperspectral image; the spectral features of the low-resolution hyperspectral image are obtained by performing sub-pixel convolution processing on the global deep features of each group of images, and then performing the first convolution processing and cascading processing on the global deep features of each group of images after the sub-pixel convolution processing.
7. The apparatus according to claim 6, characterized in that, The feature determination module is further configured to: The corresponding intermediate features are subjected to layer normalization to obtain the corresponding layer-normalized second features; The second feature after the corresponding layer normalization is subjected to the second convolution process to obtain the second feature after the second convolution process. The second feature after the second convolution is processed into layers to obtain the corresponding layered convolution module input features and gated branch input features. Based on the 3D convolutional module, the corresponding input features of the convolutional module are processed to obtain the corresponding output features of the convolutional module; The corresponding gated branch input features are processed by the GELU activation function to obtain the corresponding gated branch output features; The corresponding output features of the convolution module and the corresponding output features of the gated branch are multiplied by pixels to obtain the corresponding multiplied features. The features obtained by multiplying the corresponding pixels are then subjected to the third convolution process, and the features obtained by the third convolution process are then added to the corresponding intermediate features to obtain the input features of the second self-attention mechanism model in the processing network.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.