Hyperspectral Image Super-Resolution Reconstruction Method Based on Hybrid Neural Network
Through a super-resolution reconstruction method based on hybrid neural networks, combined with hypergraph learning and wavelet transformation, the problem of low spatial resolution of hyperspectral images is solved, higher quality image reconstruction is achieved, and application potential is improved.
Patent Information
- Application Number
- CN202510403092.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Due to the physical limitations of imaging spectrometers, hyperspectral images have low spatial resolution, which limits their application potential in tasks such as refined species classification and abnormal detection.
A remote sensing hyperspectral image super-resolution reconstruction method based on hybrid neural network is adopted, combining hypergraph learning, wavelet transformation and spectral hybrid analysis to characterize and utilize spectral information and spatial information, and feature fusion is performed through sensitive band extraction branches to improve the reconstruction effect.
By improving the spatial resolution of the image, the capture ability of subject structure and detail change characteristics in the image is enhanced, and the reconstruction effect is improved, especially in terms of spectral fidelity and detail reconstruction.
Smart Images

Figure CN119904359B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing hyperspectral images, and particularly relates to a method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network. Background Art
[0002] Benefiting from the ability to capture rich spectral signals in the observed scene, hyperspectral images can provide more accurate guidance for image interpretation. With the upgrade of hardware technology, hyperspectral images have gradually become a very important information source in remote sensing information processing technology. However, due to the physical limitations of imaging spectrometers, hyperspectral images usually face the problem of low spatial resolution during the acquisition process. The low spatial resolution greatly limits its application potential, especially for tasks such as fine species classification and anomaly detection. The image super-resolution technology can improve the spatial resolution of hyperspectral images at a relatively low cost, effectively broadening the application scope of hyperspectral images.
[0003] The hyperspectral image super-resolution technology aims to reconstruct high-quality images with higher spatial resolution from existing low-resolution images through algorithmic means. While pursuing more fine-grained spatial details, it also tries to retain spectral information as much as possible. Compared with simple and direct interpolation algorithms, hyperspectral image super-resolution reconstruction needs to consider both spatial information and spectral signals during implementation, which increases the challenge of this task. For the early research in this field, scholars focused on traditional algorithms. Wavelet transform, maximum a posteriori estimation, and mixed pixel analysis are among the hot technical routes. With the development and innovation of machine learning in the field of image processing, the method for hyperspectral image super-resolution reconstruction based on deep learning networks has become a research hotspot in this field.
[0004] Compared with traditional methods, the advantage of deep learning is that it can automatically learn complex patterns in images by leveraging a large amount of training data, usually with strong performance and generalization ability. During the reconstruction process, the rich spectral information carried in the image must be considered. The 3D convolution calculation mode has unique advantages for extracting spectral features, so it has been introduced into this field by some scholars. However, 3D convolution also brings a high computational burden, which poses certain restrictions on the design of the model. The local features in remote sensing images have obvious cross-regional and cross-scale reproducibility, so the extraction of global information seems to be very crucial. With the emergence of Transformer in multiple vision tasks, the excellent performance of the self-attention mechanism has gradually attracted attention. The self-attention mechanism associates pixels at different positions within the global range and autonomously assigns weights. However, the learning process of Transformer is all completed by the network itself, lacking the guidance of prior information and learning focus, which limits the learning efficiency and task orientation of the network.
[0005] In recent years, hypergraph learning has been gradually introduced into the field of image processing as a learning mode with great potential. A hypergraph is a more flexible and complex data representation format that can connect feature points with dependencies within the global scope based on specific prior information. This means that by mapping node information into a more explicit and task-oriented feature space, guidance and tendency can be imposed on the learning process of the network. Remote sensing image super-resolution reconstruction usually corresponds to a large observation area, and complex components are highly coupled in a single pixel, which makes it very challenging to represent the internal information of the image. The complete representation and full utilization of features in multiple dimensions are the keys to achieving accurate image reconstruction. The advantage of this feature space transformation of hypergraphs is particularly suitable for processing high-dimensional and complex data such as hyperspectral images and can make up for the disadvantages of the above-mentioned convolutional neural network and Transformer architecture. Summary of the Invention
[0006] In view of this, the present invention aims to provide a method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network. Taking hypergraph learning as the benchmark learning mode and wavelet transform and spectral mixture analysis as data analysis methods, while representing and utilizing spectral information and spatial information, hypergraph learning realizes the capture of long-range correlations between pixels, enhancing the richness and breadth in the network learning process.
[0007] To achieve the above object, the technical solution of the present invention is realized as follows:
[0008] A method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network includes:
[0009] S1: Obtain a remote sensing hyperspectral image dataset and preprocess the dataset to obtain a training set;
[0010] S2: Construct a hybrid neural network, and use the training set obtained in step S1 to train the hybrid neural network to obtain a hybrid network model;
[0011] The hybrid neural network includes: a spectral hypergraph branch for extracting spectral features in the input image; a spatial hypergraph branch for extracting spatial features in the input image; a semantic hypergraph branch for extracting semantic features in the input image; a sensitive band extraction branch for fusing spectral features, spatial features, and semantic features, and using the fused features to extract high-quality semantic features of the input image; a restoration and reconstruction branch for obtaining a reconstructed remote sensing hyperspectral image based on the preliminary super-resolution image of the input image, in combination with spectral features, spatial features, semantic features, and high-quality semantic features;
[0012] S3: Input the low-resolution remote sensing hyperspectral image to be reconstructed into the hybrid network model obtained in step S2 to obtain the corresponding reconstructed remote sensing hyperspectral image.
[0013] Further, in step S1, the preprocessing process of the dataset includes: splitting, augmenting, and degrading the dataset to obtain a low-resolution input image.
[0014] Further, in the spectral hypergraph branch of step S2: perform a one-dimensional wavelet transform on the input image to obtain approximation coefficients and detail coefficients; construct a low-frequency feature hypergraph and a high-frequency feature hypergraph based on the approximation coefficients and detail coefficients, and then respectively obtain a low-frequency feature hypergraph operator and a high-frequency feature hypergraph operator by transforming the low-frequency feature hypergraph and the high-frequency feature hypergraph through the following formula:
[0015] ;
[0016] ;
[0017] where, represents the low-frequency feature hypergraph operator, represents the high-frequency feature hypergraph operator, represents the low-frequency feature hypergraph, represents the high-frequency feature hypergraph, and respectively represent the diagonal matrices of the hyperedge degree and the vertex degree, and E is initialized as the identity matrix;
[0018] Take the input image as the input feature, input it into the spectral hypergraph mutual attention module together with the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator, and the output feature is processed by the spectral hypergraph mutual attention module no less than 1 time with the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator to obtain the spectral feature.
[0019] Further, in the spectral hypergraph mutual attention module:
[0020] Multiply the input feature with the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator respectively to obtain a low-frequency spectral feature map and a high-frequency spectral feature map, that is:
[0021] ;
[0022] ;
[0023] where, , n represents the number of spectral hypergraph mutual attention modules, N represents the total number of spectral hypergraph mutual attention modules, represents the spectral feature output by the previous spectral hypergraph mutual attention module, represents that the feature input into the spectral hypergraph mutual attention module is the input image, represents the low-frequency spectral feature map, represents the high-frequency spectral feature map, and represent the weight coefficients for calculating the low-frequency spectral feature map and the high-frequency spectral feature map;
[0024] Extract the query matrix from the low-frequency spectral feature map, extract the key matrix and the value matrix from the high-frequency spectral feature map, and then perform the multi-head mutual attention mechanism operation on the query matrix, the key matrix, and the value matrix to obtain the spectral features output by the current spectral hypergraph mutual attention module, that is:
[0025] ;
[0026] Among them, represents the spectral features output by the current spectral hypergraph mutual attention module, represents the query matrix, represents the key matrix, represents the value matrix, represents the multi-head mutual attention mechanism operation.
[0027] Furthermore, in the spatial hypergraph branch in step S2:
[0028] Reduce the dimension of the input image, and then perform a two-dimensional wavelet transform on the dimension-reduced image to obtain the low-frequency component and the high-frequency component of the input image, that is:
[0029] ;
[0030] Among them, represents the low-frequency component, represents the high-frequency components in the horizontal, vertical, and diagonal directions of the input image, represents the two-dimensional wavelet transform, represents the image retained after the input image is dimension-reduced;
[0031] Perform two-fold bicubic interpolation on the low-frequency component and the high-frequency component, expand the extracted frequency features, and construct the low-frequency spatial information hypergraph and the high-frequency spatial information hypergraph of the input image through the following formula:
[0032] ;
[0033] Among them, represents the low-frequency spatial information hypergraph, represents the high-frequency spatial information hypergraph, represents the complete function for constructing the low-frequency spatial information hypergraph and the high-frequency spatial information hypergraph;
[0034] Add the corresponding elements of the low-frequency spatial information hypergraph and the high-frequency spatial information hypergraph to obtain the spatial information hypergraph of the input image, and convert the spatial information hypergraph into a spatial information hypergraph operator through the following formula:
[0035] ;
[0036] Among them, represents the spatial information hypergraph operator, represents the spatial information hypergraph, and respectively represent the diagonal matrices of hyperedge degree and vertex degree, and E is initialized as the identity matrix;
[0037] Taking the input image as the input feature and inputting it into the spatial hypergraph self-attention module together with the spatial information hypergraph operator, the output feature is then processed by the spatial information hypergraph operator through no less than 1 spatial hypergraph self-attention module to obtain the spatial feature.
[0038] Furthermore, in the spatial hypergraph self-attention module:
[0039] Multiplying the input feature by the spatial information hypergraph operator to obtain the spatial information feature map, that is:
[0040] ;
[0041] Among them, , m represents the number of spatial hypergraph self-attention modules, M represents the total number of spatial hypergraph self-attention modules, represents the spatial feature output by the previous spatial hypergraph self-attention module, represents that the feature input into the spatial hypergraph self-attention module is the input image, represents the spatial information feature map, represents the weight coefficient for calculating the spatial information feature map;
[0042] Extracting the query matrix, key matrix, and value matrix from the spatial information feature map, and then performing the multi-head self-attention mechanism operation on the query matrix, key matrix, and value matrix to obtain the spatial feature output by the current spatial hypergraph self-attention module, that is:
[0043] ;
[0044] Among them, represents the spatial feature output by the current spatial hypergraph self-attention module, represents the query matrix, represents the key matrix, represents the value matrix, represents the multi-head self-attention mechanism operation.
[0045] Furthermore, in the semantic hypergraph branch in step S2:
[0046] Parsing the abundance matrix from the input image by the non-negative matrix factorization algorithm;
[0047] According to the abundance matrix, construct the semantic hypergraph of the input image through the following formula:
[0048] ;
[0049] where, represents the semantic hypergraph, represents the abundance matrix, represents the complete function for constructing the semantic hypergraph;
[0050] Convert the semantic hypergraph to obtain the semantic hypergraph operator through the following formula:
[0051] ;
[0052] where, represents the semantic hypergraph operator, and respectively represent the diagonal matrices of the hyperedge degree and vertex degree, and E is initialized as the identity matrix;
[0053] Take the input image as the input feature, and input it together with the semantic hypergraph operator into the semantic hyper Figure 3 D module. The output feature is then processed by the semantic hyper Figure 3 D module no less than 1 time to obtain the semantic feature.
[0054] Furthermore, in the semantic hyper Figure 3 D module:
[0055] Multiply the input feature by the semantic hypergraph operator to obtain the semantic information feature map, that is:
[0056] ;
[0057] where, , o represents the number of semantic hyper Figure 3 D modules, represents the total number of semantic hyper Figure 3 D modules, represents the semantic feature output by the previous semantic hyper Figure 3 D module, represents that the feature input to the input semantic hyper Figure 3 D module is the input image, represents the semantic information feature map, represents the weight coefficient for calculating the semantic information feature map;
[0058] Then perform three-dimensional convolution processing on the semantic information feature map to obtain the semantic feature output by the current semantic hyper Figure 3 D module.
[0059] Further, the sensitive band extraction branch in step S2 includes a weight acquisition module and a reference residual module; specifically:
[0060] In the weight acquisition module, pixel value statistics are performed on each band in the spectral feature, spatial feature, and semantic feature to estimate the pixel value probability distribution of each band; the entropy value of each band in each feature is calculated through the maximum entropy formula, and the band with the maximum entropy is selected as the sensitive band of each feature.
[0061] The reference residual module includes no less than 3 cascaded residual blocks. Between two adjacent residual blocks, the sensitive bands of the three features are used as weights and multiplied by the corresponding elements of the input features of the previous residual block. The multiplied features are then added to the input features of the subsequent residual block to obtain the output features; each residual block includes no less than 2 cascaded convolutional layers. After the convolutional operation of the convolutional layer on the input features of the current residual block, the features are added to the corresponding elements of the input features of the current residual block to obtain the output features; the output features obtained by the last residual block are the high-quality semantic features.
[0062] Further, in the restoration and reconstruction branch in step S2,
[0063] The corresponding elements of the spectral feature, spatial feature, semantic feature, and high-quality semantic feature are added together. The added features are sequentially subjected to no less than 1 convolutional operation and transposed convolutional operation. The obtained features are added to the corresponding elements of the preliminary super-resolution image obtained by upsampling the input image. After that, the added features are sequentially subjected to no less than 1 convolutional operation to obtain the reconstructed remote sensing hyperspectral image.
[0064] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0065] (1) In the remote sensing hyperspectral image super-resolution reconstruction method based on a hybrid neural network of the present invention, a remote sensing hyperspectral image super-resolution hybrid network model is provided, which uses hypergraph learning as the benchmark learning mode and wavelet transform and spectral mixture analysis as data analysis methods. Based on the high and low frequency features in space and spectrum, the main structure and detail change characteristics in the image are captured; combined with the sensitive band extraction branch, the spectral feature, spatial feature, and semantic feature are fused to strengthen the propagation and reconstruction of semantic information, thereby improving the reconstruction effect.
[0066] (2) In the method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network according to the present invention, the spectral hypergraph branch combines one-dimensional wavelet transform, hypergraph learning, and mutual attention mechanism to better represent and utilize spectral information; in the spectral hypergraph mutual attention module, the frequency characteristics of each signal node are considered and processed, which can effectively improve the fidelity of the reconstructed spectrum; within the same scene, due to the strong cross-regional reproducibility of the target, and the hypergraph represents the dependencies between nodes within the global scope, hypergraph learning realizes the capture of long-range correlations between pixels, thereby enhancing the richness and breadth in the network learning process;
[0067] (3) In the method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network according to the present invention, the spatial hypergraph branch combines two-dimensional wavelet transform, hypergraph learning, and self-attention mechanism to achieve precise reconstruction of image structure information and detailed texture; in the spatial hypergraph self-attention module, hypergraph learning based on the main structure and spatial features of the input image can efficiently utilize similar features that repeatedly appear in the image, realizing mutual enhancement between feature points and precise reconstruction of spatial information;
[0068] (4) In the method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network according to the present invention, the semantic hypergraph branch combines an unmixing, hypergraph learning, and 3D convolution semantic learning module. By mapping highly coupled complex information into a more explicit high-level semantic space, it effectively reduces the difficulty of information propagation, which is beneficial to more precise detail reconstruction;
[0069] (5) In the method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network according to the present invention, the intervention of sensitive band attention in the sensitive band extraction branch dynamically adjusts the weights of each feature during information fusion to more precisely guide the interaction between spectrum, space, and high-level semantic information. This method ensures the collaborative work of different feature domains and ultimately improves the super-resolution effect. Description of the Drawings
[0070] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0071] Figure 1 is a schematic flow chart of the method for super-resolution reconstruction of remote sensing hyperspectral images based on a hybrid neural network according to the embodiment of the present invention;
[0072] Figure 2 is a schematic network structure diagram of the hybrid neural network according to the embodiment of the present invention;
[0073] Figure 3 Schematic diagram of the spectral hypergraph mutual attention module according to an embodiment of the present invention;
[0074] Figure 4 Schematic diagram of the spatial hypergraph self-attention module according to an embodiment of the present invention;
[0075] Figure 5 For the semantic hyper Figure 3 Schematic diagram of the D module according to an embodiment of the present invention;
[0076] Figure 6 Visual comparison result graph of the reconstruction method according to an embodiment of the present invention and other reconstruction methods on the MDAS dataset;
[0077] Figure 7 Visual comparison result graph of the reconstruction method according to an embodiment of the present invention and other reconstruction methods on the Pavia Centre dataset;
[0078] Figure 8 Visual comparison result graph of the reconstruction method according to an embodiment of the present invention and other reconstruction methods on the Houston dataset. Detailed implementation manners
[0079] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention.
[0080] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0081] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.
[0082] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0083] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.
[0084] As Figures 1 to 2 shown, the super-resolution reconstruction method for remote sensing hyperspectral images based on a hybrid neural network according to an embodiment of the present invention includes:
[0085] S1: Obtain a remote sensing hyperspectral image dataset, and preprocess the dataset to obtain a training set.
[0086] In step S1, the preprocessing process of the dataset includes: segmenting, augmenting, and degrading the dataset to obtain low-resolution input images. In some embodiments, 70% of the image regions in the remote sensing hyperspectral image dataset are selected as the training set, and 30% of the regions are used as the validation set. Patches are randomly selected from each region, and different patch sizes are set according to the magnification factor. Among them, magnifications of 2, 3, and 4 times correspond to patch sizes of 64×64, 96×96, and 128×128 respectively. Each patch is randomly horizontally flipped, rotated at different angles, and scaled at different ratios. Finally, these patches are downsampled to 32×32 low-resolution hyperspectral images according to different scale factors through bicubic downsampling. The low-resolution hyperspectral images are used as the input images of the hybrid neural network, and the corresponding real remote sensing hyperspectral images are used as the output images of the hybrid neural network for subsequent training of the hybrid neural network. In one embodiment, remote sensing hyperspectral images are obtained from the MDAS dataset, the PaviaCentre dataset, and the Houston dataset to form a remote sensing hyperspectral image dataset.
[0087] S2: Construct a hybrid neural network, and use the training set obtained in step S1 to train the hybrid neural network to obtain a hybrid network model.
[0088] As Figure 2As shown in the figure, the hybrid neural network includes a parallel spectral hypergraph branch, a spatial hypergraph branch, a semantic hypergraph branch, and a sensitive band extraction branch, as well as a recovery and reconstruction branch that receives the output features of the three branches. Among them, the spectral hypergraph branch extracts the spectral features in the input image, the spatial hypergraph branch extracts the spatial features in the input image, the semantic hypergraph branch extracts the semantic features in the input image, the sensitive band extraction branch fuses the spectral features, spatial features, and semantic features, and uses the fused features to extract the high-quality semantic features of the input image. The recovery and reconstruction branch is based on the preliminary super-resolution image of the input image, and combines the spectral features, spatial features, semantic features, and high-quality semantic features to obtain the reconstructed remote sensing hyperspectral image.
[0089] Specifically, in the spectral hypergraph branch of step S2: perform one-dimensional wavelet transform on the input image to obtain the approximation coefficient and the detail coefficient, which can be expressed by the following formula:
[0090] ;
[0091] Among them, represents the approximation coefficient, represents the detail coefficient, represents the input image, represents the one-dimensional wavelet transform, H, W, and C respectively represent the height, width, and number of channels of the input image, represents a real number matrix with a size of .
[0092] According to the similarity of the low-frequency and high-frequency characteristics among the spectra, use the approximation coefficient and the detail coefficient to construct the low-frequency feature hypergraph and the high-frequency feature hypergraph. The nodes of the hypergraph are the approximation coefficients or detail coefficients of each spectrum, and each hyperedge represents the similarity between the nodes. The higher the similarity, the higher the weight assigned to the hyperedge. Specifically, each hyperedge contains ten nodes selected by the K-nearest neighbor algorithm based on the Euclidean distance, and the weights of the nodes are obtained by the negative exponential function. The hypergraph construction process can be expressed as:
[0093] ;
[0094] Among them, represents the low-frequency feature hypergraph, represents the high-frequency feature hypergraph, represents the complete function for constructing the low-frequency feature hypergraph and the high-frequency feature hypergraph, , T represents the number of nodes, represents a real number matrix with a size of .
[0095] The low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator are respectively obtained by converting the low-frequency feature hypergraph and the high-frequency feature hypergraph through the following formula:
[0096] ;
[0097] ;
[0098] Among them, represents a low-frequency feature hypergraph operator, represents a high-frequency feature hypergraph operator, and respectively represent the diagonal matrices of hyperedge degree and vertex degree, and E is initialized as the identity matrix.
[0099] Hypergraph learning is performed on the input features and the input image. Specifically: taking the input image as the input feature, it is input into the spectral hypergraph mutual attention module together with the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator. The output feature is then processed by the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator through no less than 1 spectral hypergraph mutual attention module to obtain spectral features. In one embodiment, the number of spectral hypergraph mutual attention modules is 6.
[0100] The structure of the spectral hypergraph mutual attention module is as Figure 3 shown, where: the input feature is multiplied by the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator respectively to obtain a low-frequency spectral feature map and a high-frequency spectral feature map, that is:
[0101] ;
[0102] ;
[0103] Among them, , n represents the number of spectral hypergraph mutual attention modules, N represents the total number of spectral hypergraph mutual attention modules, represents the spectral feature output by the previous spectral hypergraph mutual attention module, represents that the feature input into the spectral hypergraph mutual attention module is the input image, represents the low-frequency spectral feature map, represents the high-frequency spectral feature map, and represent the weight coefficients for calculating the low-frequency spectral feature map and the high-frequency spectral feature map, represents a real number matrix with a size of .
[0104] The cross-attention algorithm is performed on the low-frequency spectral feature map and the high-frequency spectral feature map to better utilize the dependency between high-frequency features and low-frequency structures, that is, a query matrix is extracted from the low-frequency spectral feature map, and a key matrix and a value matrix are extracted from the high-frequency spectral feature map. Then, the query matrix, the key matrix, and the value matrix are subjected to a multi-head cross-attention mechanism operation to obtain the spectral features output by the current spectral hypergraph cross-attention module, that is:
[0105] ;
[0106] Among them, represents the spectral features output by the current spectral hypergraph cross-attention module, represents the query matrix, represents the key matrix, represents the value matrix, represents the multi-head cross-attention mechanism operation.
[0107] To sum up, the feature extraction process of the complete spectral hypergraph cross-attention module can be expressed as:
[0108] ;
[0109] Among them, represents the processing operation of the nth spectral hypergraph cross-attention module. In this module, the frequency features of each signal node are considered and processed, which can effectively improve the fidelity of the reconstructed spectrum. In addition, within the same scene, the target has strong cross-region reproducibility, and the hypergraph represents the dependency relationship between nodes within the global range. Hypergraph learning realizes the capture of long-range correlations between pixels, thereby enhancing the richness and breadth in the network learning process.
[0110] In addition to spectral information, for super-resolution tasks aiming to improve spatial resolution, capturing and utilizing spatial information is also equally important. The low-frequency signals in an image represent features that change slowly in space. The pixel values in this part change relatively smoothly and usually contain the main structure information in the observed scene, such as the brightness distribution of the background; the high-frequency signals in the image represent features that change rapidly in space. The pixel values in this part change relatively rapidly and usually contain the detailed structure information in the observed scene, such as the edges and textures of the target. Therefore, in the spatial hypergraph branch of step S2 of the present invention, a spatial feature extraction module combining two-dimensional wavelet transform, hypergraph learning, and self-attention mechanism is used to achieve precise reconstruction of image structure information and detailed textures.
[0111] Specifically, hyperspectral images usually have hundreds of bands, with significant differences in the quality of each band and containing a lot of redundant information. To simplify the processing process and reduce the computational overhead, in the spatial hypergraph branch, first, the input image is dimensionally reduced, and then the dimensionally reduced image is subjected to two-dimensional wavelet transform to obtain the low-frequency component and high-frequency component of the input image, that is:
[0112] ;
[0113] ;
[0114] Among them, represents the low-frequency component, represents the high-frequency components in the horizontal, vertical, and diagonal directions of the input image, represents the two-dimensional wavelet transform. In the embodiments of the present invention, principal component analysis (PCA) is preferably used to dimensionally reduce the image, k represents the number of channels retained after the image is dimensionally reduced, represents the image retained after the input image is dimensionally reduced, represents a real number matrix of size . The low-frequency component and high-frequency component are magnified by a factor of two using bicubic interpolation to expand the extracted frequency features, and the low-frequency spatial information hypergraph and high-frequency spatial information hypergraph of the input image are constructed through the following formula:
[0115] ;
[0116] ;
[0117] Among them, represents the low-frequency spatial information hypergraph, represents the high-frequency spatial information hypergraph, represents the complete function for constructing the low-frequency spatial information hypergraph and high-frequency spatial information hypergraph;
[0118] The corresponding elements of the low-frequency spatial information hypergraph and high-frequency spatial information hypergraph are added to obtain the spatial information hypergraph of the input image, that is:
[0119] ;
[0120] Among them, represents the spatial information hypergraph;
[0121] The spatial information hypergraph is transformed through the following formula to obtain the spatial information hypergraph operator:
[0122] ;
[0123] Among them, represents the spatial information hypergraph operator.
[0124] Perform hypergraph learning on the input features or input image. Specifically, take the input image as the input feature and input it, together with the spatial information hypergraph operator, into the spatial hypergraph self-attention module. The output feature is then processed by no less than 1 spatial hypergraph self-attention module together with the spatial information hypergraph operator to obtain the spatial feature. In one embodiment, the number of spatial hypergraph self-attention modules is 6.
[0125] The structure of the spatial hypergraph self-attention module is as Figure 4 shown and specifically includes:
[0126] Multiply the input feature by the spatial information hypergraph operator to obtain the spatial information feature map, that is:
[0127] ;
[0128] Where , m represents the number of spatial hypergraph self-attention modules, M represents the total number of spatial hypergraph self-attention modules, represents the spatial feature output by the previous spatial hypergraph self-attention module, represents that the feature input into the spatial hypergraph self-attention module is the input image, represents the spatial information feature map, represents the weight coefficient for calculating the spatial information feature map.
[0129] Extract the query matrix, key matrix, and value matrix from the spatial information feature map, and then perform the multi-head attention mechanism operation on the query matrix, key matrix, and value matrix to obtain the spatial feature output by the current spatial hypergraph self-attention module, that is:
[0130] ;
[0131] Where represents the spatial feature output by the current spatial hypergraph self-attention module, represents the query matrix, represents the key matrix, represents the value matrix, represents the multi-head self-attention mechanism operation.
[0132] To sum up, the feature extraction process of the complete spatial hypergraph self-attention module can be expressed as:
[0133] ;
[0134] Where represents the feature extraction operation of the spatial hypergraph self-attention module.
[0135] In the spatial hypergraph self-attention module, hypergraph learning based on the main structure and detailed texture can efficiently utilize the similar features that repeatedly appear in the image, achieving mutual enhancement between feature points and precise reconstruction of spatial information.
[0136] One pixel point in a hyperspectral image often corresponds to a relatively large observation area, which means that the image contains a large number of mixed pixels. The different targets usually show complex interspersed distributions, and this distribution imbalance and multi-class coupling affect the expression of their respective spatial, spectral, and frequency domain information. The information decoupling of mixed pixels can perform feature decomposition and reconstruction more precisely, reducing the mutual interference of multi-dimensional information in complex environments. Therefore, in step S2 of the present invention, a semantic hypergraph branch based on unmixing, hypergraph learning, and 3D convolution is proposed to complete high-quality semantic information extraction. Specifically, in the semantic hypergraph branch, it is necessary to first parse out the endmembers representing pure substances and the abundance ratios of each endmember in each pixel from the input image. In the embodiments of the present invention, the non-negative matrix factorization (NMF) algorithm based on the gradient descent method is preferably used to estimate the endmember matrix and the abundance matrix, and its basic form can be expressed as:
[0137] ;
[0138] where represents the original input image, represents the abundance matrix of the content of each endmember in the pixel, represents the endmember matrix of the spectral characteristics of pure substances, q represents the number of endmembers, represents a matrix of size , represents a matrix of size . The objective function of NMF is:
[0139] ;
[0140] where represents the Frobenius norm, represents the element in the input image X, represents the matrix element after multiplying the abundance matrix A and the endmember matrix . In addition, the obtained abundance matrix A and the endmember matrix need to satisfy the non-negativity constraint and the sum-to-one constraint, that is:
[0141]
[0142] where, represents the abundance of the jth endmember contained in the ith node.
[0143] Construct a semantic hypergraph of the input image according to the abundance matrix by the following formula:
[0144] ;
[0145] where, represents the semantic hypergraph, represents the complete function for constructing the semantic hypergraph.
[0146] Convert the semantic hypergraph to a semantic hypergraph operator by the following formula:
[0147] ;
[0148] where, represents the semantic hypergraph operator.
[0149] Perform hypergraph learning on the input image or input features, that is, take the input image as the input feature, and input it together with the semantic hypergraph operator into the semantic hyper Figure 3 D module. The output feature is then processed by the semantic hypergraph operator through no less than 1 semantic hyper Figure 3 D module to obtain the semantic feature. In one embodiment, the number of semantic hyper Figure 3 D modules is 6.
[0150] The structure of the semantic hyper Figure 3 D module is as shown in Figure 5 . Multiply the input feature by the semantic hypergraph operator to obtain the semantic information feature map, that is:
[0151] ;
[0152] where, , o represents the number of semantic hyper Figure 3 D modules, represents the total number of semantic hyper Figure 3 D modules, represents the semantic feature output by the previous semantic hyper Figure 3 D module, represents that the input feature of the input semantic hyper Figure 3 D module is the input image, represents the semantic information feature map, represents the weight coefficient for calculating the semantic information feature map;
[0153] To better extract the spectral features, in the embodiment of the present invention, the semantic information feature map is further subjected to three-dimensional convolution processing to obtain the semantic feature output by the current semantic hyper Figure 3 D module, that is:
[0154]
[0155] where, Represents the current semantic hyper Figure 3 The semantic features output by the D module Represents three-dimensional convolutional processing Represents a real number matrix of size In the embodiments of the present invention, the three-dimensional convolutional processing includes performing convolutional processing on the features to be processed through three consecutive three-dimensional convolutional layers, and the features after convolution in each three-dimensional convolutional layer are activated by the ReLU activation function.
[0156] In summary, the complete semantic hyper Figure 3 The feature extraction process of the D module can be expressed as:
[0157]
[0158] Among them, Represents the feature extraction operation of the semantic hyper Figure 3 D module.
[0159] An important feature of hypergraph learning is that it can support cross-pixel node collaborative optimization. Pixels with similar abundance distributions are grouped in the same hyperedge, and the model can consider the collaborative relationships between different pixels during optimization. The semantic hyper Figure 3 D module provided by the present invention effectively reduces the difficulty of information propagation by mapping highly coupled complex information into a more explicit high-level semantic space, which is beneficial to more accurate detail reconstruction.
[0160] The above three branches respectively extract the spectral, spatial, and high-level semantic information in the input image, which can provide relatively comprehensive guidance for the reconstruction process. However, there are domain differences between the features extracted by each branch, and independent learning paths often lead to poor compatibility of the features used for reconstruction. The sensitive band extraction branch in step S2 of the present invention can achieve cross-domain information cross-fusion.
[0161] Specifically, the sensitive band extraction branch in step S2 includes a weight acquisition module and a reference residual module; where:
[0162] In the weight acquisition module, pixel value statistics are performed on each band in the spectral features, spatial features, and semantic features to estimate the pixel value probability distribution of each band , and the entropy value of each band in each feature is calculated through the maximum entropy formula, that is:
[0163] ;
[0164] Among them represents the set of pixel values of the i-th band, and the band with the maximum entropy is selected as the sensitive band of each feature.
[0165] The reference residual module includes no less than 3 cascaded residual blocks. Between two adjacent residual blocks, taking the sensitive bands of three features as weights, multiply with the corresponding elements of the input features of the previous residual block, and then add the multiplied features to the input features of the subsequent residual block to obtain the output features. In the embodiments of the present invention, the reference residual module has a total of six residual blocks. The output of the second residual block is fused with the sensitive bands of other branches and then fed into the third residual block. At this time, the above content can be expressed as:
[0166] ;
[0167] ;
[0168] Among them, represents the weight, , and respectively represent the sensitive bands in spectral features, spatial features, and semantic features, represents the output features of the second residual block, that is, the input features of the third residual block, represents the output features of the third residual block, represents the operation of the residual block on the feature ; The remaining residual blocks include no less than 2 cascaded convolutional layers. After the input features of the current residual block are subjected to the convolution operation of the convolutional layer, they are added to the corresponding elements of the input features of the current residual block to obtain the output features; the output features obtained by the last residual block are the high-quality semantic features. In the embodiments of the present invention, each residual block includes two cascaded convolutional layers, and the activation function between the two convolutional layers is ReLU.
[0169] The intervention of sensitive band attention in the sensitive band extraction branch dynamically adjusts the weights of each feature during information fusion to more accurately guide the interaction between spectrum, space, and high-level semantic information. This method ensures the collaborative work of different feature domains and finally improves the super-resolution effect.
[0170] In the restoration and reconstruction branch in step S2, the corresponding elements of the spectral feature, spatial feature, semantic feature, and high-quality semantic feature are added together, and the added feature is fed into a convolution operation and a transposed convolution operation that are each performed at least once in sequence. The obtained feature is added to the corresponding elements of the preliminary super-resolution image obtained by upsampling the input image. Then, after the added feature is subjected to at least once convolution operation in sequence, the reconstructed remote sensing hyperspectral image is obtained. In the embodiment of the present invention, the added feature is fed into a convolution operation and a transposed convolution operation that are each performed once in sequence, and the activation function between the convolution operation and the transposed convolution operation is the ReLU activation function; the input image is subjected to bicubic interpolation to obtain a preliminary super-resolution image with the same resolution as the reconstructed remote sensing hyperspectral image; the feature obtained by adding the preliminary super-resolution image and the feature output by the transposed convolution operation is subjected to two convolution operations, and the activation function between the two convolution operations is the ReLU activation function, to obtain the final reconstructed remote sensing hyperspectral image.
[0171] During the training process of step S2, the embodiment of the present invention preferably uses a loss function jointly constructed by an L1 loss function, a SAM loss function, and a gradient loss function to train the hybrid neural network. The constructed loss function is as follows:
[0172] ;
[0173] wherein, represents the loss function constructed in the embodiment of the present invention, y represents the true remote sensing hyperspectral image corresponding to the input image, represents the reconstructed remote sensing hyperspectral image corresponding to the input image, represents the L1 loss function, represents the SAM loss function, represents the gradient loss function, , and respectively represent the weight coefficients of the L1 loss function, the SAM loss function, and the gradient loss function. In the embodiment of the present invention, During the training process, the optimizer selects Adam (the first-order momentum decay rate is 0.9, and the second-order momentum decay rate is 0.999), the batch size is 4, and the initial learning rate is 0.0001. The number of training epochs is set to 50. The parameters of the hybrid neural network are updated according to the loss function. As the loss function gradually decreases, the super-resolution reconstruction effect of the hybrid neural network on the low-resolution input images becomes better and better. After every few rounds of update iterations of the network parameters, the performance of the hybrid network model is tested using the validation set. When the results of the objective image quality evaluation metrics of the reconstructed images in the validation set no longer increase or reach the preset upper limit of the number of training epochs, the training ends. In one embodiment, when the fluctuation range of the image evaluation metrics in the validation set is lower than 0.1 within ten epochs, the training ends.
[0174] S3: Input the low-resolution remote sensing hyperspectral image to be reconstructed into the hybrid network model obtained in step S2 to obtain the corresponding reconstructed remote sensing hyperspectral image.
[0175] To verify the effectiveness of the hyperspectral image super-resolution reconstruction method based on a hybrid neural network provided by the present invention, embodiments of the present invention use peak signal-to-noise ratio (PSNR), structural similarity (SSIM), spectral angle similarity (SAM), and relative global error separation metric (ERGAS) to objectively evaluate the quality of the reconstructed hyperspectral images, and use bicubic interpolation (Bicubic), 3D-FCNN (published in the paper "Hyperspectral Image Spatial Super-Resolution via 3D Full Convolutional Neural Network" in the journal Remote Sensing), GDRRN (published in the paper "Single Hyperspectral Image Super-resolution with Grouped Deep Recursive Residual Network" in the journal 2018 IEEE Fourth International Conference on Multimedia Big Data (BigMM)), SSPSR (published in the paper "Learning Spatial-Spectral Prior for Super-Resolution of Hyperspectral Imagery" in the journal IEEE TRANSACTIONS ON COMPUTATIONAL IMAGING), MCNet (published in the paper "Mixed 2D / 3D Convolutional Network for Hyperspectral Image Super-Resolution" in the journal Remote Sensing), ERCSR (published in the paper "Exploring the Relationship Between 2D / 3D Convolution for Hyperspectral Image Super-Resolution" in the journal IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING), GELIN (published in the paper "A Group-Based Embedding Learning and Integration Network for Hyperspectral ImageSuper-Resolution), SRDNet (the paper "Hyperspectral Image Super-Resolution via Dual-Domain Network Based on Hybrid Convolution" published in the journal IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING), and MSSR (the paper "Remote Sensing Hyperspectral Image Super-Resolution via Multidomain Spatial Information and Multiscale Spectral Information Fusion" published in the journal IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING) are respectively compared with the method provided by the present invention on the MDAS dataset, Pavia Centre dataset, and Houston dataset. The objective comparison results of the method provided by the present invention and the comparative methods on the MDAS dataset are shown in Table 1:
[0176] Table 1:
[0177]
[0178] From the experimental results, the method provided by the present invention has obtained the optimal numerical results under three magnification factors of 2×, 3×, and 4×. Specifically, in the 2× super-resolution reconstruction, compared with the sub-optimal results, the four indicators are respectively optimized by 0.042 dB, 0.0002, 0.13, and 0.077. Compared with other indicators, the SAM of the reconstructed image of the present invention has a more obvious lead. The SAMs under the three magnification factors are respectively optimized by 0.13, 0.185, and 0.155, which proves the effectiveness of the method provided by the present invention in improving spectral fidelity. The hypergraph constructed based on the high-frequency information and low-frequency information in the spectral signal effectively establishes the spectral similarity relationship in the cross-region. The proposed spectral hypergraph mutual attention module integrates the low-frequency features and high-frequency features in the spectrum globally, and the mutual attention mechanism further captures and strengthens the long-range correlation between signals. The experimental results show that the above strategies for spectral information play a very positive role in improving spectral fidelity. The comparison result graph is as Figure 6 shown, from Figure 6It can be observed that the reconstructed edges and textures of the method provided by the present invention are clearer, while other methods usually have more obvious blurring and deformation. Whether in terms of numerical values or visual effects, the method provided by the present invention has more obvious advantages.
[0179] The objective comparison results of the method provided by the present invention and the comparative method on the Pavia Centre dataset are shown in Table 2. It can be seen from Table 2 that in the 3× super-resolution reconstruction task, compared with the sub-optimal results, the method provided by the present invention has improved by 0.171dB, 0.0042, 0.097, and 0.107 respectively in four indicators; in the 4× super-resolution reconstruction task, the method provided by the present invention has obtained two optimal indicators and two sub-optimal indicators. The method provided by the present invention constructs hypergraphs with different characteristics based on spectral high / low-frequency information, spatial high / low-frequency information, and high-level semantic space respectively, and realizes more accurate reconstruction by comprehensively characterizing the pixel correlation within the global range. The comparison result graph is as Figure 7 shown. It can be seen from Figure 7 that in the area marked by the red box, the method provided by the present invention can better restore tiny details and alleviate spatial distortion.
[0180] Table 2:
[0181]
[0182] The objective comparison results of the method provided by the present invention and the comparative method on the Houston dataset are shown in Table 3. It can be seen from Table 3 that in the 2× and 3× super-resolution reconstruction tasks, all the indicators of the reconstructed images obtained by using the method provided by the present invention are the best, and the PSNR reaches 38.119dB (+0.032dB) and 35.004dB (+0.066dB) respectively. This proves that the hypergraph learning based on high and low frequency features can well capture the main structure and subtle changes in the image. The hypergraph branch of the method provided by the present invention is used to capture global information, while the residual branch is used to mine local information. The feature extraction at multiple scales enables the network to take into account the neighborhood details and long-range correlation of information nodes during the learning process. To sum up, our method has excellent performance on all three datasets, which proves the superiority and robustness of the model. The comparison result graph is as Figure 8 shown. It can be seen from Figure 8 that in the area marked by the red box, the method provided by the present invention reconstructs more realistic road details.
[0183] Table 3:
[0184]
[0185] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.
[0186] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A remote sensing hyperspectral image super-resolution reconstruction method based on hybrid neural network, characterized in that: include: S1: Obtain remote sensing hyperspectral image dataset and preprocess the dataset to obtain a training set; S2: Construct a hybrid neural network, and train the hybrid neural network using the training set obtained in step S1 to obtain a hybrid network model; the hybrid neural network includes: The spectral hypergraph branch performs a one-dimensional wavelet transform on the input image to obtain an approximation coefficient and a detail coefficient; constructs a low-frequency feature hypergraph and a high-frequency feature hypergraph according to the approximation coefficient and the detail coefficient, and then converts the low-frequency feature hypergraph and the high-frequency feature hypergraph to obtain a low-frequency feature hypergraph operator and a high-frequency feature hypergraph operator respectively; the input image is used as an input feature, and is input into the spectral hypergraph mutual attention module with the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator, and the output feature is then processed with the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator by at least one spectral hypergraph mutual attention module to obtain a spectral feature; The spatial hypergraph branch extracts spatial features from the input image; The semantic hypergraph branch extracts semantic features from the input image; A sensitive band extraction branch is used to fuse the spectral feature, the spatial feature and the semantic feature, and to extract high-quality semantic features of the input image using the fused features; The restoration and reconstruction branch is based on the preliminary super-resolution image of the input image and combines the spectral feature, the spatial feature, the semantic feature and the high-quality semantic feature to obtain a reconstructed remote sensing hyperspectral image; S3: Input the low-resolution remote sensing hyperspectral image to be reconstructed into the hybrid network model obtained in step S2 to obtain the corresponding reconstructed remote sensing hyperspectral image.
2. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 1, characterized in that: In step S1, the preprocessing process of the data set includes: The data set is segmented, amplified and degraded to obtain a low-resolution input image, which is used as an input of the hybrid neural network.
3. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 1, characterized in that: In the spectral hypergraph branch in step S2: The low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator are obtained by the following conversion: ; ; in, represents the low-frequency feature hypergraph operator, represents the high-frequency feature hypergraph operator, represents the low-frequency feature hypergraph, represents the high-frequency feature hypergraph, and are the diagonal matrices representing the hyperedge degree and vertex degree respectively, and E is initialized to the identity matrix.
4. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 1, characterized in that: In the spectral hypergraph mutual attention module: The input features are multiplied by the low-frequency feature hypergraph operator and the high-frequency feature hypergraph operator respectively to obtain a low-frequency spectral feature map and a high-frequency spectral feature map, that is: ; ; in, , n represents the number of the spectral hypergraph mutual attention modules, N represents the total number of the spectral hypergraph mutual attention modules, represents the spectral features output by the previous spectral hypergraph mutual attention module, Indicates that the feature input to the spectral hypergraph mutual attention module is the input image, represents the low-frequency spectral characteristic graph, represents the high-frequency spectral characteristic graph, and Indicates the weight coefficient for calculating the low-frequency spectral feature graph and the high-frequency spectral feature graph; Extract the query matrix from the low-frequency spectral feature map, extract the key matrix and the value matrix from the high-frequency spectral feature map, and then perform a multi-head mutual attention mechanism operation on the query matrix, the key matrix and the value matrix to obtain the spectral features output by the current spectral hypergraph mutual attention module, that is: ; in, represents the spectral features output by the current spectral hypergraph mutual attention module, represents the query matrix, represents the bond matrix, represents the value matrix, Represents the operation of the multi-head mutual attention mechanism.
5. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 1, characterized in that: In the spatial hypergraph branch in step S2: The input image is reduced in dimension, and then the reduced image is subjected to a two-dimensional wavelet transform to obtain the low-frequency component and high-frequency component of the input image, namely: ; in, represents the low-frequency component, represents the high-frequency components in the horizontal, vertical and diagonal directions of the input image, represents the two-dimensional wavelet transform, represents the image retained after the input image is reduced in dimension; The low-frequency component and the high-frequency component are subjected to double bicubic interpolation, the extracted frequency features are extended, and the low-frequency spatial information hypergraph and the high-frequency spatial information hypergraph of the input image are constructed by the following formula: ; in, represents the low-frequency spatial information hypergraph, represents the high-frequency spatial information hypergraph, Represents a complete function for constructing the low-frequency spatial information hypergraph and the high-frequency spatial information hypergraph; The corresponding elements of the low-frequency spatial information hypergraph and the high-frequency spatial information hypergraph are added to obtain the spatial information hypergraph of the input image, and the spatial information hypergraph is converted by the following formula to obtain the spatial information hypergraph operator: ; in, represents the spatial information hypergraph operator, represents the spatial information hypergraph, and The diagonal matrices representing the hyperedge degree and vertex degree respectively, E is initialized to the identity matrix; The input image is used as an input feature and is input into the spatial hypergraph self-attention module together with the spatial information hypergraph operator. The output feature is then processed together with the spatial information hypergraph operator by at least one spatial hypergraph self-attention module to obtain the spatial feature.
6. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 5, characterized in that: In the spatial hypergraph self-attention module: The input feature is multiplied by the spatial information hypergraph operator to obtain a spatial information feature graph, namely: ; in, , m represents the number of the spatial hypergraph self-attention modules, M represents the total number of the spatial hypergraph self-attention modules, represents the spatial features output by the previous spatial hypergraph self-attention module, Indicates that the feature input to the spatial hypergraph self-attention module is the input image, represents the spatial information feature map, represents the weight coefficient for calculating the spatial information feature map; Extract the query matrix, key matrix and value matrix from the spatial information feature graph, and then perform a multi-head self-attention mechanism operation on the query matrix, the key matrix and the value matrix to obtain the spatial features output by the current spatial hypergraph self-attention module, that is: ; in, represents the spatial features output by the current spatial hypergraph self-attention module, represents the query matrix, represents the bond matrix, represents the value matrix, Represents the operation of the multi-head self-attention mechanism.
7. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 1, characterized in that: In the semantic hypergraph branch in step S2: The abundance matrix is parsed from the input image using a non-negative matrix factorization algorithm; According to the abundance matrix, the semantic hypergraph of the input image is constructed by the following formula: ; in, represents the semantic hypergraph, represents the abundance matrix, represents a complete function for constructing the semantic hypergraph; The semantic hypergraph is converted into a semantic hypergraph operator by the following formula: ; in, represents the semantic hypergraph operator, and The diagonal matrices representing the hyperedge degree and vertex degree respectively, E is initialized to the identity matrix; The input image is used as an input feature and is input into the semantic hypergraph 3D module together with the semantic hypergraph operator. The output feature is then processed together with the semantic hypergraph operator by no less than one semantic hypergraph 3D module to obtain the semantic feature.
8. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 7, characterized in that: In the semantic hypergraph 3D module: The input feature is multiplied by the semantic hypergraph operator to obtain a semantic information feature graph, namely: ; in, , o represents the number of the semantic hypergraph 3D modules, represents the total number of 3D modules of the semantic hypergraph, represents the semantic features output by the previous semantic hypergraph 3D module, Indicates that the feature input to the semantic hypergraph 3D module is the input image, represents the semantic information feature graph, Represents the weight coefficient for calculating the semantic information feature map; The semantic information feature map is then subjected to three-dimensional convolution processing to obtain the semantic features output by the current semantic hypergraph 3D module.
9. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 1, characterized in that: The sensitive band extraction branch in step S2 includes a weight acquisition module and a reference residual module; wherein: In the weight acquisition module, pixel value statistics are performed on each band in the spectral feature, the spatial feature and the semantic feature to estimate the probability distribution of the pixel value of each band; the entropy value of each band in each feature is calculated by the maximum entropy formula, and the band with the largest entropy is selected as the sensitive band of each feature; The reference residual module includes no less than 3 cascaded residual blocks, wherein between two adjacent residual blocks, the sensitive bands of the three features are used as weights, multiplied with the corresponding elements of the input features of the previous residual block, and the multiplied features are added with the input features of the next residual block to obtain output features; each residual block includes no less than 2 cascaded convolutional layers, and the input features of the current residual block are convolved by the convolutional layer and then added with the corresponding elements of the input features of the current residual block to obtain output features; the output features obtained by the last residual block are the high-quality semantic features.
10. The method for super-resolution reconstruction of remote sensing hyperspectral images based on hybrid neural network according to claim 1, characterized in that: In the restoration and reconstruction branch in step S2, The spectral features, the spatial features, the semantic features and the corresponding elements of the high-quality semantic features are added, and the added features are sequentially subjected to at least one convolution operation and a transposed convolution operation, and the obtained features are added to the corresponding elements of the preliminary super-resolution image obtained by upsampling the input image, and the added features are sequentially subjected to at least one convolution operation to obtain the reconstructed remote sensing hyperspectral image.
Citation Information
Patent Citations
Remote sensing hyperspectral image super-resolution reconstruction method based on hypergraph neural network
CN118096536A