Double-domain feature fusion method for hyperspectral image classification
Through the two-domain feature fusion method, combined with deformable convolution, frequency domain Transformer and deep convolution techniques, the problems of insufficient robustness of feature extraction and low classification accuracy of hyperspectral image are solved, and higher classification accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510290144.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-17
AI Technical Summary
The hyperspectral image feature extraction is not robust to noise interference and mixed pixel processing. There are many wrong classification points in the classification of similar categories and low shape repetition rate in the classification are not high.
The two-domain feature fusion method is used to perform preliminary feature extraction through deformable convolution, a two-branch feature extraction network is constructed, and the frequency domain and spatial features are extracted respectively using the frequency domain Transformer module and the deep convolution module, and the weight is allocated in the spatial and channel dimensions through the dual-domain fusion module, and finally the deformation gate feedforward network is used to capture spatial information for classification.
The model's robustness to noise interference and mixed pixels is significantly improved, the classification accuracy of similar categories is improved, and the classification accuracy of categories with low shape repetition rate is improved.
Smart Images

Figure CN120164069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of hyperspectral image technology, and particularly to a dual-domain feature fusion method for hyperspectral image classification. Background Art
[0002] Currently, the mainstream HSI classification methods can be divided into traditional machine learning methods, methods based on Convolutional Neural Network (CNN), and methods based on Transformer (self-attention mechanism). Traditional machine learning methods use handcrafted features to train classifiers, which require high domain expertise and engineering skills. It is difficult to optimize the feature design for different datasets, resulting in difficulty in balancing the robustness and discrimination ability of features. In addition, a single feature is difficult to comprehensively represent image information, and multiple features need to be combined. However, this feature combination depends on artificial design, increasing the complexity.
[0003] The methods based on CNN extract multi-level spatial and spectral features automatically from the original hyperspectral data through an end-to-end learning process. CNN can effectively extract local spatial features through convolution operations and capture more complex and abstract feature information by stacking multiple convolutional layers. To better utilize the spectral information of hyperspectral images, many CNN models also incorporate convolutional operations in the spectral dimension, thereby enhancing the model's discrimination ability for different materials and objects. When performing hyperspectral image classification, CNN is usually divided into two parts: one is the spatial feature extraction network, which captures the spatial structure of the image through convolutional layers; the other is the spectral feature extraction network, which uses spectral information to finely classify the category of each pixel. Through joint optimization, CNN can fully utilize the spatial and spectral information of the image, thereby improving the classification accuracy. Compared with traditional methods, the hyperspectral image classification based on CNN has higher classification accuracy and stronger robustness. Especially when facing problems such as complex ground object types, noise interference, and data imbalance, CNN can automatically adjust and optimize its feature extraction process to further improve the classification effect. The hyperspectral image classification method based on convolutional neural network has advantages such as parameter sharing and sliding window, but it is insensitive to the global features of the image due to the nature of the local receptive field of convolution.
[0004] In the Transformer - based hyperspectral image classification method, Transformer first fuses the spatial and spectral information of each pixel point, maps it into a feature vector, and retains the spatial information through position encoding. The self - attention mechanism enables each pixel to consider the spectral and spatial relationships of other pixels during classification, thereby effectively extracting global features. The multi - head self - attention further enhances the model's expressive ability, can learn multiple feature sub - spaces in parallel, and improves the classification accuracy. The Transformer - based hyperspectral image classification method has the ability of dynamics and global modeling, but it cannot pay attention to the local features of the image. The self - attention mechanism has high overhead in terms of the number of parameters and computational complexity. At the same time, due to the characteristics of high data dimensions and limited feature expression ability in hyperspectral image analysis, traditional spatial - domain features may result in feature loss due to excessive dimensions. Summary of the Invention
[0005] Object of the Invention: The object of the present invention is to provide a dual - domain feature fusion method for hyperspectral image classification, which solves the problem of the robustness of hyperspectral image feature extraction against noise interference and mixed - pixel processing; solves the problem of many misclassification points in similar - category classification in hyperspectral image classification; and solves the problem of the classification accuracy of categories with low shape repetition rate in hyperspectral image classification.
[0006] Technical Solution: A dual - domain feature fusion method for hyperspectral image classification according to the present invention includes the following steps:
[0007] (1) Obtain an open - source dataset and perform pre - processing;
[0008] (2) Use deformable convolution for preliminary feature extraction;
[0009] (3) Construct a dual - branch feature extraction network to process the result obtained in step (2);
[0010] (4) Use the dual - branch feature fusion module DFF to assign weights in the spatial and channel dimensions respectively;
[0011] (5) Use the deformable gated feed - forward network DGFN to capture spatial information and fit the category shape to obtain the classification result.
[0012] Further, in step (1), the pre - processing is specifically as follows: Use principal component analysis technology to perform dimensionality reduction pre - processing in the spectral dimension of the original data.
[0013] Further, in step (3), the dual - branch feature extraction network composed of the frequency - domain FDformer and depth convolution includes the following steps:
[0014] (31) The frequency-domain FDformer is used to extract the frequency-domain features and global features of an image; first, the two-dimensional discrete Fourier transform DFT is introduced as follows: Given a two-dimensional signal X[m, n], 0 ≤ m ≤ M - 1, 0 ≤ n ≤ N - 1, the two-dimensional DFT of X[m, n] is expressed as:
[0015]
[0016] The shallow features after deformable convolution are taken as input with H×W non-overlapping small blocks, and the flat small blocks are projected into L = H×W tokens with a dimension of B and enter the frequency filtering layer; among them, the frequency filtering layer mixes and represents the tokens at different spatial positions: Given a token x ∈ R H×W×B , first perform a two-dimensional FFT along the spatial dimension to transform x to the frequency domain:
[0017] X = F[x] ∈ C H×W×B
[0018] where F[·] is the two-dimensional FFT; X is a complex tensor representing the spectrum of x.
[0019] (32) X processes high-frequency noise through a high-frequency denoiser;
[0020] (33) Modulate the spectrum by multiplying the learnable filter K ∈ C H×W×B by X:
[0021]
[0022] (34) Use the inverse FFT to transform the modulated spectrum ~X back to the spatial domain and update the token:
[0023]
[0024] The formula for learnable filtering comes from the frequency filter in digital image processing, where K can be regarded as a set of learnable frequency filters with different hidden dimensions. This filtering layer is equivalent to a deep global circular convolution with a filter size of H×W. Therefore, K is different from the standard convolution layer, which uses a relatively small filter size to strengthen the inductive bias of locality. And its complexity is O(DL log L), while the complexity of the ordinary deep global circular convolution in the spatial domain is O(DL 2 ), while greatly reducing the number of parameters while effectively obtaining global information.
[0025] (35) The deep convolutional network is used to obtain the spatial features and local features of an image, and the deep convolutional network gradually extracts features at different levels through multiple convolutional layers.
[0026] Further, step (32) is specifically as follows: First, calculate the power spectrum of X to identify the dominant frequency components; the power spectrum P is calculated by the square of the amplitude of the frequency components: P = |F|², which reflects the intensity of different frequency components in the time series; process the high-frequency components in the power spectrum P through a trainable threshold θ, and θ will be adjusted according to the spectral characteristics of the data; the formula is as follows:
[0027] X f = X ⊙ (P > θ)
[0028] where ⊙ represents element-wise multiplication, and (P > θ) is a binary mask, indicating that the frequencies with power higher than the threshold θ are retained, while other frequencies are filtered out.
[0029] Further, step (4) includes the following steps:
[0030] (41) Let the feature representation after FDformer be Y f ∈R H×W×B , and the feature representation of depth convolution be Y d ∈R H×W×B ; the representation after DDF fusion is Y f and Y d ; DDF includes two operations: spatial fusion SF and channel fusion CF; Y f is calculated by SF to obtain the spatial weight map denoted as S-Weight, with a size of R H×W×1 ; Y d is calculated by CF to obtain the channel weight map denoted as C-Weight, with a size of R 1×1×B ; the formula is as follows:
[0031] S-Weight(Y d ) = f(W2σ(W1Y d )),
[0032] C-Weight(Y f ) = f(W4σ(W3H GP Y f ))
[0033] where, H GP is global average pooling, f(·) is the sigmoid function, and σ(·) is the GELU function. W(·) represents the weight of convolution; the reduction ratios of W1 and W2 are r1 respectively, W3 has a reduction ratio of r2, and W4 has a dilation rate of r2; subsequently, apply their respective weight maps to the other input to achieve fusion; the formula is as follows:
[0034] SF = Y f ⊙ S-Weight(Y d )
[0035] CF = Y d ⊙C-Weight(Y f )
[0036] Add the elements of the two fused features, and the formula is as follows:
[0037] F add = SF + CF
[0038] Furthermore, in step (5), the deformable gate feed-forward network DGFN consists of deformable convolution and element-wise multiplication; specifically as follows: In the channel dimension, the feature map is divided into two parts, a convolution branch and a multiplication branch, and the formula is as follows:
[0039] Let the given input X ∈ R H×W×C , and the DGFN formula is:
[0040]
[0041] Among them, and represent linear projections, σ represents the GELU function, and W d is the learnable parameter of the deformable convolution; the fused feature passes through DGFN to obtain the final output feature F out ; then F out is refined through a convolutional layer, and then through a linear layer, and the softmax function is used to calculate the probability that the input belongs to a certain category, and the label with the largest probability value is the category of the sample.
[0042] A dual-domain feature fusion system for hyperspectral image classification according to the present invention includes:
[0043] Preprocessing module: used to obtain an open-source dataset and perform preprocessing;
[0044] Extraction module: used to perform preliminary feature extraction using deformable convolution;
[0045] Dual-branch module: used to construct a dual-branch feature extraction network to process the results obtained by the extraction module;
[0046] Allocation module: used to allocate weights in the spatial and channel dimensions respectively using the dual-branch feature fusion module DFF;
[0047] Classification module: used to capture spatial information using the deformable gate feed-forward network DGFN and fit the category shape to obtain the classification result.
[0048] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the dual-domain feature fusion methods for hyperspectral image classification.
[0049] A storage medium according to the present invention stores a computer program. When the computer program is executed by a processor, it implements any one of the dual-domain feature fusion methods for hyperspectral image classification.
[0050] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The frequency-domain Transformer module is used to extract the frequency-domain features of the hyperspectral image, effectively separating the useful signals and noise, efficiently resolving similar categories through the frequency information, combining the excellent spatial interaction ability of the Transformer to capture global features at the same time, and reducing the number of parameters; Using the preliminary deformable convolution can significantly improve the adaptability of the model to spatial information, enabling the network to have stronger flexibility at the earliest feature extraction stage, automatically adjusting the receptive field to adapt to the categories with low shape repeatability in the hyperspectral image, and improving the classification accuracy. Especially when dealing with the heterogeneity and complex background in the hyperspectral image, it has significant advantages; Using the deformable gate feed-forward network can eliminate the redundant information in the channel, capture spatial information, and refit the category shape again, significantly improving the classification accuracy; The DDFF proposed by the present invention passes the image after deformable convolution through the frequency-domain Transformer module and the depth convolution module respectively, and then uses the feature fusion module to fully interact the frequency-domain and spatial information in the HSI. At the same time, it can also obtain global features with the help of the good spatial information interaction ability of the Transformer and fuse them with local features. Finally, the features are further refined through the deformable gate feed-forward network. Thereby significantly improving the classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the network structure prototype of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0052] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0053] As Figure 1 shown, the embodiment of the present invention provides a hyperspectral image classification method based on multi-frequency high-dimensional feature representation, including the following steps:
[0054] (1) Obtain open-source datasets and perform preprocessing; specifically as follows: Indian Pines (IP), Pavia University (PU), Salinas Scene (SA), and Houston 2013 datasets. For the IP dataset and the Houston 2013 dataset, 10% of the pixels in each category are randomly selected as the training set. For the SA and PU datasets, 5% of the pixels in each category are randomly selected as the training set;
[0055] Considering that redundant bands in hyperspectral data affect the model performance, principal component analysis technology is used to perform dimensionality reduction preprocessing in the spectral dimension of the original data. Given a hyperspectral image X with a resolution of H×W×L as input, where H, W, and L are the height, width, and number of spectral bands respectively. Since it is a pixel-level classification task, each pixel in X forms a one-dimensional one-hot label vector Y=(y1,y2,…,y N ), where N represents the number of classification categories. Using PCA technology, the number of bands is reduced to C while keeping the spatial dimension unchanged, obtaining the dimensionality-reduced hyperspectral image X′, and the resolution becomes H×W×B. Through preprocessing, the hyperspectral data not only eliminates redundant bands but also retains the spatial information crucial for the classification task.
[0056] (2) Deformable convolution for preliminary feature extraction:
[0057] As Figure 1 (a) shown, the dimensionality-reduced HSI first passes through a deformable convolution. Introducing a deformable convolution in the initial stage of the classification task can enable the network to have stronger flexibility in the earliest feature extraction stage and automatically adjust the receptive field to adapt to various categories in the hyperspectral image.
[0058] (3) Dual-branch feature extraction network for dual-domain feature extraction:
[0059] In the present invention, in order to obtain the frequency-domain features and spatial features of an image, a dual-branch feature extraction network composed of a frequency-domain Transformer and a deep convolution is designed.
[0060] (a) Frequency-domain Transformer (FDformer):
[0061] The frequency-domain Transformer is the core component of DDFF, as Figure 1 (b) shown. It is used to extract the frequency-domain features and global features of an image. First, the two-dimensional discrete Fourier transform (DFT) is introduced. Given a two-dimensional signal X[m,n], 0≤m≤M - 1, 0≤n≤N - 1, then the two-dimensional DFT of X[m,n] is expressed as:
[0062]
[0063] The shallow features after deformable convolution are used as input in H×W non-overlapping small blocks, and the flat small blocks are projected into L = H×W tokens with dimension B and fed into the frequency filtering layer.
[0064] Frequency filtering layer. As an alternative to the self-attention layer, the frequency filtering layer can mix and represent tokens at different spatial positions. Given a token (x ∈ R H×W×B ), we first perform a two-dimensional FFT along the spatial dimension to transform x into the frequency domain:
[0065] X = F[x] ∈ C H×W×B #(2)
[0066] where F[·] is the two-dimensional FFT. At this time, X is a complex tensor representing the spectrum of x.
[0067] Then, X is processed by a high-frequency denoiser to remove high-frequency noise. High-frequency components usually represent the rapid fluctuations of the signal, and these fluctuations often deviate from the underlying trend or the pattern of interest, making them appear more random and difficult to interpret. Therefore, an adaptive high-frequency filter is proposed, which can dynamically adjust the filtering intensity according to the characteristics of the dataset and effectively remove high-frequency noise components. Especially when dealing with non-stationary data whose spectrum changes over time, this method is crucial. The filter adaptively sets a suitable frequency threshold for each specific time series according to the data characteristics. Specifically, first calculate the power spectrum of X to identify the dominant frequency components. The power spectrum P is calculated by squaring the amplitude of the frequency components: P = |F|², which reflects the intensity of different frequency components in the time series. The key to effective noise reduction lies in the processing of high-frequency components in the power spectrum P by the adaptive filter. This is achieved through a trainable threshold θ, which is adjusted according to the spectral characteristics of the data. This threshold, as a learnable parameter, is optimized through backpropagation during the training process, specifically adjusted by to ensure that θ can effectively distinguish the basic signal frequency from the noise.
[0068] X f = X ⊙ (P > θ) #(3)
[0069] where ⊙ represents element-wise multiplication, and (P > θ) is a binary mask, indicating that the frequencies with power higher than the threshold θ are retained, while other frequencies are filtered out.
[0070] The adaptive nature of the threshold θ ensures that while removing high-frequency noise, important signal information can be retained. By dynamically selecting the frequency threshold according to the characteristics of each time series dataset, the filtering process can be adjusted specifically, thereby improving the overall performance of the model in different data scenarios.
[0071] Next, modulate the spectrum by multiplying the learnable filter \(K\in\mathbb{C}\) H×W×B by \(X\):
[0072]
[0073] The filter \(K\) has the same dimension as \(X\), so it can capture global information. Finally, we use the inverse FFT to transform the modulated spectrum \(\sim X\) back to the spatial domain and update the token:
[0074]
[0075] The formula for learnable filtering is derived from frequency filters in digital image processing, where \(K\) can be regarded as a set of learnable frequency filters with different hidden dimensions. This filtering layer is equivalent to a deep global circular convolution with a filter size of \(H\times W\). Therefore, \(K\) is different from the standard convolutional layer, which uses a relatively small filter size to strengthen the inductive bias of locality. And its complexity is \(O(DL\log L)\), while the complexity of the ordinary deep global circular convolution in the spatial domain is \(O(DL 2 ), significantly reducing the number of parameters while effectively capturing global information.
[0076] (b) Depthwise Convolution (Deepwise Conv):
[0077] The depthwise convolutional network is used to obtain the spatial and local features of the image. The depthwise convolutional network gradually extracts features at different levels through multiple convolutional layers. The lower convolutional layers usually learn basic features such as edges and textures, while the higher convolutional layers capture more complex patterns, such as the shape and semantic information of objects. The systematic combination of the frequency-domain Transformer module and the depthwise convolutional network can make full use of the frequency-domain and spatial information in the HSI and fully fuse the global and local features, thus significantly improving the classification accuracy.
[0078] (c) Dual Feature Fusion Module (DFF) across Spatial and Channel Dimensions:
[0079] However, simply increasing the convolutional branches cannot efficiently and accurately achieve the fusion of global and local features. Use the dual-branch feature fusion module (DFF) to reassign weights to the features of the two branches in the spatial and channel dimensions respectively. Specifically as follows:
[0080] First, represent the features after passing through FDformer as \(Y\) f \(\in\mathbb{R}\) H×W×B , and represent the features after depthwise convolution (DW-Conv) as \(Y\) d \(\in\mathbb{R}\) H×W×B . Then, fuse \(Y\) through DDFf and Y d . Specifically, DDF contains two operations: spatial fusion (SF) and channel fusion (CF), as shown in Figure 1 (c)(d). Yf undergoes SF calculation to obtain the spatial weight map (denoted as S-Weight, with size R H ×W×1 ). Yd undergoes CF calculation to obtain the channel weight map (denoted as C-Weight, with size R 1×1×B ). The formulas are as follows:
[0081] S-Weight(Y d ) = f(W2σ(W1Y d )),
[0082] C-Weight(Y f ) = f(W4σ(W3H GP Y f ))#(6)
[0083] In the formula, H GP is global average pooling, f(·) is the sigmoid function, σ(·) is the GELU function. W(·) represents the weight of convolution. The reduction ratios of W2 and W2 are r1 respectively, W3 has a reduction ratio of r2, and W4 has a dilation rate of r2. Subsequently, the respective weight maps are applied to another input to achieve fusion. This process is expressed as:
[0084] SF = Y f ⊙S-Weight(Y d )
[0085] CF = Y d ⊙C-Weight(Y f )#(7)
[0086] Finally, the two fused features are added element-wise.
[0087] F add = SF + CF#(8)
[0088] (d) Deformable Door Feedforward Network (DGFN):
[0089] Traditional feedforward networks have one non-linear activation layer and two linear projection layers to extract features. However, they ignore the modeling of spatial information. In addition, redundant information in the channels affects the feature expression ability and cannot adapt to local deformations. To overcome the above limitations, a deformable door feedforward network (DGFN) is used to replace the traditional feedforward network. As shown in Figure 1As shown in (e), this module is a simple gate mechanism, consisting of deformable convolution and element-wise multiplication. In the channel dimension, the feature map is divided into two parts: a convolutional branch and a multiplication branch. Generally speaking, given the input X ∈ R H×W×C , the DGFN formula is:
[0090]
[0091]
[0092] where and represent linear projections, σ represents the GELU function, and W d are the learnable parameters of the deformable convolution. Compared with the FFN, DGFN can capture non-linear spatial information, reduce the channel redundancy of the fully connected layer, fit local deformations, and the fused features pass through DGFN to obtain the final output feature F out .
[0093] F out passes through a convolutional layer to refine the features, and then through a linear layer, and the softmax function is used to calculate the probability that the input belongs to a certain class. The label with the largest probability value is the class of the sample.
Claims
1. A dual-domain feature fusion method for hyperspectral image classification, characterized in that: The following steps are involved: (1) Obtain open source datasets and preprocess them; (2) Use deformable convolution to perform preliminary feature extraction; (3) constructing a dual-branch feature extraction network to process the results obtained in step (2); (4) Using the dual-branch feature fusion module DFF to assign weights in the spatial and channel dimensions respectively; (5) The deformable gate feedforward network (DGFN) is used to capture spatial information and fit the category shape to obtain the classification result.
2. The dual-domain feature fusion method for hyperspectral image classification according to claim 1, characterized in that: In step (1), the preprocessing is specifically as follows: principal component analysis technology is used to perform dimension reduction preprocessing on the spectral dimension of the original data.
3. The dual-domain feature fusion method for hyperspectral image classification according to claim 1, characterized in that: In step (3), the dual-branch feature extraction network consisting of frequency domain FDformer and deep convolution includes the following steps: (31) The frequency domain FDformer is used to extract the frequency domain features and global features of the image. First, the two-dimensional discrete Fourier transform DFT is introduced as follows: Given a two-dimensional signal X[m, n], 0≤m≤M-1, 0≤n≤N-1, the two-dimensional DFT of X[m, n] is expressed as: The shallow features after deformable convolution are taken as H×W non-overlapping small blocks as input, and the flattened small blocks are projected into L=H×W tokens of dimension B and enter the frequency filter layer; among them, the frequency filter layer mixes tokens representing different spatial positions: let the given token be x∈R H×W×B , we first perform a 2D FFT along the spatial dimension to transform x into the frequency domain: X=F[x]∈C H×W×B where F[·] is the two-dimensional FFT and X is a complex tensor representing the frequency spectrum of x. (32)X processes high frequency noise through a high frequency denoiser; (33) By using the learnable filter K∈C H×W×B Multiply by X to get the modulation spectrum: Among them, the formula of learnable filtering comes from the frequency filter in digital image processing, where K is regarded as a set of learnable frequency filters with different hidden dimensions; (34) Use inverse FFT to transform the modulation spectrum ~X back to the spatial domain and update the token: (35) Deep convolutional networks are used to obtain the spatial and local features of images. Deep convolutional networks gradually extract features at different levels through multiple convolutional layers.
4. The dual-domain feature fusion method for hyperspectral image classification according to claim 3, characterized in that: Step (32) is as follows: First, the power spectrum of X is calculated to identify the dominant frequency components; the power spectrum P is calculated by the square of the amplitude of the frequency component: P = |F| 2 , reflects the intensity of different frequency components in the time series; the high-frequency components in the power spectrum P are processed by a trainable threshold θ, and θ will be adjusted according to the spectral characteristics of the data; the formula is as follows: X f =X⊙(P>θ) where ⊙ represents element-wise multiplication and P>θ is a binary mask, indicating that frequencies with power above a threshold θ are retained while other frequencies are filtered out.
5. The dual-domain feature fusion method for hyperspectral image classification according to claim 1, characterized in that: Step (4) comprises the following steps: (41) Let the feature representation after FDformer be Y f ∈R H×W×B , the feature of the deep convolution is represented as Y d ∈R H×W×B ; After DDF fusion Y f and Y d ; DDF contains two operations: spatial fusion SF and channel fusion CF; Y f The spatial weight map obtained by SF calculation is recorded as S-Weight, and its size is R H×W×1 ; Y d The channel weight map obtained by CF calculation is recorded as C-Weight, and its size is R 1 ×1×B ; The formula is as follows: S-Weight(Y d )=f(W2σ(W1Y d )), C-Weight(Y f )=f(W4σ(W3H GP Y f )) Among them, H GP is the global average pooling, f(·) is the sigmoid function, σ(·) is the GELU function; W(·) represents the weight of the convolution; the reduction ratios of W1 and W2 are r1, W3 has a reduction ratio of r2, and W4 has an expansion ratio of r2; then, each weight map is applied to the other input to achieve fusion; the formula is as follows: SF=Y f ⊙S-Weight(Y d ) CF=Y d ⊙C-Weight(Y f ) The two fused features are added element by element. The formula is as follows: F add =SF+CF。 6. The dual-domain feature fusion method for hyperspectral image classification according to claim 1, characterized in that: In step (5), the deformable gate feedforward network DGFN consists of deformable convolution and element-wise multiplication; specifically, in the channel dimension, the feature map is divided into two parts: the convolution branch and the multiplication branch, and the formula is as follows: Given an input X∈R H×W×C , the DGFN formula is: in, and represents linear projection, σ represents GELU function, W d is the learnable parameter of the deformable convolution; the fused features are passed through DGFN to obtain the final output feature F out ; then F out After a convolutional layer to refine the features, and then through a linear layer, the softmax function is used to calculate the probability that the input belongs to a certain category. The label with the largest probability value is the category of the sample.
7. A dual-domain feature fusion system for hyperspectral image classification, characterized in that: include: Preprocessing module: used to obtain open source data sets and perform preprocessing; Extraction module: used for preliminary feature extraction using deformable convolution; Dual-branch module: used to construct a dual-branch feature extraction network to process the results obtained by the extraction module; Allocation module: used to allocate weights in the spatial and channel dimensions respectively using the dual-branch feature fusion module DFF; Classification module: used to capture spatial information using the deformation gate feedforward network DGFN and fit the category shape to obtain the classification result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the dual-domain feature fusion method for hyperspectral image classification according to any one of claims 1 to 6 is implemented.
9. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a dual-domain feature fusion method for hyperspectral image classification according to any one of claims 1 to 6 is implemented.