A cross-attention fusion hyperspectral image classification method
Patent Information
- Application Number
- CN202611024517.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-15
Smart Images

Figure CN122760918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image classification technology, and more specifically, to a hyperspectral image classification method based on cross-attention fusion. Background Technology
[0002] Hyperspectral imaging technology can simultaneously acquire spatial contour information of ground features and continuous band spectral information, which has key application value in fields such as ground feature surveying, crop classification, and land resource remote sensing. Most of the current mainstream deep learning classification schemes use three-dimensional convolution or simple stitching and fusion to process spatial spectral data. Traditional processing schemes mostly directly use the original hyperspectral images to feed into the network. Hyperspectral data naturally has hundreds of continuous bands, and the information of adjacent bands is highly redundant. After the superposition of sensor imaging noise interference, redundant data and noise will significantly increase the number of model parameters and computational burden. At the same time, the spatial spectral mixed data that has not been split is prone to the problem that spatial texture information can cover up weak spectral details. The fine spectral fingerprint features specific to ground features are easily submerged by spatial information, resulting in insufficient completeness of spectral feature extraction.
[0003] Existing networks mostly use a single-stream structure to extract spatial spectral features simultaneously. These networks cannot specifically adapt to the heterogeneous data attributes of one-dimensional spectral sequences and two-dimensional spatial patches. Traditional one-dimensional convolution can only capture the correlation between local adjacent bands and is difficult to model the spectral dependence between distant bands. Conventional convolution is also difficult to simultaneously take into account small-scale fine textures and large-scale contiguous land cover outlines. As a result, the extracted features have weak representation capabilities and poor model generalization performance when faced with small sample training scenarios.
[0004] Most existing solutions use direct feature splicing or fixed weight summation to complete spatial-spectral integration. They cannot adaptively adjust the weight ratio of spectral and spatial information according to the attributes of different land features. For difficult samples such as different spectra of the same object or the same spectrum of different objects, the fixed fusion mode cannot flexibly supplement complementary information, which can easily lead to feature confusion. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the present invention aims to provide a hyperspectral image classification method based on cross-attention fusion.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A hyperspectral image classification method based on cross-attention fusion includes the following steps: Hyperspectral images are preprocessed and heterogeneous samples are constructed to obtain independent spectral branch input data and spatial branch input data; Spectral features are extracted based on the input data of the spectral branch, and spatial features are extracted based on the input data of the spectral branch. Using the spectral and spatial features as input, interactive features are generated through a bidirectional cross-attention mechanism; spatial-spectral joint discrimination features are obtained based on the interactive features. Input the spatial-spectral joint identification features into the classifier and output the land cover category prediction results.
[0007] Preferably, the hyperspectral image is preprocessed and heterogeneous samples are constructed to obtain independent spectral branch input data and spatial branch input data, specifically as follows: Dimensionality reduction and noise reduction are performed on hyperspectral images to suppress band redundancy and noise interference; Extract the continuous spectral curve of a single pixel or a small neighborhood as the input data for the spectral branch; Extract the spatial neighborhood map tiles after dimensionality reduction as input data for spatial branches.
[0008] Preferably, the spectral features are extracted using a sequence feature extraction network based on a self-attention mechanism, and the first feature is calculated using the following formula. Layer encoder output: ; ; in, For the first The layer input features are obtained by mapping the spectral branch input data through the embedding layer; MHSA is a multi-head self-attention module; FFN is a feedforward neural network; LN is layer normalization; the final output... As a spectral feature.
[0009] Preferably, the projection of the query, key, and value, and the attention output in the multi-head self-attention module are calculated as follows: ; ; in, , , For learnable projection matrices, This is the scaling factor.
[0010] Preferably, spatial features are extracted based on the spectral branch input data, specifically as follows: ; ; ; ; in, This represents a k×k convolution operation, and GAP stands for global average pooling.
[0011] Preferably, using the spectral features and spatial features as input, interactive features are generated through a bidirectional cross-attention mechanism, specifically as follows: Using spectral features as queries and spatial features as keys, first-direction cross-attention calculation is performed to obtain spectral-spatial interaction features; Using spatial features as queries and spectral features as keys, a second-direction cross-attention calculation is performed to obtain spatial-spectral interaction features; The spectral spatial interaction feature and the spatial spectral interaction feature are combined to form the interaction feature.
[0012] Preferably, the cross-attention between the first and second directions is calculated as follows: ; ; ; .
[0013] Preferably, the spatial-spectral joint discrimination feature is obtained based on the interaction feature, specifically as follows: ; ; ; in, It is a gated multilayer sensor.
[0014] Preferably, the classifier is a multilayer perceptron.
[0015] Preferably, the spatial-spectral joint identification features are input into the classifier, and the land cover category prediction result is output, specifically: ; in, This is the classification weight matrix. This is a bias term.
[0016] Compared with the prior art, the present invention has the following beneficial effects: In the preprocessing and heterogeneous sample construction stages, this invention splits the original hyperspectral image into two independent heterogeneous input data types: spectral and spatial. This achieves physical decoupling of spatial and spectral information from the input source, eliminating data interference from massive redundant bands in the hyperspectral spectrum and imaging noise, avoiding the defect of spatial pixel texture information masking weak spectral fingerprint details, and adapting to the differentiated data input forms of the dual-branch network. This effectively reduces the computational overhead caused by redundant data and improves the accuracy of subsequent feature extraction from the data source level. In the dual-branch feature extraction stage, the spectral branch relies on a self-attention sequence network to mine long-distance spectral dependencies across all bands, breaking through the limitation of one-dimensional convolution that can only capture local band information, and fully extracting the unique spectral features of ground objects. The spatial branch uses lightweight stacked convolution to achieve multi-scale spatial feature extraction, taking into account both subtle textures and large-scale contours with fewer parameters. The split extraction mode can specifically enhance the feature representation capabilities of each dimension of the spectrum and space, avoiding feature loss caused by spatial-spectral aliasing in a single-stream network, and significantly improving the robustness of feature extraction in small sample scenarios. The bidirectional cross-attention fusion design, combined with gated adaptive weighting, achieves bidirectional complementarity of spectral and spatial information. On one end, spatial information is used to constrain spectral features to improve the problem of different spectra for the same object, while on the other end, spectral identifiers are used to enhance spatial features to distinguish different ground objects with similar shapes. At the same time, the gated network can adaptively adjust the proportion of spatial and spectral weights according to different ground object samples, abandoning the coarse fusion method of fixed splicing, and generating spatial and spectral joint features with stronger discriminative power, which greatly alleviates the classification confusion caused by different objects with the same spectrum and the same object with different spectra. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of an embodiment of the present invention; Figure 2 This is a model network diagram of an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0019] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0020] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0021] Reference Figures 1-2 As shown.
[0022] The embodiments further illustrate the hyperspectral image classification method with cross-attention fusion proposed in this invention.
[0023] A hyperspectral image classification method based on cross-attention fusion includes the following steps: Hyperspectral images are preprocessed and heterogeneous samples are constructed to obtain independent spectral branch input data and spatial branch input data; Spectral features are extracted from the spectral branch input data; spatial features are extracted from the spectral branch input data. Using spectral and spatial features as inputs, interactive features are generated through a bidirectional cross-attention mechanism; spatial-spectral joint discrimination features are obtained based on the interactive features. Input the spatial-spectral joint identification features into the classifier and output the land cover category prediction results.
[0024] The hyperspectral image is preprocessed and heterogeneous samples are constructed to obtain independent spectral branch input data and spatial branch input data, specifically as follows: Dimensionality reduction and noise reduction are performed on hyperspectral images to suppress band redundancy and noise interference; Extract the continuous spectral curve of a single pixel or a small neighborhood as the input data for the spectral branch; Extract the spatial neighborhood map tiles after dimensionality reduction as input data for spatial branches.
[0025] Using hyperspectral raw images as the basic data source, and through hierarchical and step-by-step data processing logic, the spectral branch input data and spatial branch input data that are independent of each other are generated. This completes the decoupling of spatial and spectral information at the data source level, laying the data foundation for the separate feature extraction of the dual-branch network.
[0026] First, a unified preprocessing operation of dimensionality reduction and denoising is performed on the entire hyperspectral image. Hyperspectral images have hundreds of continuous spectral bands, and there is a lot of redundant data with highly overlapping information between adjacent bands. At the same time, imaging environment and sensor hardware interference will introduce random noise and impulse noise into the image. These redundant bands and useless noise will not only significantly increase the number of computational parameters of the subsequent network, but also obscure the true spectral and spatial characteristics of the ground objects themselves. Dimensionality reduction can filter out redundant bands with extremely low correlation, compressing the data dimension while retaining key feature information. The corresponding denoising process can filter out stray noise introduced by imaging, suppressing the interference of band redundancy and noise on the construction of subsequent branch data from the source, and optimizing the data quality of the two types of branch data.
[0027] After completing the unified preprocessing of the entire image, two heterogeneous input samples are constructed to adapt to the different branch requirements based on the differences in the data representation of spectral and spatial information. For the data construction stage of the spectral branch, continuous spectral curves of single pixels or small neighborhoods within a pixel are extracted as branch input data, relying on the pixel information after preprocessing. The continuous spectral numerical sequence corresponding to a single pixel can accurately characterize the unique spectral absorption and reflection variation patterns of the ground feature at that location, i.e., the ground feature's unique spectral fingerprint feature. Selecting small neighborhood pixel spectra can slightly incorporate weak spectral correlation information from the surrounding area while retaining the main spectral features. This type of one-dimensional curve data can accurately match the task attribute of the spectral branch network, which focuses on mining band-dimensional correlation features. This allows the branch model to focus on learning the spectral variation patterns between different bands, avoiding information occlusion of fine spectral details by spatial pixel texture information.
[0028] For the data construction stage of spatial branches, a fixed-size spatial neighborhood patch is extracted around the target pixel from the image data that has already undergone dimensionality reduction to serve as the branch input data. After dimensionality reduction, redundant bands are removed from the patch, and only the effective channel information that can support spatial texture representation is retained. The neighborhood patch retains the spatial features such as the outline of the ground objects around the central pixel, texture distribution, and spatial arrangement of adjacent ground objects in the form of a two-dimensional image. This two-dimensional patch data is adapted to the operation logic of the spatial branch convolutional network to extract multi-scale spatial structure features. In this way, two types of heterogeneous samples with independent data form and information dimension are finally obtained, realizing the physical decoupling of spectral data and spatial data at the input source.
[0029] Spectral features are extracted using a sequence feature extraction network based on a self-attention mechanism, and the first feature is calculated using the following formula. Layer encoder output: ; ; in, For the first The layer input features are obtained by mapping the spectral branch input data through the embedding layer; MHSA is a multi-head self-attention module; FFN is a feedforward neural network; LN is layer normalization; the final output... As a spectral feature.
[0030] Before the operation flow of a single-level encoder starts, the feature data output from the previous level... This is the input basis for the operation of this layer. The original input features originate from the one-dimensional sequence input data of the spectral branches obtained after preprocessing. The original spectral curve data first undergoes dimension mapping and feature encoding through the embedding layer, transforming the discrete band reflectance values into a feature vector sequence adapted to the operation format of the attention network. This forms a sequence that can be fed into the first layer. The layer encoder performs the calculations. Feature tensor. Entering the... During the layer encoder operation, the input features are... The data is fed into the Multi-Head Self-Attention (MHSA) module for global correlation modeling. The MHSA module can split into multiple parallel attention heads, traversing the dependencies between all band points in the entire spectral sequence from different representation scales. It autonomously allocates attention weights based on the spectral reflectance variation patterns of ground objects, capturing the implicit spectral correlation features between distant bands and adjacent bands. This effectively overcomes the limitation of traditional one-dimensional convolution, which can only capture local adjacent band information. After the computation, the output of the MHSA module and the original input features are compared. Residual summation is performed to mitigate the vanishing gradient problem that often occurs during deep network training, leveraging the structural advantages of residual connections. This ensures that shallow spectral detail information can be propagated through network layers. After summation, a layer normalization (LN) module performs standardization. Layer normalization unifies and regularizes the data distribution within the features, preventing excessive fluctuations in feature values from interfering with subsequent network convergence. This round of computation yields intermediate features. .
[0031] Intermediate features are obtained in the first round of calculation. Based on this, the second layer of encoder operation process is started, and In the feedforward neural network (FFN), the network consists of multiple layers of nonlinear transformation units. It relies on nonlinear activation functions to map and transform features, performing refined feature upscaling and filtering on the attention-filtered spectral features. This further refines the subtle spectral fingerprint information hidden within the bands, enhancing the spectral representation capabilities specific to ground features. Subsequently, the residual connection approach is continued, integrating the output of the feedforward network with the input features. The residuals are summed, and then the data is normalized again using a layer normalization (LN) process. After this complete computation, the final output features of the encoder at this layer are obtained. Following the aforementioned two-layer step-by-step computation logic, multiple encoders are stacked layer by layer. Through multi-layer attention modeling and nonlinear feature transformation, the spectral sequence information is continuously optimized, ineffective and redundant band information is constantly eliminated, and the spectral features of key ground features are enhanced. The final output is obtained after iterative computation across all encoder levels. This is a complete spectral feature extracted in depth, which fully preserves the reflectance variation law of ground objects across the entire spectral band.
[0032] The projection of queries, keys, and values, as well as the attention output, in the multi-head self-attention module are calculated as follows: ; ; in, , , For learnable projection matrices, This is the scaling factor.
[0033] For input features Three independent sets of learnable projection matrices are connected respectively. , , A linear transformation projection operation is performed. The three sets of projection matrices are parameter weights that the network continuously optimizes during the model training phase, each undertaking a differentiated feature mapping function. It is responsible for mapping the original spectral features to generate a query matrix Q. The query matrix is used to characterize the retrieval needs of each spectral band point and represents the retrieval benchmark for the spectral features at the current location to find related information in the entire band sequence. Complete the projection generation of the key matrix K. The key matrix corresponds to the index identifier of each band feature in the entire spectral sequence, which is used to perform similarity matching with the query content. The mapping yields a value matrix V, which stores the true spectral details carried by each band location itself. This value matrix serves as the data source for subsequent weighted extraction of effective features. After linear operations on three sets of matrices, the originally homogeneous input features... The data is split into three heterogeneous features—Q, K, and V—representing different functional attributes, completing the data splitting preparation before attention calculation. After entering the attention weight calculation stage, the query matrix Q and the transpose matrix K of the key matrix are... T Matrix multiplication is performed to quantify the similarity of features between any two spectral bands through the inner product operation. The magnitude of the result directly reflects the strength of the correlation between the spectral features of the two bands; a higher value indicates a stronger correlation in the spectral variation patterns of the two bands. To avoid the problem of the inner product result becoming too large as the feature dimension increases, causing the subsequent Softmax function to fall into the gradient saturation region, a scaling factor is introduced into the formula. To perform square root scaling, use... right The calculation result is scaled by division to regularize the numerical range of similarity data and stabilize gradient changes during network training. The scaled result is then fed into the Softmax activation function for normalization. The Softmax function transforms a set of similarity values without a fixed range into a probability weight distribution with all values between 0 and 1, where the sum of all values is 1. The final weight coefficient represents the proportion of attention allocated to each band relative to the current target band. A higher weight value indicates a greater contribution of the spectral information of the corresponding band to the feature representation of the central band. In the final feature aggregation step, the normalized attention weight matrix is multiplied by the value matrix V. Based on the weight coefficients, the original spectral features of each band in the value matrix are weighted and summed for filtering. Key associated bands with high weights retain more effective spectral information, while irrelevant band information with low weights is appropriately suppressed and filtered. Finally, the output result of a single self-attention operation is generated. This result fully integrates the associated spectral features of all bands and target locations in the entire sequence, accurately capturing the spectral dependence patterns hidden between distant bands and cross-interval bands in hyperspectral data. It provides the basic operational logic of single-head attention for multi-head self-attention to extract spectral details at multiple scales.
[0034] Spatial features are extracted based on the spectral branch input data, specifically as follows: ; ; ; ; in, This represents a k×k convolution operation, and GAP stands for global average pooling.
[0035] By using convolutional operations with equivalent receptive fields of different sizes to capture both fine-grained and coarse-grained spatial information in layers, and then through continuous operations of feature concatenation, channel integration, and global pooling, a dimensionally regularized spatial feature that integrates multi-scale information is finally generated. The entire process uses neighborhood spatial tiles that have undergone dimensionality reduction processing in the early stages. As the initial data source for computation, this 2D tile data fully preserves the original spatial information such as the edge contours, local textures, and arrangement of adjacent features around the central pixel, serving as the foundational material for multi-scale convolution feature extraction. In the first step of multi-scale feature extraction, the original input... Perform a single 3×3 convolution operation directly Obtain features A single 3×3 convolution corresponds to an effective receptive field of 3×3, which can focus on a small local pixel area and accurately capture fine-grained small-scale spatial features such as subtle textures and fine edges of ground objects, corresponding to the spatial variation patterns between nearby pixels. Then, the equivalent 5×5 receptive field feature extraction operation is performed by stacking two consecutive 3×3 convolutions. Obtained by performing progressive convolution Two concatenated 3×3 convolutions achieve an equivalent 5×5 receptive field coverage with fewer parameters than a single 5×5 convolution. This scale feature can cover a wider range of neighboring pixels, effectively extracting coarse-grained large-scale spatial features such as large-area outlines and contiguous regional structures. This lightweight parameter design enables differentiated extraction of dual-scale spatial information, balancing model computational efficiency with multi-scale feature capture capabilities. This results in two independent features at different scales. and Then, the Concat channel concatenation operation is used to directly concatenate and fuse the two feature sets along the channel dimension to generate... The stitching operation fully preserves all feature information from both scales, without losing fine texture details at the small scale or deleting global contour information at the large scale, achieving a complete summary of fine-grained and coarse-grained spatial features across the feature dimensions. The fused features after stitching are shown below. First, feed in a 1×1 convolution. Channel compression and feature filtering are performed. 1×1 convolution can adjust the number of channels without changing the feature space size, filtering out redundant feature channels generated after concatenation, strengthening the representation weights of effective spatial features, and reducing the data volume of subsequent operations. The features optimized by 1×1 convolution are then fed into the global average pooling (GAP) module for global aggregation. Global average pooling calculates the average value of feature values at all spatial locations within each independent channel of the feature map, thoroughly compressing the two-dimensional spatial dimension and finally obtaining the one-dimensional spatial features. This feature condenses all spatial structure information at different scales of the original image, removes redundant spatial dimensions, and simplifies feature morphology while preserving the multi-scale spatial attributes of ground features.
[0036] Using spectral and spatial features as input, interactive features are generated through a bidirectional cross-attention mechanism, specifically: Using spectral features as queries and spatial features as keys, first-direction cross-attention calculation is performed to obtain spectral-spatial interaction features; Using spatial features as queries and spectral features as keys, a second-direction cross-attention calculation is performed to obtain spatial-spectral interaction features; Among them, the combination of spectral spatial interaction features and spatial spectral interaction features constitutes the interaction features.
[0037] A bidirectional cross-attention computational architecture is adopted to realize cross-modal information interaction and fusion of spectral and spatial features. By relying on attention calculations in two opposite retrieval directions, the information complementarity between the two types of features is completed respectively. Then, the interactive features produced by bidirectional computation are integrated and summarized to form a complete interactive feature that takes into account both spectral and spatial attributes. This opens up the correlation path of spatial and spectral information from a bidirectional dimension and improves the feature confusion problem caused by different spectra of the same object and the same spectrum of different objects.
[0038] In the cross-attention computation of the first computational direction, the refined spectral features extracted by the spectral branch self-attention network are selected as the query vector. The query vector carries the unique full-band spectral reflectance patterns and spectral fingerprint information of each land cover, representing the information retrieval needs of the current spectral dimension. Then, the aggregated spatial features output by the multi-scale convolutional network are unified as key and value data. The spatial features condense all spatial structural information such as the texture of the land cover's neighborhood, edge contours, and regional distribution, serving as an index and information source for query matching. The cross-attention computation in this direction will traverse and match spatial information with strong correlation in the spatial features based on the retrieval needs of the spectral features through similarity matching calculation. Differentiated attention weights are assigned according to the degree of feature relevance. The effective structural information of the spatial dimension is embedded into the original spectral features through weighted aggregation, ultimately generating spectral spatial interaction features. These features supplement the corresponding land cover spatial context information on the basis of the original refined spectral information, making up for the defects of single spectral features that lack spatial constraints and are prone to spectral aliasing.
[0039] In the second reverse cross-attention operation step, the data sources of the query and key are swapped. The spatial features extracted at multiple scales are used as new query data. The retrieval task is initiated based on the land feature outline and regional distribution information carried by the spatial features. Instead, the spectral features refined by the Transformer encoder are used as key material. The band reflection variation rules stored in the spectral features provide matching references. This directional operation associates spatial features with the exclusive spectral information of the corresponding location. It uses attention weights to filter spectral details that are highly matched with the current spatial structure. The filtered effective spectral information is then weighted and fused into the original spatial features to obtain spatial spectral interaction features. While retaining the original spatial texture and outline information, this feature supplements the land feature-specific spectral identification information, improving the shortcoming of a single spatial feature that is difficult to distinguish spectrally similar land features based solely on their shape. After the bidirectional features have completed the cross-attention operation, the spectral spatial interaction features and spatial spectral interaction features obtained above are combined. The two types of interaction features have completed information exchange from the two dimensions of spectral-enhanced space and spatial-enhanced spectrum, respectively. The final interaction feature generated by the combination retains the refined spectral details and multi-scale spatial structure information, realizing the bidirectional complementary fusion of spectral and spatial information, and providing a high-quality fusion basis for the final classification feature output after spatial-spectral fusion.
[0040] The cross-attention between the first and second directions is calculated as follows: ; ; ; .
[0041] In the first set of computational branches for embedding spectral information into space, the original spectral features extracted through the spectral branch network are used first. As a data source, it connects to a learnable projection matrix. Complete the linear mapping and generate query features in that direction. , It fully inherits the full-band reflectance variation patterns and spectral fingerprints of ground features carried by the original spectral characteristics, and is used to initiate cross-modal information retrieval based on spectral information; at the same time, it selects the stereotyped spatial features output by the spatial branch. As the data source for key values, each relies on an independent learnable weight matrix. , Perform linear projection transformation to generate key features sequentially. AND value characteristics , It serves the function of spatial feature index matching. It preserves complete original spatial information materials such as multi-scale textures, outlines, and neighborhood distribution of ground features, which can be fused. After completing the projection of the three types of features, the query matrix will be... transpose of the bond matrix The inner product operation is performed, relying on the matrix inner product to quantify the correlation similarity between spectral features and spatial features at each location. To avoid excessively large inner product values due to high feature dimensionality, which could lead to gradient saturation of the activation function, a method is adopted... The inner product result is numerically scaled using a scaling factor to normalize the numerical distribution range of similarity. The scaled result is then fed into a Softmax function for normalization, transforming the unbounded similarity values into an attention weight distribution with a sum of 1. The weight coefficients intuitively reflect the degree of matching between the corresponding spatial information and the current spectral features. Finally, the normalized weights are used to adjust the value features. We perform a weighted summation operation to filter out effective spatial features with high correlation and fuse them into the spectral information, ultimately obtaining the spectral spatial interaction features after spectral fusion. .
[0042] In the second set of reverse spatial-spectral supplementary information computation branches, the feature sources of the query and key value are swapped, using native spatial features. With learnable projection matrix Mapping to generate query vectors , Using spatial structure information of ground features as a retrieval benchmark, suitable band information is matched in spectral features; then the original spectral features are... Passing through , Two sets of independent learnable projection matrix transformations generate key features. AND value characteristics , Serves as a retrieval index for full-band spectral information. Store all refined spectral details. Subsequent calculations will follow the standardized cross-attention logic, starting with... and Solving for the inner product similarity using... The scaling factor constrains the numerical range, and then spectral attention weights are generated from the spatial perspective through Softmax normalization. Finally, the weights are used to... Weighted aggregation embeds the selected effective spectral features into the original spatial features, generating spatial spectral interaction features that fuse spectral details. After two sets of complete calculations in both directions, and The two types of features enable bidirectional information exchange between the spectral reinforcement space and the spatially empowered spectrum, respectively. When combined, they form an overall interactive feature that integrates bidirectional spatial-spectral correlation information, thus completing the bidirectional adaptive fusion of cross-modal spatial-spectral information.
[0043] Based on the interaction features, the joint spatial-spectral discrimination features are obtained, specifically: ; ; ; in, It is a gated multilayer sensor.
[0044] An adaptive weighted fusion process is performed on the two types of spatial-spectral interaction features output by bidirectional cross-attention. Relying on a gated multilayer perceptron, the optimal contribution weights of the two types of features are autonomously learned, abandoning the traditional fusion method of fixed-coefficient concatenation, and generating a final fused feature that takes into account the dynamic ratio of spectral and spatial information. In the first step of feature stacking and concatenation, the spectral-spatial interaction features are concatenated through a Concat channel concatenation operation. Interaction characteristics with spatial spectrum By concatenating and integrating the features along the feature channel dimension, stacked features are obtained. ,in It is an interactive feature based on native spectral information, supplemented by matching and associating spatial texture contour information, focusing on preserving the refined spectral fingerprint features of ground features and including spatial constraint information. It is an interactive feature that takes the original spatial structure information as the core, integrates the spectral details of the corresponding bands, focuses on preserving the multi-scale spatial distribution characteristics of ground features, and supplements them with exclusive spectral identification information. This process fully incorporates all the spatial spectral information generated by the bidirectional interaction, providing a complete feature data source for the gating network to learn its weights. After entering the second step of solving the gating weights, the features are stacked. Sent to the door control multilayer sensor Internally, the gated multilayer perceptron consists of a multilayer nonlinear mapping network. It autonomously mines the differences in contribution of two types of interactive features to the final classification task based on the network's trainable parameters. It adaptively extracts suitable weight reference information for different land cover samples, outputs raw weight data after multilayer feature transformation, and then processes this data through a Softmax normalization function to generate a weight combination with values ranging from 0 to 1, where the sum of the two weight values is always equal to 1. The weight correspond Feature contribution ratio, weight correspond The proportion of feature contribution, in sample scenarios where the spectral features of ground objects are more recognizable. The numerical values will adaptively improve, especially in sample scenarios where spatial contour discrimination is more critical. The weights automatically increase, thus dynamically adjusting the weights according to the sample attributes. In the third step of the weighted fusion calculation, the generated adaptive weights are used to multiply the coefficients of the two types of interaction features respectively. Then, the weighted feature results are summed to obtain the final fused feature. This weighted calculation can flexibly adjust the proportion of spectral and spatial information based on the attributes of the land cover itself. For land covers with excellent spectral feature recognition, it automatically amplifies the weight coefficient of the dominant spectral features; for land covers that can only be distinguished by their spatial morphology, it increases the weight of the dominant spatial features. The final result is... The dynamic balancing of the spatial and spectral information effectively alleviates the feature confusion caused by different spectra of the same object and the same spectra of different objects, and condenses the integrated spatial-spectral fusion feature with stronger characterization ability, which can be directly connected to the classification module to complete the determination of land cover category.
[0045] The classifier is a multilayer perceptron.
[0046] The spatial-spectral joint discriminative features are input into the classifier, and the output is the predicted land cover category, specifically: ; in, This is the classification weight matrix. This is a bias term.
[0047] The joint spatial-spectral discrimination features obtained by gated adaptive weighted fusion in the previous steps As the data source for classification input, it relies on linear mapping operations combined with the Softmax normalization function to complete the probability prediction of land cover categories. The entire process incorporates the spatial-spectral fusion results, achieving the final mapping output from refined features to category prediction results. (The input is...) This integrated spatial-spectral feature is generated through bidirectional cross-attention information interaction and dynamic weighting with gated weights. It encapsulates the unique full-band spectral fingerprints and reflectance variation patterns of each land cover, integrates multi-scale spatial structure information such as land cover edge textures and neighborhood spatial arrangement, and optimizes the spatial-spectral information ratio through adaptive weights. It possesses extremely strong class discrimination capabilities, providing the classifier with sufficiently discriminative pre-feature data. During the linear transformation stage of the classification operation, the spatial-spectral fusion feature... With learnable classification weight matrix Perform matrix multiplication. This is the parameter matrix that the model continuously iterates and optimizes during the training phase. Different parameters within the matrix correspond to the association mapping weights between different land cover categories and input feature channels, used to quantify the contribution of various features to the land cover classification results. After the multiplication operation, a classification bias term is added. The bias term is used to correct the baseline value of linear operations and compensate for the calculation deviation caused by the overall feature offset. After the overall linear transformation, the original multi-channel spatial spectral features are mapped to a dimensional space matching the number of land cover categories, resulting in unnormalized raw category score vectors. Next, the raw scores output by the linear transformation are fed into the Softmax activation function for normalization. The Softmax function converts the unconstrained range of each category score into probability values between 0 and 1, and the sum of the probabilities corresponding to all categories is fixed at 1, ultimately generating... The predicted probability distribution of land cover categories is given. The label corresponding to the dimension with the largest value in the probability vector is the predicted land cover category of the sample. This completes the full-link prediction operation from spatial-spectral joint features to land cover classification results.
[0048] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0049] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hyperspectral image classification method based on cross-attention fusion, characterized in that, Includes the following steps: Hyperspectral images are preprocessed and heterogeneous samples are constructed to obtain independent spectral branch input data and spatial branch input data; Spectral features are extracted based on the input data of the spectral branch, and spatial features are extracted based on the input data of the spectral branch. Using the spectral and spatial features as input, interactive features are generated through a bidirectional cross-attention mechanism; spatial-spectral joint discrimination features are obtained based on the interactive features. Input the spatial-spectral joint identification features into the classifier and output the land cover category prediction results.
2. The hyperspectral image classification method based on cross-attention fusion according to claim 1, characterized in that, The hyperspectral image is preprocessed and heterogeneous samples are constructed to obtain independent spectral branch input data and spatial branch input data, specifically as follows: Dimensionality reduction and noise reduction are performed on hyperspectral images to suppress band redundancy and noise interference; Extract the continuous spectral curve of a single pixel or a small neighborhood as the input data for the spectral branch; Extract the spatial neighborhood map tiles after dimensionality reduction as input data for spatial branches.
3. The hyperspectral image classification method based on cross-attention fusion according to claim 1, characterized in that, Spectral features are extracted using a sequence feature extraction network based on a self-attention mechanism, and the first feature is calculated using the following formula. Layer encoder output: ; ; in, For the first The layer input features are obtained by mapping the spectral branch input data through the embedding layer; MHSA is a multi-head self-attention module; FFN is a feedforward neural network; LN is layer normalization; the final output... As a spectral feature.
4. The hyperspectral image classification method based on cross-attention fusion according to claim 3, characterized in that, The projection of query, key, and value, and the attention output in the multi-head self-attention module are calculated as follows: ; ; in, , , For learnable projection matrices, This is the scaling factor.
5. The hyperspectral image classification method based on cross-attention fusion according to claim 4, characterized in that, Based on the spectral branch input data, spatial features are extracted, specifically as follows: ; ; ; ; in, This represents a k×k convolution operation, and GAP stands for global average pooling.
6. The hyperspectral image classification method based on cross-attention fusion according to claim 5, characterized in that, Using the aforementioned spectral and spatial features as input, interactive features are generated through a bidirectional cross-attention mechanism, specifically: Using spectral features as queries and spatial features as keys, first-direction cross-attention calculation is performed to obtain spectral-spatial interaction features; Using spatial features as queries and spectral features as keys, a second-direction cross-attention calculation is performed to obtain spatial-spectral interaction features; The spectral spatial interaction feature and the spatial spectral interaction feature are combined to form the interaction feature.
7. The hyperspectral image classification method based on cross-attention fusion according to claim 6, characterized in that, The cross-attention between the first and second directions is calculated as follows: ; ; ; 。 8. The hyperspectral image classification method based on cross-attention fusion according to claim 7, characterized in that, Based on the interaction features, the spatial-spectral joint discrimination features are obtained, specifically: ; ; ; in, It is a gated multilayer sensor.
9. The hyperspectral image classification method based on cross-attention fusion according to claim 8, characterized in that, The classifier is a multilayer perceptron.
10. The hyperspectral image classification method based on cross-attention fusion according to claim 9, characterized in that, The spatial-spectral joint discriminative features are input into the classifier, and the output is the predicted land cover category, specifically: ; in, This is the classification weight matrix. This is a bias term.